PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Summary: Introduces PyTorch Fully Sharded Data Parallel (FSDP), an industry-grade, non-intrusive sharding framework co-designed with PyTorch internals (Tensor, dispatcher, CUDA allocator) to enable training of much larger models than DDP. FSDP bundles memory and communication optimizations across hardware to achieve near-linear TFLOPS scalability and DDP-comparable throughput while drastically reducing memory footprint. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yanli Zhao (Meta)
- 2. Andrew Gu (Meta)
- 3. Rohan Varma (Meta)
- 4. Liang Luo (Meta)
- 5. Chien-Chin Huang (Meta)
- 6. Min Xu (Meta)
- 7. Less Wright (Meta)
- 8. Hamid Shojanazeri (Meta)
- 9. Myle Ott (Meta)
- 10. Sam Shleifer (Meta)
- 11. Alban Desmaison (Meta)
- 12. Can Balioglu (Meta)
- 13. Pritam Damania (Meta)
- 14. Bernard Nguyen (Meta)
- 15. Geeta Chauhan (Meta)
- 16. Yuchen Hao (Meta)
- 17. Ajit Mathews (Meta)
- 18. Shen Li (Meta)
BibTeX Citation
@article{zhao_vldb23,
title = {{PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel}},
author = {Zhao, Yanli and Gu, Andrew and Varma, Rohan and Luo, Liang and Huang, Chien-Chin and Xu, Min and Wright, Less and Shojanazeri, Hamid and Ott, Myle and Shleifer, Sam and Desmaison, Alban and Balioglu, Can and Damania, Pritam and Nguyen, Bernard and Chauhan, Geeta and Hao, Yuchen and Mathews, Ajit and Li, Shen},
journal = {PVLDB},
series = {{VLDB} '23},
volume = {16},
number = {12},
pages = {3848--3860},
doi = {10.14778/3611540.3611569},
url = {https://doi.org/10.14778/3611540.3611569},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 521 | PyTorch Distributed: Experiences on Accelerating Data Parallel Training | 2020 | VLDB | 0.0001713368 |
| 3,530 | MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud | 2023 | VLDB | 7.3379782e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,799 | Scalable Asynchronous Gradient Descent Optimization for Out-of-Core Models | 2017 | VLDB |
| 2 | 10,113 | Distributed Learning of Fully Connected Neural Networks using Independent Subnet Training | 2022 | VLDB |
| 3 | 2,640 | Scalable and Efficient Full-Graph GNN Training for Large Graphs | 2023 | SIGMOD |
| 4 | 10,219 | DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization | 2026 | SIGMOD |
| 5 | 3,764 | FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline | 2023 | VLDB |
| 6 | 4,956 | Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce | 2021 | SIGMOD |
| 7 | 9,414 | Model-Parallel Model Selection for Deep Learning Systems | 2021 | SIGMOD |
| 8 | 8,905 | TensorSocket: Shared Data Loading for Deep Learning Training | 2026 | SIGMOD |
| 9 | 8,101 | SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training | 2023 | VLDB |
| 10 | 521 | PyTorch Distributed: Experiences on Accelerating Data Parallel Training | 2020 | VLDB |