PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Summary: Introduces PyTorch Fully Sharded Data Parallel (FSDP), an industry-grade, non-intrusive sharding framework co-designed with PyTorch internals (Tensor, dispatcher, CUDA allocator) to enable training of much larger models than DDP. FSDP bundles memory and communication optimizations across hardware to achieve near-linear TFLOPS scalability and DDP-comparable throughput while drastically reducing memory footprint. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yanli Zhao (Meta)
- 2. Andrew Gu (Meta)
- 3. Rohan Varma (Meta)
- 4. Liang Luo (Meta)
- 5. Chien-Chin Huang (Meta)
- 6. Min Xu (Meta)
- 7. Less Wright (Meta)
- 8. Hamid Shojanazeri (Meta)
- 9. Myle Ott (Meta)
- 10. Sam Shleifer (Meta)
- 11. Alban Desmaison (Meta)
- 12. Can Balioglu (Meta)
- 13. Pritam Damania (Meta)
- 14. Bernard Nguyen (Meta)
- 15. Geeta Chauhan (Meta)
- 16. Yuchen Hao (Meta)
- 17. Ajit Mathews (Meta)
- 18. Shen Li (Meta)
BibTeX Citation
@article{zhao_vldb23,
title = {{PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel}},
author = {Zhao, Yanli and Gu, Andrew and Varma, Rohan and Luo, Liang and Huang, Chien-Chin and Xu, Min and Wright, Less and Shojanazeri, Hamid and Ott, Myle and Shleifer, Sam and Desmaison, Alban and Balioglu, Can and Damania, Pritam and Nguyen, Bernard and Chauhan, Geeta and Hao, Yuchen and Mathews, Ajit and Li, Shen},
journal = {PVLDB},
series = {{VLDB} '23},
volume = {16},
number = {12},
pages = {3848--3860},
doi = {10.14778/3611540.3611569},
url = {https://doi.org/10.14778/3611540.3611569},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 13 of 13 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 522 | PyTorch Distributed: Experiences on Accelerating Data Parallel Training | 2020 | VLDB | 0.00016912378 |
| 3,450 | MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud | 2023 | VLDB | 7.2921762e-05 |
Previous
Page 1 / 1
Next