DBScholar

Back to papers

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Summary: Introduces PyTorch Fully Sharded Data Parallel (FSDP), an industry-grade, non-intrusive sharding framework co-designed with PyTorch internals (Tensor, dispatcher, CUDA allocator) to enable training of much larger models than DDP. FSDP bundles memory and communication optimizations across hardware to achieve near-linear TFLOPS scalability and DDP-comparable throughput while drastically reducing memory footprint. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
h0ecbdae73a7b577f
Venue
VLDB
Year
2023
Pagerank
8.7637336e-05
Overall Rank
2,246 | 84.91%
DOI
10.14778/3611540.3611569

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{zhao_vldb23,
        title = {{PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel}},
        author = {Zhao, Yanli and Gu, Andrew and Varma, Rohan and Luo, Liang and Huang, Chien-Chin and Xu, Min and Wright, Less and Shojanazeri, Hamid and Ott, Myle and Shleifer, Sam and Desmaison, Alban and Balioglu, Can and Damania, Pritam and Nguyen, Bernard and Chauhan, Geeta and Hao, Yuchen and Mathews, Ajit and Li, Shen},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {12},
        pages = {3848--3860},
        doi = {10.14778/3611540.3611569},
        url = {https://doi.org/10.14778/3611540.3611569},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 13 of 13 citing papers.

Rank Citing Paper Year Venue Pagerank
3,250 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.4938546e-05
8,887 mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs 2025 VLDB 5.2559789e-05
10,042 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.0921006e-05
10,435 DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization 2026 SIGMOD 4.9793485e-05
10,577 Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment 2026 SIGMOD 4.9793485e-05
10,659 Mixtera: A Data Plane for Foundation Model Training 2026 SIGMOD 4.9793485e-05
10,861 UniTG: A Unified System for Efficient and Seamless Textual Graph Learning 2026 VLDB 4.9793485e-05
11,124 Sequoia: An Accessible and Extensible Framework for Privacy-Preserving Machine Learning over Distributed Data 2025 SIGMOD 4.9793485e-05
11,189 Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization 2025 SIGMOD 4.9793485e-05
11,250 GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research 2025 VLDB 4.9793485e-05
11,282 LobRA: Multi-tenant Fine-tuning over Heterogeneous Data 2025 VLDB 4.9793485e-05
11,328 Evoschema: Towards Text-To-Sql Robustness Against Schema Evolution 2025 VLDB 4.9793485e-05
13,648 DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems 2025 VLDB -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
522 PyTorch Distributed: Experiences on Accelerating Data Parallel Training 2020 VLDB 0.00016912378
3,450 MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud 2023 VLDB 7.2921762e-05
Previous Page 1 / 1 Next

Semantically Similar Papers