DBScholar

Back to papers

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Summary: Introduces PyTorch Fully Sharded Data Parallel (FSDP), an industry-grade, non-intrusive sharding framework co-designed with PyTorch internals (Tensor, dispatcher, CUDA allocator) to enable training of much larger models than DDP. FSDP bundles memory and communication optimizations across hardware to achieve near-linear TFLOPS scalability and DDP-comparable throughput while drastically reducing memory footprint. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
h0ecbdae73a7b577f
Venue
VLDB
Year
2023
Pagerank
8.7597965e-05
Overall Rank
2,247 | 84.90%
DOI
10.14778/3611540.3611569
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{zhao_vldb23,
        title = {{PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel}},
        author = {Zhao, Yanli and Gu, Andrew and Varma, Rohan and Luo, Liang and Huang, Chien-Chin and Xu, Min and Wright, Less and Shojanazeri, Hamid and Ott, Myle and Shleifer, Sam and Desmaison, Alban and Balioglu, Can and Damania, Pritam and Nguyen, Bernard and Chauhan, Geeta and Hao, Yuchen and Mathews, Ajit and Li, Shen},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {12},
        pages = {3848--3860},
        doi = {10.14778/3611540.3611569},
        url = {https://doi.org/10.14778/3611540.3611569},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 13 of 13 citing papers.

Rank Citing Paper Year Venue Pagerank
3,236 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.4996147e-05
8,896 mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs 2025 VLDB 5.2534908e-05
10,047 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.0896901e-05
10,447 DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization 2026 SIGMOD 4.9769913e-05
10,588 Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment 2026 SIGMOD 4.9769913e-05
10,670 Mixtera: A Data Plane for Foundation Model Training 2026 SIGMOD 4.9769913e-05
10,870 UniTG: A Unified System for Efficient and Seamless Textual Graph Learning 2026 VLDB 4.9769913e-05
11,133 Sequoia: An Accessible and Extensible Framework for Privacy-Preserving Machine Learning over Distributed Data 2025 SIGMOD 4.9769913e-05
11,198 Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization 2025 SIGMOD 4.9769913e-05
11,258 GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research 2025 VLDB 4.9769913e-05
11,290 LobRA: Multi-tenant Fine-tuning over Heterogeneous Data 2025 VLDB 4.9769913e-05
11,336 Evoschema: Towards Text-To-Sql Robustness Against Schema Evolution 2025 VLDB 4.9769913e-05
13,654 DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems 2025 VLDB -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers