DBScholar

Back to papers

PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Summary: Design, implementation, and evaluation of PyTorch Distributed Data Parallel for scalable data-parallel training on GPUs. Unique practical optimizations—gradient bucketing, compute/communication overlap, and skipped gradient sync—achieving near-linear scaling to 256 GPUs. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hfe9e667121f66f81
Venue
VLDB
Year
2020
Pagerank
0.000169044
Overall Rank
523 | 96.49%
DOI
10.14778/3415478.3415530
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{li_vldb20,
        title = {{PyTorch Distributed: Experiences on Accelerating Data Parallel Training}},
        author = {Li, Shen and Zhao, Yanli and Varma, Rohan and Salpekar, Omkar and Noordhuis, Pieter and Li, Teng and Paszke, Adam and Smith, Jeff and Vaughan, Brian and Damania, Pritam and Chintala, Soumith},
        journal = {PVLDB},
        series = {{VLDB} '20},
        volume = {13},
        number = {12},
        pages = {3005--3018},
        doi = {10.14778/3415478.3415530},
        url = {https://doi.org/10.14778/3415478.3415530},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 29 of 29 citing papers.

Rank Citing Paper Year Venue Pagerank
2,247 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel 2023 VLDB 8.7597965e-05
2,281 NeutronStar: Distributed GNN Training with Hybrid Dependency Management 2022 SIGMOD 8.7021423e-05
2,460 Query Processing on Tensor Computation Runtimes 2022 VLDB 8.4308406e-05
2,521 HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework 2022 VLDB 8.3481558e-05
3,242 Towards Demystifying Serverless Machine Learning Training 2021 SIGMOD 7.4967268e-05
4,352 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 6.6396584e-05
4,949 HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training 2022 SIGMOD 6.3405861e-05
4,965 Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce 2021 SIGMOD 6.3354086e-05
5,011 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 6.3131083e-05
5,292 Tensor Relational Algebra for Distributed Machine Learning System Design 2021 VLDB 6.1911633e-05
5,999 BAGUA: Scaling up Distributed Learning with System Relaxations 2022 VLDB 5.9170896e-05
7,071 Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads 2024 VLDB 5.6056639e-05
8,257 SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 5.3661896e-05
8,562 Harmony: Overcoming the Hurdles of GPU Memory Capacity to Train Massive DNN Models on Commodity Servers 2022 VLDB 5.3127204e-05
8,820 FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement 2023 SIGMOD 5.269787e-05
8,896 mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs 2025 VLDB 5.2534908e-05
9,006 ANN Softmax: Acceleration of Extreme Classification Training 2022 VLDB 5.2375792e-05
9,153 Cerebro: A Layered Data Platform for Scalable Deep Learning 2021 CIDR 5.217676e-05
9,244 Effective and Efficient Distributed Temporal Graph Learning through Hotspot Memory Sharing 2025 VLDB 5.2032182e-05
9,564 Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning 2021 VLDB 5.1547835e-05
9,658 How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study 2024 VLDB 5.142891e-05
9,663 BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach 2023 SIGMOD 5.142891e-05
10,024 EinDecomp: Decomposition of Declaratively-Specified Machine Learning and Numerical Computations for Parallel Execution 2025 VLDB 5.091949e-05
10,447 DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization 2026 SIGMOD 4.9769913e-05
10,588 Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment 2026 SIGMOD 4.9769913e-05
11,258 GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research 2025 VLDB 4.9769913e-05
11,290 LobRA: Multi-tenant Fine-tuning over Heterogeneous Data 2025 VLDB 4.9769913e-05
11,297 Heta: Distributed Training of Heterogeneous Graph Neural Networks 2025 VLDB 4.9769913e-05
13,654 DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems 2025 VLDB -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 0 of 0 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Semantically Similar Papers