DBScholar

Back to papers

PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Summary: Design, implementation, and evaluation of PyTorch Distributed Data Parallel for scalable data-parallel training on GPUs. Unique practical optimizations—gradient bucketing, compute/communication overlap, and skipped gradient sync—achieving near-linear scaling to 256 GPUs. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hfe9e667121f66f81
Venue
VLDB
Year
2020
Pagerank
0.00016912378
Overall Rank
522 | 96.50%
DOI
10.14778/3415478.3415530

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{li_vldb20,
        title = {{PyTorch Distributed: Experiences on Accelerating Data Parallel Training}},
        author = {Li, Shen and Zhao, Yanli and Varma, Rohan and Salpekar, Omkar and Noordhuis, Pieter and Li, Teng and Paszke, Adam and Smith, Jeff and Vaughan, Brian and Damania, Pritam and Chintala, Soumith},
        journal = {PVLDB},
        series = {{VLDB} '20},
        volume = {13},
        number = {12},
        pages = {3005--3018},
        doi = {10.14778/3415478.3415530},
        url = {https://doi.org/10.14778/3415478.3415530},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 29 of 29 citing papers.

Rank Citing Paper Year Venue Pagerank
2,246 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel 2023 VLDB 8.7637336e-05
2,279 NeutronStar: Distributed GNN Training with Hybrid Dependency Management 2022 SIGMOD 8.7062637e-05
2,460 Query Processing on Tensor Computation Runtimes 2022 VLDB 8.4348335e-05
2,519 HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework 2022 VLDB 8.3521095e-05
3,240 Towards Demystifying Serverless Machine Learning Training 2021 SIGMOD 7.5002772e-05
4,351 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 6.642803e-05
4,945 HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training 2022 SIGMOD 6.3435891e-05
4,963 Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce 2021 SIGMOD 6.3384091e-05
5,009 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 6.3158866e-05
5,289 Tensor Relational Algebra for Distributed Machine Learning System Design 2021 VLDB 6.1940481e-05
5,997 BAGUA: Scaling up Distributed Learning with System Relaxations 2022 VLDB 5.9198921e-05
7,069 Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads 2024 VLDB 5.6083188e-05
8,251 SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 5.3687311e-05
8,555 Harmony: Overcoming the Hurdles of GPU Memory Capacity to Train Massive DNN Models on Commodity Servers 2022 VLDB 5.3152366e-05
8,812 FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement 2023 SIGMOD 5.2722828e-05
8,887 mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs 2025 VLDB 5.2559789e-05
8,998 ANN Softmax: Acceleration of Extreme Classification Training 2022 VLDB 5.2400492e-05
9,144 Cerebro: A Layered Data Platform for Scalable Deep Learning 2021 CIDR 5.2201471e-05
9,234 Effective and Efficient Distributed Temporal Graph Learning through Hotspot Memory Sharing 2025 VLDB 5.2056825e-05
9,556 Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning 2021 VLDB 5.1572248e-05
9,651 How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study 2024 VLDB 5.1453267e-05
9,656 BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach 2023 SIGMOD 5.1453267e-05
10,019 EinDecomp: Decomposition of Declaratively-Specified Machine Learning and Numerical Computations for Parallel Execution 2025 VLDB 5.0943606e-05
10,435 DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization 2026 SIGMOD 4.9793485e-05
10,577 Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment 2026 SIGMOD 4.9793485e-05
11,250 GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research 2025 VLDB 4.9793485e-05
11,282 LobRA: Multi-tenant Fine-tuning over Heterogeneous Data 2025 VLDB 4.9793485e-05
11,289 Heta: Distributed Training of Heterogeneous Graph Neural Networks 2025 VLDB 4.9793485e-05
13,648 DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems 2025 VLDB -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 0 of 0 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Semantically Similar Papers