Back to papers
PyTorch Distributed: Experiences on Accelerating Data Parallel Training
Summary: Design, implementation, and evaluation of PyTorch Distributed Data Parallel for scalable data-parallel training on GPUs. Unique practical optimizations—gradient bucketing, compute/communication overlap, and skipped gradient sync—achieving near-linear scaling to 256 GPUs.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 12187
- Venue
- VLDB
- Year
- 2020
- Pagerank
- 0.00023881138
- Overall Rank
- 411 | 97.15%
- DOI
-
10.14778/3415478.3415530
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 28 of 28 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 2,678 |
HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework |
2022 |
VLDB |
8.3224016e-05 |
| 2,795 |
Towards Demystifying Serverless Machine Learning Training |
2021 |
SIGMOD |
8.1135606e-05 |
| 2,907 |
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel |
2023 |
VLDB |
7.9322286e-05 |
| 3,028 |
NeutronStar: Distributed GNN Training with Hybrid Dependency Management |
2022 |
SIGMOD |
7.6833093e-05 |
| 3,260 |
Query Processing on Tensor Computation Runtimes |
2022 |
VLDB |
7.3091312e-05 |
| 5,169 |
HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training |
2022 |
SIGMOD |
5.642415e-05 |
| 5,332 |
Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce |
2021 |
SIGMOD |
5.5640779e-05 |
| 5,730 |
BAGUA: Scaling up Distributed Learning with System Relaxations |
2022 |
VLDB |
5.3476346e-05 |
| 5,836 |
Tensor Relational Algebra for Distributed Machine Learning System Design |
2021 |
VLDB |
5.3079723e-05 |
| 6,361 |
Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism |
2023 |
VLDB |
5.0903244e-05 |
| 7,143 |
Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity |
2024 |
VLDB |
4.8143774e-05 |
| 8,117 |
SDPipe: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training |
2023 |
VLDB |
4.5788485e-05 |
| 8,518 |
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs |
2025 |
VLDB |
4.4893996e-05 |
| 8,605 |
Harmony: Overcoming the Hurdles of GPU Memory Capacity to Train Massive DNN Models on Commodity Servers |
2022 |
VLDB |
4.4813623e-05 |
| 8,709 |
ANN Softmax: Acceleration of Extreme Classification Training |
2022 |
VLDB |
4.4584416e-05 |
| 8,807 |
FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement |
2023 |
SIGMOD |
4.4413307e-05 |
| 8,864 |
Cerebro: A Layered Data Platform for Scalable Deep Learning |
2021 |
CIDR |
4.4283952e-05 |
| 9,225 |
Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning |
2021 |
VLDB |
4.3656789e-05 |
| 9,324 |
How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study |
2024 |
VLDB |
4.351469e-05 |
| 9,331 |
BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach |
2023 |
SIGMOD |
4.351469e-05 |
| 9,603 |
Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads |
2024 |
VLDB |
4.3136057e-05 |
| 9,693 |
EinDecomp: Decomposition of Declaratively-Specified Machine Learning and Numerical Computations for Parallel Execution |
2025 |
VLDB |
4.2984337e-05 |
| 10,089 |
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,589 |
GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research |
2025 |
VLDB |
4.1905499e-05 |
| 10,634 |
LobRA: Multi-tenant Fine-tuning over Heterogeneous Data |
2025 |
VLDB |
4.1905499e-05 |
| 10,646 |
Heta: Distributed Training of Heterogeneous Graph Neural Networks |
2025 |
VLDB |
4.1905499e-05 |
| 10,664 |
Effective and Efficient Distributed Temporal Graph Learning through Hotspot Memory Sharing |
2025 |
VLDB |
4.1905499e-05 |
| 13,136 |
DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems |
2025 |
VLDB |
- |
Outgoing Citations (Sorted by Pagerank)
Showing 0 of 0 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 6,482 |
Dynamic Parameter Allocation in Parameter Servers |
2020 |
VLDB |
5.0396586e-05 |
| 1,102 |
Large Graph Convolutional Network Training with GPU-Oriented Data Communication Architecture |
2021 |
VLDB |
0.00014011556 |
| 5,383 |
Parallel Training of Knowledge Graph Embedding Models: A Comparison of Techniques |
2022 |
VLDB |
5.5357645e-05 |
| 9,400 |
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism |
2025 |
VLDB |
4.3399748e-05 |
| 8,117 |
SDPipe: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training |
2023 |
VLDB |
4.5788485e-05 |
| 9,596 |
Scalable Graph Convolutional Network Training on Distributed-Memory Systems |
2023 |
VLDB |
4.3150788e-05 |
| 8,731 |
TensorSocket: Shared Data Loading for Deep Learning Training |
2026 |
SIGMOD |
4.4520434e-05 |
| 9,964 |
Distributed Learning of Fully Connected Neural Networks using Independent Subnet Training |
2022 |
VLDB |
4.2229209e-05 |
| 5,332 |
Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce |
2021 |
SIGMOD |
5.5640779e-05 |
| 2,907 |
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel |
2023 |
VLDB |
7.9322286e-05 |