Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism
Summary: Galvatron automatically searches hybrid Transformer parallelism across multiple GPUs and memory budgets. A decision-tree decomposition/pruning stage combined with dynamic programming efficiently navigates the large strategy space, outperforming prior limited-parallelism systems in throughput. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Xupeng Miao (Carnegie Mellon University; Peking University)
- 2. Yujie Wang (Peking University)
- 3. Youhe Jiang (Peking University)
- 4. Chunan Shi (Peking University)
- 5. Xiaonan Nie (Peking University)
- 6. Hailin Zhang (Peking University)
- 7. Bin Cui (Peking University)
BibTeX Citation
@article{miao_vldb23,
title = {{Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism}},
author = {Miao, Xupeng and Wang, Yujie and Jiang, Youhe and Shi, Chunan and Nie, Xiaonan and Zhang, Hailin and Cui, Bin},
journal = {PVLDB},
series = {{VLDB} '23},
volume = {16},
number = {3},
pages = {470--479},
doi = {10.14778/3570690.3570697},
url = {https://doi.org/10.14778/3570690.3570697},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 14 of 14 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 521 | PyTorch Distributed: Experiences on Accelerating Data Parallel Training | 2020 | VLDB | 0.0001713368 |
| 1,157 | Cerebro: A Data System for Optimized Deep Learning Model Selection | 2020 | VLDB | 0.00011924049 |
| 2,485 | HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework | 2022 | VLDB | 8.5145736e-05 |
| 4,912 | HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training | 2022 | SIGMOD | 6.4481656e-05 |
| 4,956 | Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce | 2021 | SIGMOD | 6.4290135e-05 |
Previous
Page 1 / 1
Next