DBScholar

Back to papers

Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism

Summary: Galvatron automatically searches hybrid Transformer parallelism across multiple GPUs and memory budgets. A decision-tree decomposition/pruning stage combined with dynamic programming efficiently navigates the large strategy space, outperforming prior limited-parallelism systems in throughput. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
13492
Venue
VLDB
Year
2023
Pagerank
6.4065691e-05
Overall Rank
5,003 | 65.68%
DOI
10.14778/3570690.3570697

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{miao_vldb23,
        title = {{Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism}},
        author = {Miao, Xupeng and Wang, Yujie and Jiang, Youhe and Shi, Chunan and Nie, Xiaonan and Zhang, Hailin and Cui, Bin},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {3},
        pages = {470--479},
        doi = {10.14778/3570690.3570697},
        url = {https://doi.org/10.14778/3570690.3570697},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 14 of 14 citing papers.

Rank Citing Paper Year Venue Pagerank
3,536 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.3297343e-05
6,900 Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving 2025 SIGMOD 5.7430032e-05
7,075 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 5.7099047e-05
8,034 Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent 2023 VLDB 5.502946e-05
8,101 SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 5.4870581e-05
8,883 FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement 2023 SIGMOD 5.351513e-05
9,475 BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach 2023 SIGMOD 5.2634238e-05
9,834 EinDecomp: Decomposition of Declaratively-Specified Machine Learning and Numerical Computations for Parallel Execution 2025 VLDB 5.2112879e-05
9,880 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.2040783e-05
10,219 DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization 2026 SIGMOD 5.093636e-05
10,380 Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment 2026 SIGMOD 5.093636e-05
10,769 Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization 2025 SIGMOD 5.093636e-05
10,880 LobRA: Multi-tenant Fine-tuning over Heterogeneous Data 2025 VLDB 5.093636e-05
11,266 DARKER: Efficient Transformer with Data-driven Attention Mechanism for Time Series 2024 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers