DBScholar

Back to papers

Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism

Summary: Galvatron automatically searches hybrid Transformer parallelism across multiple GPUs and memory budgets. A decision-tree decomposition/pruning stage combined with dynamic programming efficiently navigates the large strategy space, outperforming prior limited-parallelism systems in throughput. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hea38cacf948371f7
Venue
VLDB
Year
2023
Pagerank
6.3158866e-05
Overall Rank
5,009 | 66.33%
DOI
10.14778/3570690.3570697

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{miao_vldb23,
        title = {{Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism}},
        author = {Miao, Xupeng and Wang, Yujie and Jiang, Youhe and Shi, Chunan and Nie, Xiaonan and Zhang, Hailin and Cui, Bin},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {3},
        pages = {470--479},
        doi = {10.14778/3570690.3570697},
        url = {https://doi.org/10.14778/3570690.3570697},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 14 of 14 citing papers.

Rank Citing Paper Year Venue Pagerank
3,250 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.4938546e-05
4,351 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 6.642803e-05
6,166 Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving 2025 SIGMOD 5.8631131e-05
8,197 Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent 2023 VLDB 5.3794747e-05
8,251 SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 5.3687311e-05
8,812 FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement 2023 SIGMOD 5.2722828e-05
9,656 BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach 2023 SIGMOD 5.1453267e-05
10,019 EinDecomp: Decomposition of Declaratively-Specified Machine Learning and Numerical Computations for Parallel Execution 2025 VLDB 5.0943606e-05
10,042 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.0921006e-05
10,435 DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization 2026 SIGMOD 4.9793485e-05
10,577 Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment 2026 SIGMOD 4.9793485e-05
11,189 Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization 2025 SIGMOD 4.9793485e-05
11,282 LobRA: Multi-tenant Fine-tuning over Heterogeneous Data 2025 VLDB 4.9793485e-05
11,594 DARKER: Efficient Transformer with Data-driven Attention Mechanism for Time Series 2024 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers