DBScholar

Back to papers

Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization

Summary: Malleus enables straggler-resilient hybrid training via per-GPU profiling and a planning algorithm that optimizes GPU groups, pipelines, layers, and data. It re-plans and migrates state on the fly to sustain stability, operating under dynamic straggler distributions and delivering 2.63–5.28x efficiency on LLMs up to 110B. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h7a4cd5cbab7f25f9
Venue
SIGMOD
Year
2025
Pagerank
4.9793485e-05
Overall Rank
11,189 | 24.78%
DOI
10.1145/3725322

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{li_sigmod25,
        title = {{Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization}},
        author = {Li, Haoyang and Fu, Fangcheng and Ge, Hao and Lin, Sheng and Wang, Xuanyu and Niu, Jiawen and Wang, Yujie and Zhang, Hailin and Nie, Xiaonan and Cui, Bin},
        series = {{SIGMOD} '25},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3725322},
        url = {https://dl.acm.org/doi/10.1145/3725322},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
2,184 Heterogeneity-aware Distributed Parameter Servers 2017 SIGMOD 8.8958335e-05
2,231 GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian Optimization 2024 VLDB 8.7982985e-05
2,246 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel 2023 VLDB 8.7637336e-05
2,348 The Dawn of Natural Language to SQL: Are We Fully Ready? 2024 VLDB 8.6009821e-05
3,601 Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines 2022 SIGMOD 7.1757877e-05
3,872 SketchML: Accelerating Distributed Machine Learning with Data Sketches 2018 SIGMOD 6.9540368e-05
4,351 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 6.642803e-05
4,589 ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models 2024 VLDB 6.5144711e-05
5,009 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 6.3158866e-05
5,381 GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning 2023 SIGMOD 6.1550396e-05
5,572 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0802555e-05
5,997 BAGUA: Scaling up Distributed Learning with System Relaxations 2022 VLDB 5.9198921e-05
7,242 Data and AI Model Markets: Opportunities for Data and Model Sharing, Discovery, and Integration 2023 VLDB 5.5773854e-05
7,978 Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines 2024 VLDB 5.4136835e-05
8,197 Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent 2023 VLDB 5.3794747e-05
8,812 FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement 2023 SIGMOD 5.2722828e-05
9,651 How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study 2024 VLDB 5.1453267e-05
9,656 BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach 2023 SIGMOD 5.1453267e-05
Previous Page 1 / 1 Next

Semantically Similar Papers