Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
Summary: Mitigates data-induced imbalances in Transformer training—uneven sequence-length sampling and packing mismatch between attention time (quadratic) and memory (linear)—by jointly optimizing parallel strategy and data assignment. Hydraulis applies dynamic heterogeneous parallelism and a two-stage data assignment to balance intra- and inter-replica workloads, boosting throughput 1.32–2.66×. (summarized by gpt-5-mini on Feb 11 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Haoyang Li (Peking University)
- 2. Fangcheng Fu (Shanghai Jiao Tong University)
- 3. Sheng Lin (Peking University)
- 4. Hao Ge (Peking University)
- 5. Xuanyu Wang (Peking University)
- 6. Jiawen Niu (Peking University)
- 7. Jinbao Xue (Tencent)
- 8. Yangyu Tao (Tencent)
- 9. Di Wang (Tencent)
- 10. Jie Jiang (Tencent)
- 11. Bin Cui (Peking University)
BibTeX Citation
@inproceedings{li_sigmod26,
title = {{Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment}},
author = {Li, Haoyang and Fu, Fangcheng and Lin, Sheng and Ge, Hao and Wang, Xuanyu and Niu, Jiawen and Xue, Jinbao and Tao, Yangyu and Wang, Di and Jiang, Jie and Cui, Bin},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3769802},
url = {https://dl.acm.org/doi/10.1145/3769802},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 521 | PyTorch Distributed: Experiences on Accelerating Data Parallel Training | 2020 | VLDB | 0.0001713368 |
| 2,473 | PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel | 2023 | VLDB | 8.5326287e-05 |
| 5,003 | Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism | 2023 | VLDB | 6.4065691e-05 |
| 9,880 | MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training | 2025 | SIGMOD | 5.2040783e-05 |
| 10,769 | Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization | 2025 | SIGMOD | 5.093636e-05 |
| 10,880 | LobRA: Multi-tenant Fine-tuning over Heterogeneous Data | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next