DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the Cloud
Summary: Elastic framework for DLRM training that builds a DLRM-specific resource–performance model and a three-stage heuristic to auto-allocate and dynamically adjust GPU/CPU/memory to boost utilization. Adds cloud-instability mitigation; deployed at AntGroup with 31% lower JCT and +15% CPU/+20% memory. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Qinlong Wang (Independent)
- 2. Tingfeng Lan (Sichuan University)
- 3. Yinghao Tang (Sichuan University)
- 4. Bo Sang (Independent)
- 5. Ziling Huang (Sichuan University)
- 6. Yiheng Du (Sichuan University)
- 7. Haitao Zhang (Independent)
- 8. Jian Sha (Independent)
- 9. Hui Lu (University of Texas Arlington)
- 10. Yuanchun Zhou (Chinese Academy of Sciences)
- 11. Ke Zhang (Independent)
- 12. Mingjie Tang (Sichuan University)
BibTeX Citation
@article{wang_vldb24,
title = {{DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the Cloud}},
author = {Wang, Qinlong and Lan, Tingfeng and Tang, Yinghao and Sang, Bo and Huang, Ziling and Du, Yiheng and Zhang, Haitao and Sha, Jian and Lu, Hui and Zhou, Yuanchun and Zhang, Ke and Tang, Mingjie},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {12},
pages = {4130--4144},
doi = {10.14778/3685800.3685832},
url = {https://doi.org/10.14778/3685800.3685832},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,804 | IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model Training | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,298 | GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian Optimization | 2024 | VLDB | 8.7886538e-05 |
| 2,688 | Accelerating Recommendation System Training by Leveraging Popular Choices | 2022 | VLDB | 8.2564305e-05 |
| 3,764 | FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline | 2023 | VLDB | 7.1446931e-05 |
| 4,049 | Resource Elasticity for Large-Scale Machine Learning | 2015 | SIGMOD | 6.9369379e-05 |
| 4,415 | Eigen: End-to-end Resource Optimization for Large-Scale Databases on the Cloud | 2023 | VLDB | 6.7152625e-05 |
| 8,034 | Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent | 2023 | VLDB | 5.502946e-05 |
Previous
Page 1 / 1
Next