DBScholar

Back to papers

DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the Cloud

Summary: Elastic framework for DLRM training that builds a DLRM-specific resource–performance model and a three-stage heuristic to auto-allocate and dynamically adjust GPU/CPU/memory to boost utilization. Adds cloud-instability mitigation; deployed at AntGroup with 31% lower JCT and +15% CPU/+20% memory. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
13800
Venue
VLDB
Year
2024
Pagerank
5.7303405e-05
Overall Rank
6,954 | 52.30%
DOI
10.14778/3685800.3685832

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{wang_vldb24,
        title = {{DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the Cloud}},
        author = {Wang, Qinlong and Lan, Tingfeng and Tang, Yinghao and Sang, Bo and Huang, Ziling and Du, Yiheng and Zhang, Haitao and Sha, Jian and Lu, Hui and Zhou, Yuanchun and Zhang, Ke and Tang, Mingjie},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {12},
        pages = {4130--4144},
        doi = {10.14778/3685800.3685832},
        url = {https://doi.org/10.14778/3685800.3685832},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 6 of 6 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers