Agile and Accurate CTR Prediction Model Training for Massive-Scale Online Advertising Systems
Summary: Industrial-scale CTR training on GPUs with quantization to handle hundreds of billions of features and samples. Quantization enlarges embedding capacity without extra storage, enabling agile deployment and yielding 1% revenue lift and 1.8% relative CTR gain in production. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Zhiqiang Xu (Baidu)
- 2. Dong Li (Baidu)
- 3. Weijie Zhao (Baidu)
- 4. Xing Shen (Baidu)
- 5. Tianbo Huang (Baidu)
- 6. Xiaoyun Li (Baidu)
- 7. Ping Li (Baidu)
BibTeX Citation
@inproceedings{xu_sigmod21,
title = {{Agile and Accurate CTR Prediction Model Training for Massive-Scale Online Advertising Systems}},
author = {Xu, Zhiqiang and Li, Dong and Zhao, Weijie and Shen, Xing and Huang, Tianbo and Li, Xiaoyun and Li, Ping},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457236},
url = {https://dl.acm.org/doi/10.1145/3448016.3457236},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,912 | HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training | 2022 | SIGMOD | 6.4481656e-05 |
| 9,520 | Experimental Analysis of Large-scale Learnable Vector Storage Compression | 2024 | VLDB | 5.2561283e-05 |
| 9,556 | CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models | 2024 | SIGMOD | 5.2528121e-05 |
| 9,881 | The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format | 2024 | SIGMOD | 5.2040783e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,445 | Sketching Linear Classifiers over Data Streams | 2018 | SIGMOD | 7.4089141e-05 |
| 4,083 | SketchML: Accelerating Distributed Machine Learning with Data Sketches | 2018 | SIGMOD | 6.9160949e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,907 | Scheduling Data Processing Pipelines for Incremental Training on MLP-based Recommendation Models | 2025 | SIGMOD |
| 2 | 2,199 | Real-time Targeted Influence Maximization for Online Advertisements | 2015 | VLDB |
| 3 | 1,863 | ByteGNN: Efficient Graph Neural Network Training at Large Scale | 2022 | VLDB |
| 4 | 6,954 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the Cloud | 2024 | VLDB |
| 5 | 5,143 | An Online Cost Sensitive Decision-Making Method in Crowdsourcing Systems | 2013 | SIGMOD |
| 6 | 2,688 | Accelerating Recommendation System Training by Leveraging Popular Choices | 2022 | VLDB |
| 7 | 8,813 | Efficient and Effective Algorithms for Revenue Maximization in Social Advertising | 2021 | SIGMOD |
| 8 | 8,034 | Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent | 2023 | VLDB |
| 9 | 9,215 | Deep Query Optimization | 2019 | SIGMOD |
| 10 | 13,725 | Data Management and Mining in Internet Ad Systems | 2010 | VLDB |