IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model Training
Summary: IncrCP does incremental checkpointing for massive recommender models by recording per-iteration changed parameters and their indexes into independent chunk files. A 2-D chunk orchestration plus selective extraction and concatenation reduces I/O/dedup and yields up to 6.6× faster recovery and ~60% storage reduction. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Qingyin Lin (Peng Cheng Laboratory; Sun Yat-Sen University)
- 2. Jiangsu Du (Sun Yat-Sen University)
- 3. Rui Li (Peng Cheng Laboratory)
- 4. Zhiguang Chen (Sun Yat-Sen University)
- 5. Wenguang Chen (Peng Cheng Laboratory; Tsinghua University)
- 6. Nong Xiao (Sun Yat-Sen University)
BibTeX Citation
@article{lin_vldb25,
title = {{IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model Training}},
author = {Lin, Qingyin and Du, Jiangsu and Li, Rui and Chen, Zhiguang and Chen, Wenguang and Xiao, Nong},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {4},
pages = {1049--1062},
doi = {10.14778/3717755.3717765},
url = {https://doi.org/10.14778/3717755.3717765},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,954 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the Cloud | 2024 | VLDB | 5.7303405e-05 |
| 6,958 | Efficient Fault Tolerance for Recommendation Model Training via Erasure Coding | 2023 | VLDB | 5.7303405e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,556 | CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models | 2024 | SIGMOD |
| 2 | 9,164 | FEC: Efficient Deep Recommendation Model Training with Flexible Embedding Communication | 2023 | SIGMOD |
| 3 | 8,099 | Index Checkpoints for Instant Recovery in In-Memory Database Systems | 2022 | VLDB |
| 4 | 2,688 | Accelerating Recommendation System Training by Leveraging Popular Choices | 2022 | VLDB |
| 5 | 9,520 | Experimental Analysis of Large-scale Learnable Vector Storage Compression | 2024 | VLDB |
| 6 | 2,929 | StreamRec: A Real-Time Recommender System | 2011 | SIGMOD |
| 7 | 13,321 | The Limits of Graph Samplers for Training Inductive Recommender Systems | 2025 | VLDB |
| 8 | 8,907 | Scheduling Data Processing Pipelines for Incremental Training on MLP-based Recommendation Models | 2025 | SIGMOD |
| 9 | 6,958 | Efficient Fault Tolerance for Recommendation Model Training via Erasure Coding | 2023 | VLDB |
| 10 | 13,327 | DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems | 2025 | VLDB |