In-Database Machine Learning with CorgiPile: Stochastic Gradient Descent without Full Data Shuffle
Summary: Proposes CorgiPile, a hierarchical data shuffling method for in-database SGD that avoids full shuffles yet preserves convergence. Systematic study of existing shuffles, convergence theory, and PostgreSQL integration via three new operators; achieves 1.6–12.8× speedups over MADlib/Bismarck on HDD/SSD. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Lijie Xu (Chinese Academy of Sciences; ETH Zurich)
- 2. Shuang Qiu (University of Chicago)
- 3. Binhang Yuan (ETH Zurich)
- 4. Jiawei Jiang (ETH Zurich)
- 5. Cedric Renggli (ETH Zurich)
- 6. Shaoduo Gan (ETH Zurich)
- 7. Kaan Kara (ETH Zurich)
- 8. Guoliang Li (Tsinghua University)
- 9. Ji Liu (Kwai Inc.)
- 10. Wentao Wu (Microsoft)
- 11. Jieping Ye (University of Michigan)
- 12. Ce Zhang (ETH Zurich)
BibTeX Citation
@inproceedings{xu_sigmod22,
title = {{In-Database Machine Learning with CorgiPile: Stochastic Gradient Descent without Full Data Shuffle}},
author = {Xu, Lijie and Qiu, Shuang and Yuan, Binhang and Jiang, Jiawei and Renggli, Cedric and Gan, Shaoduo and Kara, Kaan and Li, Guoliang and Liu, Ji and Wu, Wentao and Ye, Jieping and Zhang, Ce},
series = {{SIGMOD} '22},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3514221.3526150},
url = {https://dl.acm.org/doi/10.1145/3514221.3526150},
year = {2022}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,470 | Powering In-Database Dynamic Model Slicing for Structured Data Analytics | 2024 | VLDB | 5.2634238e-05 |
| 10,114 | Towards Communication-efficient Vertical Federated Learning Training via Cache-enabled Local Updates | 2022 | VLDB | 5.1319012e-05 |
| 10,386 | NeurStore: Efficient In-database Deep Learning Model Management System | 2026 | SIGMOD | 5.093636e-05 |
| 10,844 | GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research | 2025 | VLDB | 5.093636e-05 |
| 11,209 | Database Native Model Selection: Harnessing Deep Neural Networks in Database Systems | 2024 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,446 | Analyzing and Mitigating Data Stalls in DNN Training | 2021 | VLDB |
| 2 | 6,271 | Is Your Learned Query Optimizer Behaving As You Expect? A Machine Learning Perspective | 2024 | VLDB |
| 3 | 6,708 | Serving Deep Learning Models with Deduplication from Relational Databases | 2022 | VLDB |
| 4 | 8,040 | PerfGuard: Deploying ML-for-Systems without Performance Regressions, Almost! | 2021 | VLDB |
| 5 | 6,485 | Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent | 2019 | SIGMOD |
| 6 | 9,371 | Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning | 2021 | VLDB |
| 7 | 9,962 | Structure-Aware Machine Learning over Multi-Relational Databases | 2021 | SIGMOD |
| 8 | 7,112 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB |
| 9 | 6,143 | ColumnML: Column-Store Machine Learning with On-The-Fly Data Transformation | 2019 | VLDB |
| 10 | 4,799 | Scalable Asynchronous Gradient Descent Optimization for Out-of-Core Models | 2017 | VLDB |