Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent
Summary: Proposes tuple-oriented compression (TOC) for mini-batch SGD, preserving tuple boundaries while using LZW-inspired coding. Enables compressed-domain matrix operations on TOC, delivering up to 51x compression and 10.2x speedups for MGD workloads with no decompression overhead. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Fengan Li (University of Wisconsin)
- 2. Lingjiao Chen (University of Wisconsin)
- 3. Yijing Zeng (University of Wisconsin)
- 4. Arun Kumar (University of California San Diego)
- 5. Xi Wu (University of Wisconsin)
- 6. Jeffrey F. Naughton (University of Wisconsin)
- 7. Jignesh M. Patel (University of Wisconsin)
BibTeX Citation
@inproceedings{li_sigmod19,
title = {{Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent}},
author = {Li, Fengan and Chen, Lingjiao and Zeng, Yijing and Kumar, Arun and Wu, Xi and Naughton, Jeffrey F. and Patel, Jignesh M.},
series = {{SIGMOD} '19},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3299869.3300070},
url = {https://dl.acm.org/doi/10.1145/3299869.3300070},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,157 | Cerebro: A Data System for Optimized Deep Learning Model Selection | 2020 | VLDB | 0.00011924049 |
| 2,273 | SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging | 2021 | SIGMOD | 8.8230899e-05 |
| 4,067 | Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches | 2021 | VLDB | 6.9293511e-05 |
| 7,687 | ExDRa: Exploratory Data Science on Federated Raw Data | 2021 | SIGMOD | 5.5671645e-05 |
| 8,794 | AWARE: Workload-aware, Redundancy-exploiting Linear Algebra | 2023 | SIGMOD | 5.370464e-05 |
| 8,979 | Cerebro: A Layered Data Platform for Scalable Deep Learning | 2021 | CIDR | 5.3399615e-05 |
| 10,584 | QStore: Quantization-Aware Compressed Model Storage | 2026 | VLDB | 5.093636e-05 |
| 10,589 | Morphing-based Compression for Data-centric ML Pipelines | 2026 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 18 of 18 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,190 | Progressive Compressed Records: Taking a Byte out of Deep Learning Data | 2021 | VLDB |
| 2 | 8,794 | AWARE: Workload-aware, Redundancy-exploiting Linear Algebra | 2023 | SIGMOD |
| 3 | 3,092 | DeepSqueeze: Deep Semantic Compression for Tabular Data | 2020 | SIGMOD |
| 4 | 7,173 | CompressDB: Enabling Efficient Compressed Data Direct Processing for Various Databases | 2022 | SIGMOD |
| 5 | 11,770 | An Evaluation of Methods of Compressing Doubles | 2020 | SIGMOD |
| 6 | 921 | Query Optimization In Compressed Database Systems | 2001 | SIGMOD |
| 7 | 9,673 | High-Ratio Compression for Machine-Generated Data | 2023 | SIGMOD |
| 8 | 9,520 | Experimental Analysis of Large-scale Learnable Vector Storage Compression | 2024 | VLDB |
| 9 | 8,617 | Improving Matrix-vector Multiplication via Lossless Grammar-Compressed Matrices | 2022 | VLDB |
| 10 | 1,644 | Compressed Linear Algebra for Large-Scale Machine Learning | 2016 | VLDB |