Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent
Summary: Proposes tuple-oriented compression (TOC) for mini-batch SGD, preserving tuple boundaries while using LZW-inspired coding. Enables compressed-domain matrix operations on TOC, delivering up to 51x compression and 10.2x speedups for MGD workloads with no decompression overhead. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Fengan Li
- 2. Lingjiao Chen
- 3. Yijing Zeng
- 4. Arun Kumar
- 5. Xi Wu
- 6. Jeffrey F. Naughton
- 7. Jignesh M. Patel
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 684 | Cerebro: A Data System for Optimized Deep Learning Model Selection | 2020 | VLDB | 0.00018152321 |
| 1,942 | SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging | 2021 | SIGMOD | 0.00010010569 |
| 4,601 | Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches | 2021 | VLDB | 6.05274e-05 |
| 7,702 | ExDRa: Exploratory Data Science on Federated Raw Data | 2021 | SIGMOD | 4.6689015e-05 |
| 8,782 | AWARE: Workload-aware, Redundancy-exploiting Linear Algebra | 2023 | SIGMOD | 4.4478585e-05 |
| 8,864 | Cerebro: A Layered Data Platform for Scalable Deep Learning | 2021 | CIDR | 4.4283952e-05 |
| 10,298 | QStore: Quantization-Aware Compressed Model Storage | 2026 | VLDB | 4.1905499e-05 |
| 10,303 | Morphing-based Compression for Data-centric ML Pipelines | 2026 | VLDB | 4.1905499e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 18 of 18 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| Overall Rank | Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,059 | Progressive Compressed Records: Taking a Byte out of Deep Learning Data | 2021 | VLDB | 5.2272544e-05 |
| 8,782 | AWARE: Workload-aware, Redundancy-exploiting Linear Algebra | 2023 | SIGMOD | 4.4478585e-05 |
| 3,741 | DeepSqueeze: Deep Semantic Compression for Tabular Data | 2020 | SIGMOD | 6.7952067e-05 |
| 7,431 | CompressDB: Enabling Efficient Compressed Data Direct Processing for Various Databases | 2022 | SIGMOD | 4.7274757e-05 |
| 1,098 | Query Optimization In Compressed Database Systems | 2001 | SIGMOD | 0.00014070252 |
| 11,578 | An Evaluation of Methods of Compressing Doubles | 2020 | SIGMOD | 4.1905499e-05 |
| 9,595 | High-Ratio Compression for Machine-Generated Data | 2023 | SIGMOD | 4.3153078e-05 |
| 9,414 | Experimental Analysis of Large-scale Learnable Vector Storage Compression | 2024 | VLDB | 4.3399748e-05 |
| 8,655 | Improving Matrix-vector Multiplication via Lossless Grammar-Compressed Matrices | 2022 | VLDB | 4.4687769e-05 |
| 1,970 | Compressed Linear Algebra for Large-Scale Machine Learning | 2016 | VLDB | 9.9024431e-05 |