Progressive Compressed Records: Taking a Byte out of Deep Learning Data
Summary: Progressive Compressed Records (PCRs) combine progressive compression with a storage layout exposing one dataset at multiple fidelities without increasing total size. Automatic runtime level selection cuts training bandwidth by up to 50%, potentially doubling time-to-accuracy. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Michael Kuchnik (Carnegie Mellon University)
- 2. George Amvrosiadis (Carnegie Mellon University)
- 3. Virginia Smith (Carnegie Mellon University)
BibTeX Citation
@article{kuchnik_vldb21,
title = {{Progressive Compressed Records: Taking a Byte out of Deep Learning Data}},
author = {Kuchnik, Michael and Amvrosiadis, George and Smith, Virginia},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {11},
pages = {2627--2641},
doi = {10.14778/3476249.3476308},
url = {https://doi.org/10.14778/3476249.3476308},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,764 | FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline | 2023 | VLDB | 7.1446931e-05 |
| 8,794 | AWARE: Workload-aware, Redundancy-exploiting Linear Algebra | 2023 | SIGMOD | 5.370464e-05 |
| 11,425 | Homomorphic Compression: Making Text Processing on Compression Unlimited | 2023 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 60 | Integrating Compression and Execution in Column-Oriented Database Systems | 2006 | SIGMOD | 0.0003955489 |
| 1,446 | Analyzing and Mitigating Data Stalls in DNN Training | 2021 | VLDB | 0.0001076818 |
| 1,666 | HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics | 2016 | VLDB | 0.00010068964 |
| 2,018 | tf.data: A Machine Learning Data Processing Framework | 2021 | VLDB | 9.3001937e-05 |
| 4,798 | Accelerating Generalized Linear Models with MLWeaving: A One-Size-Fits-All System for Any-Precision Learning | 2019 | VLDB | 6.5024772e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,173 | CompressDB: Enabling Efficient Compressed Data Direct Processing for Various Databases | 2022 | SIGMOD |
| 2 | 11,770 | An Evaluation of Methods of Compressing Doubles | 2020 | SIGMOD |
| 3 | 10,580 | GPU Acceleration of SQL Analytics on Compressed Data | 2026 | VLDB |
| 4 | 6,485 | Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent | 2019 | SIGMOD |
| 5 | 8,634 | Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to Ask | 2024 | VLDB |
| 6 | 3,092 | DeepSqueeze: Deep Semantic Compression for Tabular Data | 2020 | SIGMOD |
| 7 | 9,673 | High-Ratio Compression for Machine-Generated Data | 2023 | SIGMOD |
| 8 | 9,292 | COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy Compression | 2022 | VLDB |
| 9 | 1,644 | Compressed Linear Algebra for Large-Scale Machine Learning | 2016 | VLDB |
| 10 | 9,520 | Experimental Analysis of Large-scale Learnable Vector Storage Compression | 2024 | VLDB |