Progressive Compressed Records: Taking a Byte out of Deep Learning Data
Summary: Progressive Compressed Records (PCRs) combine progressive compression with a storage layout exposing one dataset at multiple fidelities without increasing total size. Automatic runtime level selection cuts training bandwidth by up to 50%, potentially doubling time-to-accuracy. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Michael Kuchnik (Carnegie Mellon University)
- 2. George Amvrosiadis (Carnegie Mellon University)
- 3. Virginia Smith (Carnegie Mellon University)
BibTeX Citation
@article{kuchnik_vldb21,
title = {{Progressive Compressed Records: Taking a Byte out of Deep Learning Data}},
author = {Kuchnik, Michael and Amvrosiadis, George and Smith, Virginia},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {11},
pages = {2627--2641},
doi = {10.14778/3476249.3476308},
url = {https://doi.org/10.14778/3476249.3476308},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,811 | FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline | 2023 | VLDB | 7.0087478e-05 |
| 8,384 | AWARE: Workload-aware, Redundancy-exploiting Linear Algebra | 2023 | SIGMOD | 5.3421754e-05 |
| 11,739 | Homomorphic Compression: Making Text Processing on Compression Unlimited | 2023 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 61 | Integrating Compression and Execution in Column-Oriented Database Systems | 2006 | SIGMOD | 0.000392237 |
| 1,475 | Analyzing and Mitigating Data Stalls in DNN Training | 2021 | VLDB | 0.00010556672 |
| 1,541 | HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics | 2016 | VLDB | 0.00010313459 |
| 2,053 | tf.data: A Machine Learning Data Processing Framework | 2021 | VLDB | 9.1140407e-05 |
| 4,879 | Accelerating Generalized Linear Models with MLWeaving: A One-Size-Fits-All System for Any-Precision Learning | 2019 | VLDB | 6.3715334e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,737 | CompressDB: Enabling Efficient Compressed Data Direct Processing for Various Databases | 2022 | SIGMOD |
| 2 | 12,072 | An Evaluation of Methods of Compressing Doubles | 2020 | SIGMOD |
| 3 | 9,863 | GPU Acceleration of SQL Analytics on Compressed Data | 2026 | VLDB |
| 4 | 6,592 | Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent | 2019 | SIGMOD |
| 5 | 8,791 | Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to Ask | 2024 | VLDB |
| 6 | 3,140 | DeepSqueeze: Deep Semantic Compression for Tabular Data | 2020 | SIGMOD |
| 7 | 9,843 | High-Ratio Compression for Machine-Generated Data | 2023 | SIGMOD |
| 8 | 9,457 | COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy Compression | 2022 | VLDB |
| 9 | 1,614 | Compressed Linear Algebra for Large-Scale Machine Learning | 2016 | VLDB |
| 10 | 9,700 | Experimental Analysis of Large-scale Learnable Vector Storage Compression | 2024 | VLDB |