Back to papers
AWARE: Workload-aware, Redundancy-exploiting Linear Algebra
Summary: Introduces AWARE, a workload-aware compression framework for ML pipelines that summarizes their workload and optimizes compression plus execution plans to minimize runtime. It exploits redundancy beyond sparsity with new schemes and kernels, delivering up to 10,000x per-op and 6.6x ML gains over uncompressed baselines.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h4044a49f1641de1b
Venue
SIGMOD
Year
2023
Pagerank
5.3421754e-05
Overall Rank
8,384 | 43.64%
DOI
10.1145/3588682
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{baunsgaard_sigmod23,
title = {{AWARE: Workload-aware, Redundancy-exploiting Linear Algebra}},
author = {Baunsgaard, Sebastian and Boehm, Matthias},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3588682},
url = {https://dl.acm.org/doi/10.1145/3588682},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 38 of 38 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
61
Integrating Compression and Execution in Column-Oriented Database Systems
2006
SIGMOD
0.000392237
163
DB2 with BLU Acceleration: So Much More than Just a Column Store
2013
VLDB
0.0002749118
216
Small Materialized Aggregates: A Light Weight Index Structure for Data Warehousing
1998
VLDB
0.00024485024
219
SIMD-Scan: Ultra Fast in-Memory Table Scan using on-Chip Vector Processing Units
2009
VLDB
0.00024363532
415
SystemML: Declarative Machine Learning on Spark
2016
VLDB
0.0001865959
521
Learning Linear Regression Models over Factorized Joins
2016
SIGMOD
0.00016929744
672
The TileDB Array Data Storage Manager
2017
VLDB
0.0001489638
730
Learning Generalized Linear Models Over Normalized Data
2015
SIGMOD
0.00014406936
906
Data Blocks: Hybrid OLTP and OLAP on Compressed Storage using both Vectorization and Compilation
2016
SIGMOD
0.00013160654
910
Dictionary-based Order-preserving String Compression for Main Memory Column Stores
2009
SIGMOD
0.00013123912
1,132
Qd-tree: Learning Data Layouts for Big Data Analytics
2020
SIGMOD
0.00011898257
1,256
Towards Linear Algebra over Normalized Data
2017
VLDB
0.00011314687
1,327
Column-oriented Database Systems
2009
VLDB
0.00011003776
1,347
Chimp: Efficient Lossless Floating Point Compression for Time Series Databases
2022
VLDB
0.00010950342
1,614
Compressed Linear Algebra for Large-Scale Machine Learning
2016
VLDB
0.00010071891
1,770
How to Wring a Table Dry: Entropy Compression of Relations and Querying of Compressed Relations
2006
VLDB
9.686541e-05
1,973
Decomposed Bounded Floats for Fast Compression and Queries
2021
VLDB
9.2834606e-05
1,984
How to Barter Bits for Chronons: Compression and Bandwidth Trade Offs for Database Scans
2007
SIGMOD
9.2557255e-05
2,187
Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities
2021
SIGMOD
8.8896655e-05
2,663
A Layered Aggregate Engine for Analytics Workloads
2019
SIGMOD
8.1581558e-05
2,882
Optimal Column Layout for Hybrid Workloads
2019
VLDB
7.9116043e-05
3,101
On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML
2018
VLDB
7.649219e-05
3,140
DeepSqueeze: Deep Semantic Compression for Tabular Data
2020
SIGMOD
7.6023603e-05
3,613
Plato: Approximate Analytics over Compressed Time Series with Tight Deterministic Error Guarantees
2020
VLDB
7.162283e-05
3,872
SketchML: Accelerating Distributed Machine Learning with Data Sketches
2018
SIGMOD
6.9540368e-05
4,161
The Relational Data Borg is Learning
2020
VLDB
6.7700593e-05
4,330
MNC: Structure-Exploiting Sparsity Estimation for Matrix Expressions
2019
SIGMOD
6.6595681e-05
4,879
Accelerating Generalized Linear Models with MLWeaving: A One-Size-Fits-All System for Any-Precision Learning
2019
VLDB
6.3715334e-05
5,853
Compression Aware Physical Database Design
2011
VLDB
5.9710387e-05
5,877
BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees
2019
SIGMOD
5.9627218e-05
6,169
Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear Algebra
2021
SIGMOD
5.8624857e-05
6,311
Progressive Compressed Records: Taking a Byte out of Deep Learning Data
2021
VLDB
5.8175614e-05
6,417
MorphStore: Analytical Query Engine with a Holistic Compression-Enabled Processing Model
2020
VLDB
5.791405e-05
6,592
Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent
2019
SIGMOD
5.7421684e-05
7,839
ExDRa: Exploratory Data Science on Federated Raw Data
2021
SIGMOD
5.4432099e-05
8,563
Robust and Budget-Constrained Encoding Configurations for In-Memory Database Systems
2022
VLDB
5.3138647e-05
8,773
Improving Matrix-vector Multiplication via Lossless Grammar-Compressed Matrices
2022
VLDB
5.2804275e-05
9,457
COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy Compression
2022
VLDB
5.1735029e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
8,791
Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to Ask
2024
VLDB
2
4,116
Resource Elasticity for Large-Scale Machine Learning
2015
SIGMOD
3
9,700
Experimental Analysis of Large-scale Learnable Vector Storage Compression
2024
VLDB
4
6,592
Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent
2019
SIGMOD
5
6,169
Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear Algebra
2021
SIGMOD
6
4,974
SPORES: Sum-Product Optimization via Relational Equality Saturation for Large Scale Linear Algebra
2020
VLDB
7
3,815
Comprehensive and Efficient Workload Compression
2021
VLDB
8
8,773
Improving Matrix-vector Multiplication via Lossless Grammar-Compressed Matrices
2022
VLDB
9
10,945
Morphing-based Compression for Data-centric ML Pipelines
2026
VLDB
10
1,614
Compressed Linear Algebra for Large-Scale Machine Learning
2016
VLDB