DBScholar

Back to papers

AWARE: Workload-aware, Redundancy-exploiting Linear Algebra

Summary: Introduces AWARE, a workload-aware compression framework for ML pipelines that summarizes their workload and optimizes compression plus execution plans to minimize runtime. It exploits redundancy beyond sparsity with new schemes and kernels, delivering up to 10,000x per-op and 6.6x ML gains over uncompressed baselines. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h4044a49f1641de1b
Venue
SIGMOD
Year
2023
Pagerank
5.3396465e-05
Overall Rank
8,389 | 43.62%
DOI
10.1145/3588682

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{baunsgaard_sigmod23,
        title = {{AWARE: Workload-aware, Redundancy-exploiting Linear Algebra}},
        author = {Baunsgaard, Sebastian and Boehm, Matthias},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3588682},
        url = {https://dl.acm.org/doi/10.1145/3588682},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 38 of 38 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
61 Integrating Compression and Execution in Column-Oriented Database Systems 2006 SIGMOD 0.00039236924
163 DB2 with BLU Acceleration: So Much More than Just a Column Store 2013 VLDB 0.00027480091
216 Small Materialized Aggregates: A Light Weight Index Structure for Data Warehousing 1998 VLDB 0.00024485637
219 SIMD-Scan: Ultra Fast in-Memory Table Scan using on-Chip Vector Processing Units 2009 VLDB 0.00024367137
416 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.00018650998
521 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.00016923519
673 The TileDB Array Data Storage Manager 2017 VLDB 0.00014891415
731 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014400356
907 Data Blocks: Hybrid OLTP and OLAP on Compressed Storage using both Vectorization and Compilation 2016 SIGMOD 0.00013157412
910 Dictionary-based Order-preserving String Compression for Main Memory Column Stores 2009 SIGMOD 0.00013122392
1,128 Qd-tree: Learning Data Layouts for Big Data Analytics 2020 SIGMOD 0.00011901941
1,257 Towards Linear Algebra over Normalized Data 2017 VLDB 0.0001130959
1,328 Column-oriented Database Systems 2009 VLDB 0.00010998855
1,347 Chimp: Efficient Lossless Floating Point Compression for Time Series Databases 2022 VLDB 0.00010945231
1,614 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010067153
1,770 How to Wring a Table Dry: Entropy Compression of Relations and Querying of Compressed Relations 2006 VLDB 9.6822825e-05
1,974 Decomposed Bounded Floats for Fast Compression and Queries 2021 VLDB 9.2806652e-05
1,984 How to Barter Bits for Chronons: Compression and Bandwidth Trade Offs for Database Scans 2007 SIGMOD 9.2529061e-05
2,189 Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities 2021 SIGMOD 8.8854572e-05
2,663 A Layered Aggregate Engine for Analytics Workloads 2019 SIGMOD 8.1542952e-05
2,880 Optimal Column Layout for Hybrid Workloads 2019 VLDB 7.9118308e-05
3,103 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6456038e-05
3,141 DeepSqueeze: Deep Semantic Compression for Tabular Data 2020 SIGMOD 7.5987721e-05
3,607 Plato: Approximate Analytics over Compressed Time Series with Tight Deterministic Error Guarantees 2020 VLDB 7.1645803e-05
3,873 SketchML: Accelerating Distributed Machine Learning with Data Sketches 2018 SIGMOD 6.9507449e-05
4,161 The Relational Data Borg is Learning 2020 VLDB 6.7669004e-05
4,330 MNC: Structure-Exploiting Sparsity Estimation for Matrix Expressions 2019 SIGMOD 6.6564176e-05
4,881 Accelerating Generalized Linear Models with MLWeaving: A One-Size-Fits-All System for Any-Precision Learning 2019 VLDB 6.368585e-05
5,856 Compression Aware Physical Database Design 2011 VLDB 5.9682143e-05
5,877 BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees 2019 SIGMOD 5.9599042e-05
6,171 Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear Algebra 2021 SIGMOD 5.8597104e-05
6,314 Progressive Compressed Records: Taking a Byte out of Deep Learning Data 2021 VLDB 5.8148251e-05
6,419 MorphStore: Analytical Query Engine with a Holistic Compression-Enabled Processing Model 2020 VLDB 5.7886634e-05
6,594 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 5.7394502e-05
7,843 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4406331e-05
8,570 Robust and Budget-Constrained Encoding Configurations for In-Memory Database Systems 2022 VLDB 5.3113492e-05
8,781 Improving Matrix-vector Multiplication via Lossless Grammar-Compressed Matrices 2022 VLDB 5.2779278e-05
9,466 COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy Compression 2022 VLDB 5.1710538e-05
Previous Page 1 / 1 Next

Semantically Similar Papers