DBScholar

Back to papers

Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems

Summary: Sampling-based AQP for large-scale analytics; error bars often fail on real workloads. Fast diagnostics of error-estimation failures and a pipeline that yields approximate answers with reliable bootstrap-based error bars at interactive speeds. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h25a21afc71b72c19
Venue
SIGMOD
Year
2014
Pagerank
0.00010693831
Overall Rank
1,428 | 90.41%
DOI
10.1145/2588555.2593667

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{agarwal_sigmod14,
        title = {{Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems}},
        author = {Agarwal, Sameer and Milner, Henry and Kleiner, Ariel and Talwalkar, Ameet and Jordan, Michael and Madden, Samuel and Mozafari, Barzan and Stoica, Ion},
        series = {{SIGMOD} '14},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2588555.2593667},
        url = {https://dl.acm.org/doi/10.1145/2588555.2593667},
        year = {2014}
}

Incoming Citations (Sorted by Pagerank)

Showing 39 of 39 citing papers.

Rank Citing Paper Year Venue Pagerank
541 BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics 2020 VLDB 0.00016657685
596 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00015785583
784 VerdictDB: Universalizing Approximate Query Processing 2018 SIGMOD 0.00014012614
840 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.0001354605
1,082 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012122749
1,215 Overview of Data Exploration Techniques 2015 SIGMOD 0.00011492648
1,868 G-OLA: Generalized On-Line Aggregation for Interactive Analysis on Big Data 2015 SIGMOD 9.4754064e-05
1,916 The Analytical Bootstrap: a New Method for Fast Error Estimation in Approximate Query Processing 2014 SIGMOD 9.3837729e-05
2,000 Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee 2016 SIGMOD 9.2112617e-05
2,027 Database Learning: Toward a Database that Becomes Smarter Every Time 2017 SIGMOD 9.1618139e-05
2,224 DAQ: A New Paradigm for Approximate Query Processing 2015 VLDB 8.80823e-05
2,776 Approximate Selection with Guarantees using Proxies 2020 VLDB 8.0309448e-05
2,873 SnappyData: A Unified Cluster for Streaming, Transactions, and Interactive Analytics 2017 CIDR 7.9225385e-05
3,424 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.3117029e-05
3,597 Two Birds, One Stone: A Fast, yet Lightweight, Indexing Scheme for Modern Database Systems 2017 VLDB 7.1788912e-05
4,457 Lightweight and Accurate Cardinality Estimation by Neural Network Gaussian Process 2022 SIGMOD 6.5913732e-05
4,856 Neighbor-Sensitive Hashing 2016 VLDB 6.3798143e-05
5,316 CliffGuard: A Principled Framework for Finding Robust Database Designs 2015 SIGMOD 6.1836681e-05
5,373 At-the-time and Back-in-time Persistent Sketches 2021 SIGMOD 6.1569337e-05
5,822 SnappyData: A Hybrid Transactional Analytical Store Built On Spark 2016 SIGMOD 5.9826812e-05
5,831 Joins on Samples: A Theoretical Guide for Practitioners 2020 VLDB 5.9782109e-05
5,877 BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees 2019 SIGMOD 5.9627218e-05
6,004 Approximate Query Engines: Commercial Challenges and Research Opportunities 2017 SIGMOD 5.9166815e-05
6,227 iOLAP: Managing Uncertainty for Efficient Incremental OLAP 2016 SIGMOD 5.8452386e-05
6,954 Querying Big Data by Accessing Small Data 2015 PODS 5.6350266e-05
7,364 Weighted Distinct Sampling: Cardinality Estimation for SPJ Queries 2021 SIGMOD 5.5418075e-05
7,375 PairwiseHist: Fast, Accurate and Space-Efficient Approximate Query Processing with Data Compression 2024 VLDB 5.5400509e-05
7,890 TSCache: An Efficient Flash-based Caching Scheme for Time-series Data Workloads 2021 VLDB 5.4328058e-05
7,978 Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines 2024 VLDB 5.4136835e-05
8,166 Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints 2020 SIGMOD 5.3849926e-05
8,216 Wander Join: Online Aggregation for Joins 2016 SIGMOD 5.3764262e-05
8,490 THEMIS: Fairness in Federated Stream Processing under Overload 2016 SIGMOD 5.3315439e-05
8,770 One Size Does Not Fit All: A Bandit-Based Sampler Combination Framework with Theoretical Guarantees 2022 SIGMOD 5.2812395e-05
8,870 Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views 2015 VLDB 5.2601766e-05
9,539 Controlled Intentional Degradation in Analytical Video Systems 2022 SIGMOD 5.1621607e-05
9,762 Hephaestus: Data Reuse for Accelerating Scientific Discovery 2015 CIDR 5.1347497e-05
10,834 ConANN: Conformal Approximate Nearest Neighbor Search 2026 VLDB 4.9793485e-05
12,215 Demonstration of VerdictDB, the Platform-Independent AQP System 2018 SIGMOD 4.9793485e-05
12,328 A Study of Sorting Algorithms on Approximate Memory 2016 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 15 of 15 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers