Database Paper Browser

Back to papers

Approximate Query Processing: Taming the TeraBytes! A Tutorial

Summary: Survey of approximate query processing for terabytes, contrasting online aggregation with precomputed synopses for fast, bounded results. Covers multi-dimensional and join synopses, set-valued queries, AQUA-style rewrite, maintenance, streaming data. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
8825
Venue
VLDB
Year
2001
Pagerank
0.00020183267
Overall Rank
364 | 97.48%
DOI
-

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 41 of 41 citing papers.

Rank Citing Paper Year Venue Pagerank
125 Trio: A System for Integrated Management of Data, Accuracy, and Lineage 2005 CIDR 0.0003120745
332 Model-Driven Data Acquisition in Sensor Networks 2004 VLDB 0.00020989244
426 Mining Database Structure; Or, How to Build a Data Quality Browser 2002 SIGMOD 0.00018745626
564 BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics 2020 VLDB 0.0001653392
799 The Design of an Acquisitional Query Processor For Sensor Networks 2003 SIGMOD 0.00013973755
800 Processing Complex Aggregate Queries over Data Streams 2002 SIGMOD 0.00013970432
1,098 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012272786
1,192 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011836578
1,212 Blink and It's Done: Interactive Queries on Very Large Data 2012 VLDB 0.00011741576
1,608 Rapid Sampling for Visualizations with Ordering Guarantees 2015 VLDB 0.00010307016
1,758 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.9011485e-05
1,939 Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee 2016 SIGMOD 9.5245725e-05
2,256 Partial Results for Online Query Processing 2002 SIGMOD 8.9310801e-05
2,325 Using Probabilistic Models for Data Management in Acquisitional Environments 2005 CIDR 8.8281302e-05
2,887 Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations 2019 SIGMOD 8.0623053e-05
3,051 Lux: Always-on Visualization Recommendations for Exploratory Dataframe Workflows 2022 VLDB 7.8758656e-05
3,253 I've Seen "Enough": Incrementally Improving Visualizations to Support Rapid Decision Making 2017 VLDB 7.6654919e-05
3,914 Exploiting Correlations for Expensive Predicate Evaluation 2015 SIGMOD 7.0851333e-05
4,240 A Method for Optimizing Opaque Filter Queries 2020 SIGMOD 6.8750034e-05
4,514 Mining Graph Patterns Efficiently via Randomized Summaries 2009 VLDB 6.7219831e-05
4,681 Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation 2025 SIGMOD 6.6286556e-05
4,793 PrivateClean: Data Cleaning and Differential Privacy 2016 SIGMOD 6.5720031e-05
5,776 A Random Walk Approach to Sampling Hidden Databases 2007 SIGMOD 6.1554293e-05
5,981 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 6.0813339e-05
6,601 Capturing Data Uncertainty in High-Volume Stream Processing 2009 CIDR 5.8860028e-05
7,288 Benchmarking Spreadsheet Systems 2020 SIGMOD 5.7096966e-05
7,781 Mining a Search Engine’s Corpus: Efficient Yet Unbiased Sampling and Aggregate Estimation 2011 SIGMOD 5.6039053e-05
8,605 Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views 2015 VLDB 5.4602302e-05
8,883 Efficient Approximations of Conjunctive Queries 2012 PODS 5.4156287e-05
9,440 Aggregate Estimation Over Dynamic Hidden Web Databases 2014 VLDB 5.3325274e-05
9,602 Auto-Approximation of Graph Computing 2014 VLDB 5.3086438e-05
9,991 Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First 2026 CIDR 5.1725247e-05
10,049 Approximate Query Processing under Updates 2026 SIGMOD 5.1725247e-05
10,116 Stochastic Submodular Data Forgetting 2026 SIGMOD 5.1725247e-05
10,215 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 5.1725247e-05
10,890 FaDE: More Than a Million What-ifs Per Second 2025 VLDB 5.1725247e-05
11,506 In the Land of Data Streams where Synopses are Missing, One Framework to Bring Them All 2021 VLDB 5.1725247e-05
11,696 Enabling Data Science for the Majority 2019 VLDB 5.1725247e-05
11,840 A Study of Sorting Algorithms on Approximate Memory 2016 SIGMOD 5.1725247e-05
12,027 When Data Management Systems Meet Approximate Hardware: Challenges and Opportunities 2014 VLDB 5.1725247e-05
12,515 AQAX: A System for Approximate XML Query Answers 2006 VLDB 5.1725247e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 50 of 52 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1 Access Path Selection in a Relational Database Management System 1979 SIGMOD 0.0024179717
10 Online Aggregation 1997 SIGMOD 0.00077936311
35 Improved Histograms for Selectivity Estimation of Range Predicates 1996 SIGMOD 0.00048862242
37 Accurate Estimation Of The Number Of Tuples Satisfying A Condition 1984 SIGMOD 0.00048770736
54 Statistical Estimators for Relational Algebra Expressions 1988 PODS 0.00041048998
56 On Random Sampling over Joins 1999 SIGMOD 0.00040582148
75 Practical Selectivity Estimation through Adaptive Sampling 1990 SIGMOD 0.00037396456
76 Sampling-Based Estimation of the Number of Distinct Values of an Attribute 1995 VLDB 0.00037322071
83 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00036461682
98 On the Propagation of Errors in the Size of Join Results 1991 SIGMOD 0.00034732081
100 Selectivity Estimation Without the Attribute Value Independence Assumption 1997 VLDB 0.00034549997
120 Equi-Depth Histograms For Estimating Selectivity Factors For Multi-Dimensional Queries 1988 SIGMOD 0.00032260989
136 Ripple Joins for Online Aggregation 1999 SIGMOD 0.00030254656
140 Join Synopses for Approximate Query Answering 1999 SIGMOD 0.00029955359
151 New Sampling-Based Summary Statistics for Improving Approximate Query Answers 1998 SIGMOD 0.00029351676
154 An Efficient, Cost-Driven Index Selection Tool for Microsoft SQL Server 1997 VLDB 0.00028960058
167 Wavelet-Based Histograms for Selectivity Estimation 1998 SIGMOD 0.000277743
180 Processing Aggregate Relational Queries with Hard Time Constraints 1989 SIGMOD 0.00027097952
212 Approximate Computation of Multidimensional Aggregates of Sparse Data Using Wavelets 1999 SIGMOD 0.00025045482
221 Adaptive Selectivity Estimation Using Query Feedback 1994 SIGMOD 0.00024405057
234 Fast Incremental Maintenance of Approximate Histograms 1997 VLDB 0.00024050031
255 Distinct Sampling for Highly-Accurate Answers to Distinct Values Queries and Event Reports 2001 VLDB 0.00023376534
266 Optimal Histograms with Quality Guarantees 1998 VLDB 0.00023039606
274 Selectivity Estimation using Probabilistic Models 2001 SIGMOD 0.00022678252
277 Balancing Histogram Optimality and Practicality for Query Result Size Estimation 1995 SIGMOD 0.00022592757
281 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00022539374
300 Approximate Query Processing Using Wavelets 2000 VLDB 0.00022034489
302 Surfing Wavelets on Streams: One-Pass Summaries for Approximate Aggregate Queries 2001 VLDB 0.00022030492
360 STHoles: A Multidimensional Workload-Aware Histogram 2001 SIGMOD 0.00020269545
377 AutoAdmin "What-if" Index Analysis Utility 1998 SIGMOD 0.00019763016
425 Histogram-Based Approximation of Set-Valued Query Answers 1999 VLDB 0.00018757897
427 Tracking Join and Self-Join Sizes in Limited Storage 1999 PODS 0.00018740148
438 Self-tuning Histograms: Building Histograms Without Looking at Data 1999 SIGMOD 0.00018515832
503 Random Sampling for Histogram Construction: How much is enough? 1998 SIGMOD 0.00017452048
541 On Computing Correlated Aggregates Over Continual Data Streams 2001 SIGMOD 0.00016888064
554 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016672835
607 Efficiently Supporting Ad Hoc Queries in Large Datasets of Time Sequences 1997 SIGMOD 0.00015929565
644 Progressive Approximate Aggregate Queries with a Multi-Resolution Tree Structure 2001 SIGMOD 0.00015520834
685 Independence is Good: Dependency-Based Histogram Synopses for High-Dimensional Data 2001 SIGMOD 0.00015100765
731 Bifocal Sampling for Skew-Resistant Join Size Estimation 1996 SIGMOD 0.00014686276
769 Universality of Serial Histograms 1993 VLDB 0.00014233685
837 Approximating Multi-Dimensional Aggregate Range Queries Over Real Attributes 2000 SIGMOD 0.00013759323
951 Dynamic Maintenance of Wavelet-Based Histograms 2000 VLDB 0.00013057833
1,046 Multi-dimensional Selectivity Estimation Using Compressed Histogram Information 1999 SIGMOD 0.00012552939
1,156 ICICLES: Self-tuning Samples for Approximate Query Answering 2000 VLDB 0.00011999802
1,619 Combining Histograms and Parametric Curve Fitting for Feedback-Driven Query Result-Size Estimation 1999 VLDB 0.00010273734
1,704 Semantic Compression and Pattern Extraction with Fascicles 1999 VLDB 0.00010060319
2,577 SPARTAN: A Model-Based Semantic Compression System for Massive Data Tables 2001 SIGMOD 8.4646103e-05
2,583 A Robust, Optimization-Based Approach for Approximate Answering of Aggregate Queries 2001 SIGMOD 8.4563178e-05
3,197 Optimal and Approximate Computation of Summary Statistics for Range Aggregates 2001 PODS 7.718299e-05
Previous Page 1 / 2 Next

Semantically Similar Papers