DBScholar

Back to papers

Approximate Query Processing: Taming the TeraBytes! A Tutorial

Summary: Tutorial on approximate aggregate query processing via online sampling and precomputed synopses with explicit error guarantees. Covers multidimensional data, joins, set-valued queries, Aqua-style architectures, streaming, dependencies, and workload-tuned synopsis design. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h07b4f63cac96fca1
Venue
VLDB
Year
2001
Pagerank
0.0001971778
Overall Rank
372 | 97.51%
DOI
-

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{garofalakis_vldb01,
        title = {{Approximate Query Processing: Taming the TeraBytes! A Tutorial}},
        author = {Garofalakis, Minos and Gibbons, Phillip B.},
        journal = {PVLDB},
        series = {{VLDB} '01},
        pages = {169},
        year = {2001}
}

Incoming Citations (Sorted by Pagerank)

Showing 41 of 41 citing papers.

Rank Citing Paper Year Venue Pagerank
126 Trio: A System for Integrated Management of Data, Accuracy, and Lineage 2005 CIDR 0.00030430114
343 Model-Driven Data Acquisition in Sensor Networks 2004 VLDB 0.00020510274
435 Mining Database Structure; Or, How to Build a Data Quality Browser 2002 SIGMOD 0.0001831946
541 BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics 2020 VLDB 0.00016663833
843 Processing Complex Aggregate Queries over Data Streams 2002 SIGMOD 0.00013534623
850 The Design of an Acquisitional Query Processor For Sensor Networks 2003 SIGMOD 0.00013482116
1,061 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012208639
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011793347
1,246 Blink and It's Done: Interactive Queries on Very Large Data 2012 VLDB 0.000113499
1,662 Rapid Sampling for Visualizations with Ordering Guarantees 2015 VLDB 9.9502569e-05
1,722 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.7921604e-05
2,003 Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee 2016 SIGMOD 9.2071735e-05
2,327 Partial Results for Online Query Processing 2002 SIGMOD 8.6271354e-05
2,416 Using Probabilistic Models for Data Management in Acquisitional Environments 2005 CIDR 8.4946354e-05
2,638 Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation 2025 SIGMOD 8.1800556e-05
3,042 Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations 2019 SIGMOD 7.7179591e-05
3,100 Lux: Always-on Visualization Recommendations for Exploratory Dataframe Workflows 2022 VLDB 7.64814e-05
3,341 I've Seen "Enough": Incrementally Improving Visualizations to Support Rapid Decision Making 2017 VLDB 7.4028264e-05
3,988 Exploiting Correlations for Expensive Predicate Evaluation 2015 SIGMOD 6.8694751e-05
4,292 A Method for Optimizing Opaque Filter Queries 2020 SIGMOD 6.6792739e-05
4,650 Mining Graph Patterns Efficiently via Randomized Summaries 2009 VLDB 6.4837167e-05
4,962 PrivateClean: Data Cleaning and Differential Privacy 2016 SIGMOD 6.3360478e-05
5,340 Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First 2026 CIDR 6.1719029e-05
5,985 A Random Walk Approach to Sampling Hidden Databases 2007 SIGMOD 5.9233486e-05
6,146 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 5.869373e-05
6,850 Capturing Data Uncertainty in High-Volume Stream Processing 2009 CIDR 5.6635511e-05
7,522 Benchmarking Spreadsheet Systems 2020 SIGMOD 5.5020393e-05
8,071 Mining a Search Engine’s Corpus: Efficient Yet Unbiased Sampling and Aggregate Estimation 2011 SIGMOD 5.3926469e-05
8,408 FaDE: More Than a Million What-ifs Per Second 2025 VLDB 5.3364407e-05
8,876 Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views 2015 VLDB 5.2587627e-05
9,177 Efficient Approximations of Conjunctive Queries 2012 PODS 5.2117841e-05
9,832 Aggregate Estimation Over Dynamic Hidden Web Databases 2014 VLDB 5.1227072e-05
9,904 Auto-Approximation of Graph Computing 2014 VLDB 5.1101476e-05
10,198 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 5.0599411e-05
10,555 Approximate Query Processing under Updates 2026 SIGMOD 4.9769913e-05
10,608 Stochastic Submodular Data Forgetting 2026 SIGMOD 4.9769913e-05
12,011 In the Land of Data Streams where Synopses are Missing, One Framework to Bring Them All 2021 VLDB 4.9769913e-05
12,192 Enabling Data Science for the Majority 2019 VLDB 4.9769913e-05
12,334 A Study of Sorting Algorithms on Approximate Memory 2016 SIGMOD 4.9769913e-05
12,513 When Data Management Systems Meet Approximate Hardware: Challenges and Opportunities 2014 VLDB 4.9769913e-05
12,995 AQAX: A System for Approximate XML Query Answers 2006 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 50 of 52 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1 Access Path Selection in a Relational Database Management System 1979 SIGMOD 0.0023943337
9 Online Aggregation 1997 SIGMOD 0.00076265429
36 Accurate Estimation Of The Number Of Tuples Satisfying A Condition 1984 SIGMOD 0.00047864281
37 Improved Histograms for Selectivity Estimation of Range Predicates 1996 SIGMOD 0.0004772731
57 On Random Sampling over Joins 1999 SIGMOD 0.00040095727
58 Statistical Estimators for Relational Algebra Expressions 1988 PODS 0.00040025054
77 Sampling-Based Estimation of the Number of Distinct Values of an Attribute 1995 VLDB 0.00036817139
79 Practical Selectivity Estimation through Adaptive Sampling 1990 SIGMOD 0.00036476265
83 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00035962466
91 On the Propagation of Errors in the Size of Join Results 1991 SIGMOD 0.00034748721
103 Selectivity Estimation Without the Attribute Value Independence Assumption 1997 VLDB 0.00033884854
119 Equi-Depth Histograms For Estimating Selectivity Factors For Multi-Dimensional Queries 1988 SIGMOD 0.0003136296
135 Ripple Joins for Online Aggregation 1999 SIGMOD 0.00029858107
138 Join Synopses for Approximate Query Answering 1999 SIGMOD 0.00029618887
151 An Efficient, Cost-Driven Index Selection Tool for Microsoft SQL Server 1997 VLDB 0.00028664776
153 New Sampling-Based Summary Statistics for Improving Approximate Query Answers 1998 SIGMOD 0.00028621958
169 Wavelet-Based Histograms for Selectivity Estimation 1998 SIGMOD 0.00027126333
181 Processing Aggregate Relational Queries with Hard Time Constraints 1989 SIGMOD 0.00026384065
222 Approximate Computation of Multidimensional Aggregates of Sparse Data Using Wavelets 1999 SIGMOD 0.00024210103
232 Adaptive Selectivity Estimation Using Query Feedback 1994 SIGMOD 0.0002378554
242 Fast Incremental Maintenance of Approximate Histograms 1997 VLDB 0.00023354266
267 Distinct Sampling for Highly-Accurate Answers to Distinct Values Queries and Event Reports 2001 VLDB 0.00022713652
275 Optimal Histograms with Quality Guarantees 1998 VLDB 0.00022404363
284 Balancing Histogram Optimality and Practicality for Query Result Size Estimation 1995 SIGMOD 0.00022205848
286 Selectivity Estimation using Probabilistic Models 2001 SIGMOD 0.00022112534
295 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00021908194
309 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021376597
312 Surfing Wavelets on Streams: One-Pass Summaries for Approximate Aggregate Queries 2001 VLDB 0.0002130211
371 STHoles: A Multidimensional Workload-Aware Histogram 2001 SIGMOD 0.00019822444
378 AutoAdmin "What-if" Index Analysis Utility 1998 SIGMOD 0.00019541534
429 Tracking Join and Self-Join Sizes in Limited Storage 1999 PODS 0.00018445263
448 Histogram-Based Approximation of Set-Valued Query Answers 1999 VLDB 0.00018121533
454 Self-tuning Histograms: Building Histograms Without Looking at Data 1999 SIGMOD 0.00017955913
518 Random Sampling for Histogram Construction: How much is enough? 1998 SIGMOD 0.00016938992
564 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016297598
566 On Computing Correlated Aggregates Over Continual Data Streams 2001 SIGMOD 0.00016288241
622 Efficiently Supporting Ad Hoc Queries in Large Datasets of Time Sequences 1997 SIGMOD 0.00015485529
666 Progressive Approximate Aggregate Queries with a Multi-Resolution Tree Structure 2001 SIGMOD 0.00014989211
700 Independence is Good: Dependency-Based Histogram Synopses for High-Dimensional Data 2001 SIGMOD 0.00014675903
746 Bifocal Sampling for Skew-Resistant Join Size Estimation 1996 SIGMOD 0.00014282427
806 Universality of Serial Histograms 1993 VLDB 0.00013786471
867 Approximating Multi-Dimensional Aggregate Range Queries Over Real Attributes 2000 SIGMOD 0.00013376165
1,004 Dynamic Maintenance of Wavelet-Based Histograms 2000 VLDB 0.00012588952
1,078 Multi-dimensional Selectivity Estimation Using Compressed Histogram Information 1999 SIGMOD 0.00012149796
1,182 ICICLES: Self-tuning Samples for Approximate Query Answering 2000 VLDB 0.00011615497
1,743 Combining Histograms and Parametric Curve Fitting for Feedback-Driven Query Result-Size Estimation 1999 VLDB 9.7342409e-05
1,758 Semantic Compression and Pattern Extraction with Fascicles 1999 VLDB 9.7095435e-05
2,654 A Robust, Optimization-Based Approach for Approximate Answering of Aggregate Queries 2001 SIGMOD 8.1670397e-05
2,657 SPARTAN: A Model-Based Semantic Compression System for Massive Data Tables 2001 SIGMOD 8.162525e-05
3,303 Optimal and Approximate Computation of Summary Statistics for Range Aggregates 2001 PODS 7.4400795e-05
Previous Page 1 / 2 Next

Semantically Similar Papers