DBScholar

Back to papers

Approximate Query Processing: Taming the TeraBytes! A Tutorial

Summary: Tutorial on approximate aggregate query processing via online sampling and precomputed synopses with explicit error guarantees. Covers multidimensional data, joins, set-valued queries, Aqua-style architectures, streaming, dependencies, and workload-tuned synopsis design. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h07b4f63cac96fca1
Venue
VLDB
Year
2001
Pagerank
0.00019720059
Overall Rank
372 | 97.51%
DOI
-

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{garofalakis_vldb01,
        title = {{Approximate Query Processing: Taming the TeraBytes! A Tutorial}},
        author = {Garofalakis, Minos and Gibbons, Phillip B.},
        journal = {PVLDB},
        series = {{VLDB} '01},
        pages = {169},
        year = {2001}
}

Incoming Citations (Sorted by Pagerank)

Showing 41 of 41 citing papers.

Rank Citing Paper Year Venue Pagerank
126 Trio: A System for Integrated Management of Data, Accuracy, and Lineage 2005 CIDR 0.00030439683
343 Model-Driven Data Acquisition in Sensor Networks 2004 VLDB 0.00020519525
435 Mining Database Structure; Or, How to Build a Data Quality Browser 2002 SIGMOD 0.0001832766
541 BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics 2020 VLDB 0.00016657685
842 Processing Complex Aggregate Queries over Data Streams 2002 SIGMOD 0.00013540697
850 The Design of an Acquisitional Query Processor For Sensor Networks 2003 SIGMOD 0.00013488409
1,082 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012122749
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011798912
1,243 Blink and It's Done: Interactive Queries on Very Large Data 2012 VLDB 0.0001135375
1,661 Rapid Sampling for Visualizations with Ordering Guarantees 2015 VLDB 9.9535453e-05
1,720 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.7965659e-05
2,000 Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee 2016 SIGMOD 9.2112617e-05
2,324 Partial Results for Online Query Processing 2002 SIGMOD 8.6310477e-05
2,415 Using Probabilistic Models for Data Management in Acquisitional Environments 2005 CIDR 8.4985276e-05
2,638 Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation 2025 SIGMOD 8.1839298e-05
3,040 Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations 2019 SIGMOD 7.7216143e-05
3,098 Lux: Always-on Visualization Recommendations for Exploratory Dataframe Workflows 2022 VLDB 7.6517623e-05
3,341 I've Seen "Enough": Incrementally Improving Visualizations to Support Rapid Decision Making 2017 VLDB 7.4063139e-05
3,988 Exploiting Correlations for Expensive Predicate Evaluation 2015 SIGMOD 6.871854e-05
4,293 A Method for Optimizing Opaque Filter Queries 2020 SIGMOD 6.6819917e-05
4,647 Mining Graph Patterns Efficiently via Randomized Summaries 2009 VLDB 6.4867875e-05
4,960 PrivateClean: Data Cleaning and Differential Privacy 2016 SIGMOD 6.3390487e-05
5,391 Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First 2026 CIDR 6.1504174e-05
5,985 A Random Walk Approach to Sampling Hidden Databases 2007 SIGMOD 5.926154e-05
6,143 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 5.8721471e-05
6,846 Capturing Data Uncertainty in High-Volume Stream Processing 2009 CIDR 5.6662335e-05
7,517 Benchmarking Spreadsheet Systems 2020 SIGMOD 5.5046452e-05
8,065 Mining a Search Engine’s Corpus: Efficient Yet Unbiased Sampling and Aggregate Estimation 2011 SIGMOD 5.3952009e-05
8,870 Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views 2015 VLDB 5.2601766e-05
8,892 FaDE: More Than a Million What-ifs Per Second 2025 VLDB 5.2559789e-05
9,169 Efficient Approximations of Conjunctive Queries 2012 PODS 5.2141412e-05
9,825 Aggregate Estimation Over Dynamic Hidden Web Databases 2014 VLDB 5.1251333e-05
9,897 Auto-Approximation of Graph Computing 2014 VLDB 5.1125679e-05
10,544 Approximate Query Processing under Updates 2026 SIGMOD 4.9793485e-05
10,597 Stochastic Submodular Data Forgetting 2026 SIGMOD 4.9793485e-05
10,690 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 4.9793485e-05
12,005 In the Land of Data Streams where Synopses are Missing, One Framework to Bring Them All 2021 VLDB 4.9793485e-05
12,186 Enabling Data Science for the Majority 2019 VLDB 4.9793485e-05
12,328 A Study of Sorting Algorithms on Approximate Memory 2016 SIGMOD 4.9793485e-05
12,507 When Data Management Systems Meet Approximate Hardware: Challenges and Opportunities 2014 VLDB 4.9793485e-05
12,989 AQAX: A System for Approximate XML Query Answers 2006 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 50 of 52 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1 Access Path Selection in a Relational Database Management System 1979 SIGMOD 0.0023947656
9 Online Aggregation 1997 SIGMOD 0.00076195956
36 Accurate Estimation Of The Number Of Tuples Satisfying A Condition 1984 SIGMOD 0.00047863192
37 Improved Histograms for Selectivity Estimation of Range Predicates 1996 SIGMOD 0.00047731453
57 On Random Sampling over Joins 1999 SIGMOD 0.00040108301
58 Statistical Estimators for Relational Algebra Expressions 1988 PODS 0.00040035279
77 Sampling-Based Estimation of the Number of Distinct Values of an Attribute 1995 VLDB 0.00036828234
79 Practical Selectivity Estimation through Adaptive Sampling 1990 SIGMOD 0.00036487763
83 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00035978046
91 On the Propagation of Errors in the Size of Join Results 1991 SIGMOD 0.0003475226
103 Selectivity Estimation Without the Attribute Value Independence Assumption 1997 VLDB 0.00033894985
119 Equi-Depth Histograms For Estimating Selectivity Factors For Multi-Dimensional Queries 1988 SIGMOD 0.0003137356
135 Ripple Joins for Online Aggregation 1999 SIGMOD 0.00029866033
138 Join Synopses for Approximate Query Answering 1999 SIGMOD 0.00029627449
151 An Efficient, Cost-Driven Index Selection Tool for Microsoft SQL Server 1997 VLDB 0.00028672526
153 New Sampling-Based Summary Statistics for Improving Approximate Query Answers 1998 SIGMOD 0.00028633995
169 Wavelet-Based Histograms for Selectivity Estimation 1998 SIGMOD 0.00027134723
181 Processing Aggregate Relational Queries with Hard Time Constraints 1989 SIGMOD 0.00026389403
222 Approximate Computation of Multidimensional Aggregates of Sparse Data Using Wavelets 1999 SIGMOD 0.00024218831
232 Adaptive Selectivity Estimation Using Query Feedback 1994 SIGMOD 0.00023792809
242 Fast Incremental Maintenance of Approximate Histograms 1997 VLDB 0.00023363722
267 Distinct Sampling for Highly-Accurate Answers to Distinct Values Queries and Event Reports 2001 VLDB 0.00022722971
275 Optimal Histograms with Quality Guarantees 1998 VLDB 0.00022413521
283 Balancing Histogram Optimality and Practicality for Query Result Size Estimation 1995 SIGMOD 0.00022214789
286 Selectivity Estimation using Probabilistic Models 2001 SIGMOD 0.0002211981
295 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00021914399
309 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021384073
312 Surfing Wavelets on Streams: One-Pass Summaries for Approximate Aggregate Queries 2001 VLDB 0.00021311793
371 STHoles: A Multidimensional Workload-Aware Histogram 2001 SIGMOD 0.00019829769
378 AutoAdmin "What-if" Index Analysis Utility 1998 SIGMOD 0.00019549382
428 Tracking Join and Self-Join Sizes in Limited Storage 1999 PODS 0.0001845349
448 Histogram-Based Approximation of Set-Valued Query Answers 1999 VLDB 0.00018129161
454 Self-tuning Histograms: Building Histograms Without Looking at Data 1999 SIGMOD 0.00017962189
519 Random Sampling for Histogram Construction: How much is enough? 1998 SIGMOD 0.00016942879
564 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016296665
566 On Computing Correlated Aggregates Over Continual Data Streams 2001 SIGMOD 0.00016295476
621 Efficiently Supporting Ad Hoc Queries in Large Datasets of Time Sequences 1997 SIGMOD 0.00015492309
664 Progressive Approximate Aggregate Queries with a Multi-Resolution Tree Structure 2001 SIGMOD 0.00014995058
701 Independence is Good: Dependency-Based Histogram Synopses for High-Dimensional Data 2001 SIGMOD 0.00014680907
745 Bifocal Sampling for Skew-Resistant Join Size Estimation 1996 SIGMOD 0.00014288286
806 Universality of Serial Histograms 1993 VLDB 0.00013792174
866 Approximating Multi-Dimensional Aggregate Range Queries Over Real Attributes 2000 SIGMOD 0.00013381261
1,004 Dynamic Maintenance of Wavelet-Based Histograms 2000 VLDB 0.00012594522
1,077 Multi-dimensional Selectivity Estimation Using Compressed Histogram Information 1999 SIGMOD 0.00012154948
1,183 ICICLES: Self-tuning Samples for Approximate Query Answering 2000 VLDB 0.00011616705
1,741 Combining Histograms and Parametric Curve Fitting for Feedback-Driven Query Result-Size Estimation 1999 VLDB 9.7382372e-05
1,756 Semantic Compression and Pattern Extraction with Fascicles 1999 VLDB 9.7140303e-05
2,655 A Robust, Optimization-Based Approach for Approximate Answering of Aggregate Queries 2001 SIGMOD 8.1706092e-05
2,657 SPARTAN: A Model-Based Semantic Compression System for Massive Data Tables 2001 SIGMOD 8.1663451e-05
3,305 Optimal and Approximate Computation of Summary Statistics for Range Aggregates 2001 PODS 7.4424817e-05
Previous Page 1 / 2 Next

Semantically Similar Papers