DBScholar

Back to papers

Approximate Query Processing: Taming the TeraBytes! A Tutorial

Summary: Tutorial on approximate aggregate query processing via online sampling and precomputed synopses with explicit error guarantees. Covers multidimensional data, joins, set-valued queries, Aqua-style architectures, streaming, dependencies, and workload-tuned synopsis design. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
9015
Venue
VLDB
Year
2001
Pagerank
0.0002005475
Overall Rank
363 | 97.52%
DOI
-

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{garofalakis_vldb01,
        title = {{Approximate Query Processing: Taming the TeraBytes! A Tutorial}},
        author = {Garofalakis, Minos and Gibbons, Phillip B.},
        journal = {PVLDB},
        series = {{VLDB} '01},
        pages = {169},
        year = {2001}
}

Incoming Citations (Sorted by Pagerank)

Showing 41 of 41 citing papers.

Rank Citing Paper Year Venue Pagerank
125 Trio: A System for Integrated Management of Data, Accuracy, and Lineage 2005 CIDR 0.00030932014
337 Model-Driven Data Acquisition in Sensor Networks 2004 VLDB 0.00020783399
432 Mining Database Structure; Or, How to Build a Data Quality Browser 2002 SIGMOD 0.00018572055
569 BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics 2020 VLDB 0.00016348191
817 Processing Complex Aggregate Queries over Data Streams 2002 SIGMOD 0.00013823702
828 The Design of an Acquisitional Query Processor For Sensor Networks 2003 SIGMOD 0.00013769869
1,108 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012145154
1,147 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011974846
1,227 Blink and It's Done: Interactive Queries on Very Large Data 2012 VLDB 0.00011582387
1,634 Rapid Sampling for Visualizations with Ordering Guarantees 2015 VLDB 0.00010163938
1,736 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.8984415e-05
1,962 Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee 2016 SIGMOD 9.3978414e-05
2,280 Partial Results for Online Query Processing 2002 SIGMOD 8.8129961e-05
2,367 Using Probabilistic Models for Data Management in Acquisitional Environments 2005 CIDR 8.6855754e-05
2,816 Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation 2025 SIGMOD 8.0959781e-05
2,987 Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations 2019 SIGMOD 7.8907997e-05
3,072 Lux: Always-on Visualization Recommendations for Exploratory Dataframe Workflows 2022 VLDB 7.7873864e-05
3,288 I've Seen "Enough": Incrementally Improving Visualizations to Support Rapid Decision Making 2017 VLDB 7.5578177e-05
3,945 Exploiting Correlations for Expensive Predicate Evaluation 2015 SIGMOD 7.0055154e-05
4,264 A Method for Optimizing Opaque Filter Queries 2020 SIGMOD 6.7937529e-05
4,581 Mining Graph Patterns Efficiently via Randomized Summaries 2009 VLDB 6.6198548e-05
4,837 PrivateClean: Data Cleaning and Differential Privacy 2016 SIGMOD 6.4845444e-05
5,868 A Random Walk Approach to Sampling Hidden Databases 2007 SIGMOD 6.0613238e-05
6,038 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 5.9990929e-05
6,210 Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First 2026 CIDR 5.9425753e-05
6,712 Capturing Data Uncertainty in High-Volume Stream Processing 2009 CIDR 5.7962014e-05
7,389 Benchmarking Spreadsheet Systems 2020 SIGMOD 5.6266966e-05
7,903 Mining a Search Engine’s Corpus: Efficient Yet Unbiased Sampling and Aggregate Estimation 2011 SIGMOD 5.5182155e-05
8,714 Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views 2015 VLDB 5.3778009e-05
9,009 Efficient Approximations of Conjunctive Queries 2012 PODS 5.3335287e-05
9,650 Aggregate Estimation Over Dynamic Hidden Web Databases 2014 VLDB 5.2427667e-05
9,749 Auto-Approximation of Graph Computing 2014 VLDB 5.227679e-05
10,342 Approximate Query Processing under Updates 2026 SIGMOD 5.093636e-05
10,404 Stochastic Submodular Data Forgetting 2026 SIGMOD 5.093636e-05
10,504 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 5.093636e-05
11,109 FaDE: More Than a Million What-ifs Per Second 2025 VLDB 5.093636e-05
11,700 In the Land of Data Streams where Synopses are Missing, One Framework to Bring Them All 2021 VLDB 5.093636e-05
11,886 Enabling Data Science for the Majority 2019 VLDB 5.093636e-05
12,033 A Study of Sorting Algorithms on Approximate Memory 2016 SIGMOD 5.093636e-05
12,216 When Data Management Systems Meet Approximate Hardware: Challenges and Opportunities 2014 VLDB 5.093636e-05
12,699 AQAX: A System for Approximate XML Query Answers 2006 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 50 of 52 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1 Access Path Selection in a Relational Database Management System 1979 SIGMOD 0.0024089429
9 Online Aggregation 1997 SIGMOD 0.00077458002
35 Improved Histograms for Selectivity Estimation of Range Predicates 1996 SIGMOD 0.00048481081
36 Accurate Estimation Of The Number Of Tuples Satisfying A Condition 1984 SIGMOD 0.00048351457
54 On Random Sampling over Joins 1999 SIGMOD 0.00040810225
55 Statistical Estimators for Relational Algebra Expressions 1988 PODS 0.00040746149
75 Sampling-Based Estimation of the Number of Distinct Values of an Attribute 1995 VLDB 0.00037277061
76 Practical Selectivity Estimation through Adaptive Sampling 1990 SIGMOD 0.00037054261
82 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00036378991
89 On the Propagation of Errors in the Size of Join Results 1991 SIGMOD 0.00035031529
101 Selectivity Estimation Without the Attribute Value Independence Assumption 1997 VLDB 0.00034376651
118 Equi-Depth Histograms For Estimating Selectivity Factors For Multi-Dimensional Queries 1988 SIGMOD 0.00031922279
131 Ripple Joins for Online Aggregation 1999 SIGMOD 0.00030424509
136 Join Synopses for Approximate Query Answering 1999 SIGMOD 0.00030123303
149 New Sampling-Based Summary Statistics for Improving Approximate Query Answers 1998 SIGMOD 0.00029226907
156 An Efficient, Cost-Driven Index Selection Tool for Microsoft SQL Server 1997 VLDB 0.00028636811
168 Wavelet-Based Histograms for Selectivity Estimation 1998 SIGMOD 0.00027541029
178 Processing Aggregate Relational Queries with Hard Time Constraints 1989 SIGMOD 0.00026881845
213 Approximate Computation of Multidimensional Aggregates of Sparse Data Using Wavelets 1999 SIGMOD 0.00024723025
222 Adaptive Selectivity Estimation Using Query Feedback 1994 SIGMOD 0.00024193708
235 Fast Incremental Maintenance of Approximate Histograms 1997 VLDB 0.00023783792
255 Distinct Sampling for Highly-Accurate Answers to Distinct Values Queries and Event Reports 2001 VLDB 0.00023174541
267 Optimal Histograms with Quality Guarantees 1998 VLDB 0.00022798161
274 Balancing Histogram Optimality and Practicality for Query Result Size Estimation 1995 SIGMOD 0.00022645621
280 Selectivity Estimation using Probabilistic Models 2001 SIGMOD 0.00022454217
288 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00022296371
307 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021792475
311 Surfing Wavelets on Streams: One-Pass Summaries for Approximate Aggregate Queries 2001 VLDB 0.00021760621
365 STHoles: A Multidimensional Workload-Aware Histogram 2001 SIGMOD 0.00020041735
387 AutoAdmin "What-if" Index Analysis Utility 1998 SIGMOD 0.00019442332
418 Tracking Join and Self-Join Sizes in Limited Storage 1999 PODS 0.00018812821
435 Histogram-Based Approximation of Set-Valued Query Answers 1999 VLDB 0.000185063
448 Self-tuning Histograms: Building Histograms Without Looking at Data 1999 SIGMOD 0.00018292618
508 Random Sampling for Histogram Construction: How much is enough? 1998 SIGMOD 0.00017275873
551 On Computing Correlated Aggregates Over Continual Data Streams 2001 SIGMOD 0.00016635191
553 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016590619
610 Efficiently Supporting Ad Hoc Queries in Large Datasets of Time Sequences 1997 SIGMOD 0.00015783259
648 Progressive Approximate Aggregate Queries with a Multi-Resolution Tree Structure 2001 SIGMOD 0.00015324657
692 Independence is Good: Dependency-Based Histogram Synopses for High-Dimensional Data 2001 SIGMOD 0.00014919816
730 Bifocal Sampling for Skew-Resistant Join Size Estimation 1996 SIGMOD 0.00014539362
786 Universality of Serial Histograms 1993 VLDB 0.00014053885
850 Approximating Multi-Dimensional Aggregate Range Queries Over Real Attributes 2000 SIGMOD 0.00013619394
977 Dynamic Maintenance of Wavelet-Based Histograms 2000 VLDB 0.00012864017
1,053 Multi-dimensional Selectivity Estimation Using Compressed Histogram Information 1999 SIGMOD 0.00012401532
1,166 ICICLES: Self-tuning Samples for Approximate Query Answering 2000 VLDB 0.00011850439
1,721 Semantic Compression and Pattern Extraction with Fascicles 1999 VLDB 9.9258415e-05
1,729 Combining Histograms and Parametric Curve Fitting for Feedback-Driven Query Result-Size Estimation 1999 VLDB 9.908788e-05
2,608 A Robust, Optimization-Based Approach for Approximate Answering of Aggregate Queries 2001 SIGMOD 8.347674e-05
2,611 SPARTAN: A Model-Based Semantic Compression System for Massive Data Tables 2001 SIGMOD 8.34729e-05
3,241 Optimal and Approximate Computation of Summary Statistics for Range Aggregates 2001 PODS 7.6057671e-05
Previous Page 1 / 2 Next

Semantically Similar Papers