Back to papers
Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing
Summary: Introduces PASS, Precomputation-Assisted Stratified Sampling: a partitioned tree of partial aggregates to speed up AQP. Exact answers for partition-aligned predicates via DFS; partial overlaps are approximated by stratified samples with an algorithm for near-optimal partitioning.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 6169
- Venue
- SIGMOD
- Year
- 2021
- Pagerank
- 4.9449472e-05
- Overall Rank
- 6,724 | 53.27%
- DOI
-
10.1145/3448016.3457277
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 11 of 11 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 5,206 |
ThalamusDB: Approximate Query Processing on Multi-Modal Data |
2024 |
SIGMOD |
5.625641e-05 |
| 5,405 |
ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic Workloads |
2024 |
VLDB |
5.5243727e-05 |
| 8,341 |
Hierarchical Residual Encoding for Multiresolution Time Series Compression |
2023 |
SIGMOD |
4.5373173e-05 |
| 8,414 |
PairwiseHist: Fast, Accurate and Space-Efficient Approximate Query Processing with Data Compression |
2024 |
VLDB |
4.5135713e-05 |
| 8,521 |
Computing A Well-Representative Summary of Conjunctive Query Results |
2024 |
PODS |
4.4893996e-05 |
| 9,116 |
Towards Observability for Production Machine Learning Pipelines |
2022 |
VLDB |
4.3886184e-05 |
| 9,848 |
Saving Money for Analytical Workloads in the Cloud |
2024 |
VLDB |
4.2680295e-05 |
| 10,223 |
On Fair Epsilon Net and Geometric Hitting Set |
2026 |
VLDB |
4.1905499e-05 |
| 10,371 |
Smallest Synthetic Witnesses for Conjunctive Queries |
2025 |
PODS |
4.1905499e-05 |
| 10,491 |
FAAQP: Fast and Accurate Approximate Query Processing based on Bitmap-augmented Sum-Product Network |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,616 |
Approximation-First Timeseries Query At Scale |
2025 |
VLDB |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 29 of 29 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 14 |
Online Aggregation |
1997 |
SIGMOD |
0.0010813443 |
| 46 |
Simple Random Sampling from Relational Databases |
1986 |
VLDB |
0.00071588702 |
| 326 |
Optimal Histograms with Quality Guarantees |
1998 |
VLDB |
0.0002737538 |
| 398 |
Mergeable Summaries |
2012 |
PODS |
0.00024383201 |
| 606 |
DeepDB: Learn from Data, not from Queries! |
2020 |
VLDB |
0.00019251186 |
| 649 |
Progressive Approximate Aggregate Queries with a Multi-Resolution Tree Structure |
2001 |
SIGMOD |
0.00018652362 |
| 736 |
Congressional Samples for Approximate Answering of Group-By Queries |
2000 |
SIGMOD |
0.00017414831 |
| 752 |
Deep Unsupervised Cardinality Estimation |
2020 |
VLDB |
0.00017138049 |
| 1,116 |
Global Optimization of Histograms |
2001 |
SIGMOD |
0.00013863484 |
| 1,161 |
VerdictDB: Universalizing Approximate Query Processing |
2018 |
SIGMOD |
0.00013579831 |
| 1,257 |
Dynamic Sample Selection for Approximate Query Processing |
2003 |
SIGMOD |
0.00013002384 |
| 1,320 |
Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters |
2016 |
SIGMOD |
0.00012606067 |
| 1,331 |
ICICLES: Self-tuning Samples for Approximate Query Answering |
2000 |
VLDB |
0.00012553948 |
| 1,372 |
Random Sampling over Joins Revisited |
2018 |
SIGMOD |
0.0001233325 |
| 1,473 |
Fine-grained Partitioning for Aggressive Data Skipping |
2014 |
SIGMOD |
0.00011786148 |
| 1,574 |
Approximate Query Processing: No Silver Bullet |
2017 |
SIGMOD |
0.00011289028 |
| 2,177 |
A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data |
2014 |
SIGMOD |
9.371335e-05 |
| 2,583 |
Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee |
2016 |
SIGMOD |
8.4973431e-05 |
| 2,589 |
Database Learning: Toward a Database that Becomes Smarter Every Time |
2017 |
SIGMOD |
8.4868591e-05 |
| 2,813 |
A Robust, Optimization-Based Approach for Approximate Answering of Aggregate Queries |
2001 |
SIGMOD |
8.0816314e-05 |
| 3,944 |
AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics |
2018 |
SIGMOD |
6.6056349e-05 |
| 4,018 |
Optimal Histograms for Hierarchical Range Queries (Extended Abstract) |
2000 |
PODS |
6.5250686e-05 |
| 4,020 |
Revisiting Reuse for Approximate Query Processing |
2017 |
VLDB |
6.5209063e-05 |
| 6,485 |
Robust Estimation With Sampling and Approximate Pre-Aggregation |
2003 |
VLDB |
5.0386161e-05 |
| 7,246 |
Learning to Sample: Counting with Complex Queries |
2020 |
VLDB |
4.7847433e-05 |
| 8,139 |
Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints |
2020 |
SIGMOD |
4.5727142e-05 |
| 8,235 |
Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters |
2019 |
VLDB |
4.5481384e-05 |
| 8,669 |
CoopStore: Optimizing Precomputed Summaries for Aggregation |
2020 |
VLDB |
4.4667395e-05 |
| 8,703 |
Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views |
2015 |
VLDB |
4.4596255e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 649 |
Progressive Approximate Aggregate Queries with a Multi-Resolution Tree Structure |
2001 |
SIGMOD |
0.00018652362 |
| 2,813 |
A Robust, Optimization-Based Approach for Approximate Answering of Aggregate Queries |
2001 |
SIGMOD |
8.0816314e-05 |
| 1,257 |
Dynamic Sample Selection for Approximate Query Processing |
2003 |
SIGMOD |
0.00013002384 |
| 1,867 |
Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems |
2014 |
SIGMOD |
0.00010264932 |
| 11,287 |
Approximate Queries over Concurrent Updates |
2023 |
VLDB |
4.1905499e-05 |
| 10,049 |
Approximate Query Processing under Updates |
2026 |
SIGMOD |
4.1905499e-05 |
| 3,944 |
AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics |
2018 |
SIGMOD |
6.6056349e-05 |
| 6,481 |
Joins on Samples: A Theoretical Guide for Practitioners |
2020 |
VLDB |
5.039683e-05 |
| 10,349 |
Efficient Approximate Query Processing with Block Sampling |
2025 |
CIDR |
4.1905499e-05 |
| 2,583 |
Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee |
2016 |
SIGMOD |
8.4973431e-05 |