DBScholar

Back to papers

Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee

Summary: Proposes distribution precision as a strict error guarantee for AQP of group-by aggregates, enabling distribution-level accuracy rather than point estimates. Introduces measure-biased sampling and two in-memory indexes to support selective predicates and any aggregate estimate, delivering ~100x speedups with ~5% distribution error vs a commercial DB. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h00be92bca8c23e8f
Venue
SIGMOD
Year
2016
Pagerank
9.2112617e-05
Overall Rank
2,000 | 86.56%
DOI
10.1145/2882903.2915249

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{ding_sigmod16,
        title = {{Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee}},
        author = {Ding, Bolin and Huang, Silu and Chaudhuri, Surajit and Chakrabarti, Kaushik and Wang, Chi},
        series = {{SIGMOD} '16},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2882903.2915249},
        url = {https://dl.acm.org/doi/10.1145/2882903.2915249},
        year = {2016}
}

Incoming Citations (Sorted by Pagerank)

Showing 29 of 29 citing papers.

Rank Citing Paper Year Venue Pagerank
795 Random Sampling over Joins Revisited 2018 SIGMOD 0.00013938779
1,082 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012122749
1,369 Towards Scalable Dataframe Systems 2020 VLDB 0.00010899832
3,131 Every Row Counts: Combining Sketches and Sampling for Accurate Group-By Result Estimates 2019 CIDR 7.6141006e-05
3,424 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.3117029e-05
4,253 Sample Debiasing in the Themis Open World Database System 2020 SIGMOD 6.7015395e-05
4,304 Data Series Progressive Similarity Search with Probabilistic Quality Guarantees 2020 SIGMOD 6.6771697e-05
5,021 Adaptive Sampling for Rapidly Matching Histograms 2018 VLDB 6.3106761e-05
5,373 At-the-time and Back-in-time Persistent Sketches 2021 SIGMOD 6.1569337e-05
5,930 Hillview: A trillion-cell spreadsheet for big data 2019 VLDB 5.9438354e-05
6,221 Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing 2021 SIGMOD 5.8463347e-05
6,316 Visualization-aware Time Series Min-Max Caching with Error Bound Guarantees 2024 VLDB 5.8160796e-05
6,694 MOST: Model-Based Compression with Outlier Storage for Time Series Data 2023 SIGMOD 5.7068014e-05
6,875 Towards Democratizing Relational Data Visualization 2019 SIGMOD 5.6589849e-05
7,210 Marviq: Quality-Aware Geospatial Visualization of Range-Selection Queries Using Materialization 2020 SIGMOD 5.5844776e-05
7,656 Continuous Prefetch for Interactive Data Applications 2020 VLDB 5.4772833e-05
7,978 Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines 2024 VLDB 5.4136835e-05
8,283 Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters 2019 VLDB 5.3627138e-05
8,363 Probabilistic Database Summarization for Interactive Data Exploration 2017 VLDB 5.3472423e-05
8,659 ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data Generation 2023 VLDB 5.2930951e-05
8,770 One Size Does Not Fit All: A Bandit-Based Sampler Combination Framework with Theoretical Guarantees 2022 SIGMOD 5.2812395e-05
8,846 Data Driven Approximation with Bounded Resources 2017 VLDB 5.2649328e-05
8,944 Practical Dynamic Extension for Sampling Indexes 2023 SIGMOD 5.2542285e-05
10,315 AB-tree: Index for Concurrent Random Sampling and Updates 2022 VLDB 5.0377739e-05
10,606 Visualization-Oriented Progressive Time Series Transformation 2026 SIGMOD 4.9793485e-05
10,724 Secure Multi-Party Sampling over Joins 2026 VLDB 4.9793485e-05
11,184 FAAQP: Fast and Accurate Approximate Query Processing based on Bitmap-augmented Sum-Product Network 2025 SIGMOD 4.9793485e-05
11,794 Approximate Queries over Concurrent Updates 2023 VLDB 4.9793485e-05
12,039 FlashP: An Analytical Pipeline for Real-time Forecasting of Time-Series Relational Data 2021 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
9 Online Aggregation 1997 SIGMOD 0.00076195956
88 Automated Selection of Materialized Views and Indexes for SQL Databases 2000 VLDB 0.00035351639
222 Approximate Computation of Multidimensional Aggregates of Sparse Data Using Wavelets 1999 SIGMOD 0.00024218831
295 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00021914399
309 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021384073
336 The Aqua Approximate Query Answering System 1999 SIGMOD 0.00020657819
372 Approximate Query Processing: Taming the TeraBytes! A Tutorial 2001 VLDB 0.00019720059
564 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016296665
931 Dynamic Sample Selection for Approximate Query Processing 2003 SIGMOD 0.00013011667
1,022 Online Aggregation for Large MapReduce Jobs 2011 VLDB 0.00012438826
1,183 ICICLES: Self-tuning Samples for Approximate Query Answering 2000 VLDB 0.00011616705
1,428 Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems 2014 SIGMOD 0.00010693831
1,607 SciBORQ: Scientific data management with Bounds On Runtime and Quality 2011 CIDR 0.0001008742
1,661 Rapid Sampling for Visualizations with Ordering Guarantees 2015 VLDB 9.9535453e-05
1,916 The Analytical Bootstrap: a New Method for Fast Error Estimation in Approximate Query Processing 2014 SIGMOD 9.3837729e-05
2,224 DAQ: A New Paradigm for Approximate Query Processing 2015 VLDB 8.80823e-05
2,655 A Robust, Optimization-Based Approach for Approximate Answering of Aggregate Queries 2001 SIGMOD 8.1706092e-05
3,305 Optimal and Approximate Computation of Summary Statistics for Range Aggregates 2001 PODS 7.4424817e-05
3,560 Interactive Analysis of Web-Scale Data 2009 CIDR 7.2072209e-05
5,534 Fast and Near–Optimal Algorithms for Approximating Distributions by Histograms 2015 PODS 6.0923159e-05
Previous Page 1 / 1 Next

Semantically Similar Papers