DBScholar

Back to papers

Random Sampling for Histogram Construction: How much is enough?

Summary: Optimal sampling bounds for equi-height histograms; region-sensitive error metric; adaptive page-level sampling leveraging value clustering. Distinct-value estimation is hard; practical estimator for optimizers; SQL Server 7.0 prototype confirms accuracy. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
3096
Venue
SIGMOD
Year
1998
Pagerank
0.00017275873
Overall Rank
508 | 96.52%
DOI
10.1145/276304.276343

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{chaudhuri_sigmod98,
        title = {{Random Sampling for Histogram Construction: How much is enough?}},
        author = {Chaudhuri, Surajit and Motwani, Rajeev and Narasayya, Vivek},
        series = {{SIGMOD} '98},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/276304.276343},
        url = {https://dl.acm.org/doi/10.1145/276304.276343},
        year = {1998}
}

Incoming Citations (Sorted by Pagerank)

Showing 41 of 41 citing papers.

Rank Citing Paper Year Venue Pagerank
21 Similarity Search in High Dimensions via Hashing 1999 VLDB 0.00056760516
26 Models and Issues in Data Stream Systems 2002 PODS 0.00052982574
82 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00036378991
255 Distinct Sampling for Highly-Accurate Answers to Distinct Values Queries and Event Reports 2001 VLDB 0.00023174541
257 The History of Histograms (abridged) 2003 VLDB 0.00023154793
290 An Overview of Query Optimization in Relational Systems 1998 PODS 0.0002227038
363 Approximate Query Processing: Taming the TeraBytes! A Tutorial 2001 VLDB 0.0002005475
387 AutoAdmin "What-if" Index Analysis Utility 1998 SIGMOD 0.00019442332
443 Random Sampling Techniques for Space Efficient Online Computation of Order Statistics of Large Datasets 1999 SIGMOD 0.00018373044
593 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00016027871
819 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013815639
974 To Search or to Crawl? Towards a Query Optimizer for Text-Centric Tasks 2006 SIGMOD 0.00012870746
1,108 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012145154
1,729 Combining Histograms and Parametric Curve Fitting for Feedback-Driven Query Result-Size Estimation 1999 VLDB 9.908788e-05
1,806 Effective Use of Block-Level Sampling in Statistics Estimation 2004 SIGMOD 9.7112151e-05
2,128 Summarizing and Mining Inverse Distributions on Data Streams via Dynamic Inverse Sampling 2005 VLDB 9.1271562e-05
2,404 Cardinality Estimation Using Sample Views with Quality Assurance 2007 SIGMOD 8.6225576e-05
3,150 Comparing Data Streams Using Hamming Norms (How to Zero In) 2002 VLDB 7.7055991e-05
3,241 Optimal and Approximate Computation of Summary Statistics for Range Aggregates 2001 PODS 7.6057671e-05
3,930 Data Canopy: Accelerating Exploratory Statistical Analysis 2017 SIGMOD 7.0102082e-05
4,630 Efficient Join Synopsis Maintenance for Data Warehouse 2020 SIGMOD 6.5955933e-05
4,637 Columnstore and B+ tree – Are Hybrid Physical Designs Important? 2018 SIGMOD 6.592197e-05
4,846 A Comparison of Selectivity Estimators for Range Queries on Metric Attributes 1999 SIGMOD 6.4807645e-05
4,926 Adaptive Sampling for Rapidly Matching Histograms 2018 VLDB 6.4428362e-05
5,216 Fast and Approximate Stream Mining of Quantiles and Frequencies Using Graphics Processors 2005 SIGMOD 6.3121293e-05
5,287 PREDIcT: Towards Predicting the Runtime of Large Scale Iterative Analytics 2013 VLDB 6.2815429e-05
5,421 Fast and Near–Optimal Algorithms for Approximating Distributions by Histograms 2015 PODS 6.2245831e-05
5,459 Bloom Histogram: Path Selectivity Estimation for XML Data with Updates 2004 VLDB 6.2108888e-05
5,477 PolarDB-IMCI: A Cloud-Native HTAP Database System at Alibaba 2023 SIGMOD 6.2051869e-05
5,604 Approximate Quantiles and the Order of the Stream 2006 PODS 6.1533846e-05
6,059 Evaluating Interactive Data Systems: Workloads, Metrics, and Guidelines 2018 SIGMOD 5.9915079e-05
6,510 Approximating and Testing k-Histogram Distributions in Sub-linear Time 2012 PODS 5.8569033e-05
6,562 Query Sampling in DB2 Universal Database 2004 SIGMOD 5.838575e-05
7,142 Efficient Dynamic Weighted Set Sampling and Its Extension 2024 VLDB 5.6917227e-05
7,256 Weighted Distinct Sampling: Cardinality Estimation for SPJ Queries 2021 SIGMOD 5.6625146e-05
8,420 Histograms as a Side Effect of Data Movement for Big Data 2014 SIGMOD 5.4279873e-05
9,318 Estimating Quantiles from the Union of Historical and Streaming Data 2017 VLDB 5.289545e-05
10,504 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 5.093636e-05
10,774 PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models 2025 SIGMOD 5.093636e-05
10,806 AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse Estimators 2025 VLDB 5.093636e-05
12,024 Are Few Bins Enough: Testing Histogram Distributions 2016 PODS 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 13 of 13 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers