DBScholar

Back to papers

Random Sampling for Histogram Construction: How much is enough?

Summary: Optimal sampling bounds for equi-height histograms; region-sensitive error metric; adaptive page-level sampling leveraging value clustering. Distinct-value estimation is hard; practical estimator for optimizers; SQL Server 7.0 prototype confirms accuracy. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
he41024073c98c7b0
Venue
SIGMOD
Year
1998
Pagerank
0.00016942879
Overall Rank
519 | 96.52%
DOI
10.1145/276304.276343

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{chaudhuri_sigmod98,
        title = {{Random Sampling for Histogram Construction: How much is enough?}},
        author = {Chaudhuri, Surajit and Motwani, Rajeev and Narasayya, Vivek},
        series = {{SIGMOD} '98},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/276304.276343},
        url = {https://dl.acm.org/doi/10.1145/276304.276343},
        year = {1998}
}

Incoming Citations (Sorted by Pagerank)

Showing 41 of 41 citing papers.

Rank Citing Paper Year Venue Pagerank
20 Similarity Search in High Dimensions via Hashing 1999 VLDB 0.00057568153
26 Models and Issues in Data Stream Systems 2002 PODS 0.00052121228
83 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00035978046
255 The History of Histograms (abridged) 2003 VLDB 0.00022981861
267 Distinct Sampling for Highly-Accurate Answers to Distinct Values Queries and Event Reports 2001 VLDB 0.00022722971
272 An Overview of Query Optimization in Relational Systems 1998 PODS 0.00022509573
372 Approximate Query Processing: Taming the TeraBytes! A Tutorial 2001 VLDB 0.00019720059
378 AutoAdmin "What-if" Index Analysis Utility 1998 SIGMOD 0.00019549382
450 Random Sampling Techniques for Space Efficient Online Computation of Order Statistics of Large Datasets 1999 SIGMOD 0.00018032912
596 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00015785583
840 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.0001354605
997 To Search or to Crawl? Towards a Query Optimizer for Text-Centric Tasks 2006 SIGMOD 0.00012633809
1,082 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012122749
1,741 Combining Histograms and Parametric Curve Fitting for Feedback-Driven Query Result-Size Estimation 1999 VLDB 9.7382372e-05
1,826 Effective Use of Block-Level Sampling in Statistics Estimation 2004 SIGMOD 9.5575424e-05
2,171 Summarizing and Mining Inverse Distributions on Data Streams via Dynamic Inverse Sampling 2005 VLDB 8.9239861e-05
2,433 Cardinality Estimation Using Sample Views with Quality Assurance 2007 SIGMOD 8.4766785e-05
3,211 Comparing Data Streams Using Hamming Norms (How to Zero In) 2002 VLDB 7.536045e-05
3,305 Optimal and Approximate Computation of Summary Statistics for Range Aggregates 2001 PODS 7.4424817e-05
3,984 Data Canopy: Accelerating Exploratory Statistical Analysis 2017 SIGMOD 6.8741188e-05
4,642 Efficient Join Synopsis Maintenance for Data Warehouse 2020 SIGMOD 6.4898745e-05
4,690 Columnstore and B+ tree – Are Hybrid Physical Designs Important? 2018 SIGMOD 6.4687131e-05
4,933 A Comparison of Selectivity Estimators for Range Queries on Metric Attributes 1999 SIGMOD 6.3474817e-05
5,021 Adaptive Sampling for Rapidly Matching Histograms 2018 VLDB 6.3106761e-05
5,288 PolarDB-IMCI: A Cloud-Native HTAP Database System at Alibaba 2023 SIGMOD 6.1944376e-05
5,327 Fast and Approximate Stream Mining of Quantiles and Frequencies Using Graphics Processors 2005 SIGMOD 6.1783009e-05
5,386 PREDIcT: Towards Predicting the Runtime of Large Scale Iterative Analytics 2013 VLDB 6.1522468e-05
5,534 Fast and Near–Optimal Algorithms for Approximating Distributions by Histograms 2015 PODS 6.0923159e-05
5,580 Bloom Histogram: Path Selectivity Estimation for XML Data with Updates 2004 VLDB 6.0769576e-05
5,698 Approximate Quantiles and the Order of the Stream 2006 PODS 6.0317949e-05
6,182 Evaluating Interactive Data Systems: Workloads, Metrics, and Guidelines 2018 SIGMOD 5.8588933e-05
6,639 Approximating and Testing k-Histogram Distributions in Sub-linear Time 2012 PODS 5.7259633e-05
6,685 Query Sampling in DB2 Universal Database 2004 SIGMOD 5.7086005e-05
7,246 Efficient Dynamic Weighted Set Sampling and Its Extension 2024 VLDB 5.5750011e-05
7,364 Weighted Distinct Sampling: Cardinality Estimation for SPJ Queries 2021 SIGMOD 5.5418075e-05
8,576 Histograms as a Side Effect of Data Movement for Big Data 2014 SIGMOD 5.3111243e-05
9,496 Estimating Quantiles from the Union of Historical and Streaming Data 2017 VLDB 5.1708619e-05
10,690 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 4.9793485e-05
11,192 PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models 2025 SIGMOD 4.9793485e-05
11,219 AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse Estimators 2025 VLDB 4.9793485e-05
12,319 Are Few Bins Enough: Testing Histogram Distributions 2016 PODS 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 13 of 13 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers