Towards Estimation Error Guarantees for Distinct Values
Summary: Prove any sublinear-sample estimator for distinct counts must suffer large error on some natural distributions unless it reads a large fraction of the data. Give an estimator matching this lower bound and practical heuristics for typical distributions, validated empirically. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Moses Charikar (Stanford University)
- 2. Surajit Chaudhuri (Microsoft)
- 3. Rajeev Motwani (Stanford University)
- 4. Vivek Narasayya (Microsoft)
BibTeX Citation
@inproceedings{charikar_pods00,
address = {New York, NY, USA},
series = {{PODS} '00},
title = {{Towards Estimation Error Guarantees for Distinct Values}},
url = {https://dl.acm.org/doi/10.1145/335168.335230},
doi = {10.1145/335168.335230},
booktitle = {Proceedings of the {ACM} {SIGMOD} Symposium on {Principles} of {Database} {Systems}},
publisher = {Association for Computing Machinery},
author = {Charikar, Moses and Chaudhuri, Surajit and Motwani, Rajeev and Narasayya, Vivek},
year = {2000}
}
Incoming Citations (Sorted by Pagerank)
Showing 50 of 59 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 35 | Improved Histograms for Selectivity Estimation of Range Predicates | 1996 | SIGMOD | 0.00048481081 |
| 55 | Statistical Estimators for Relational Algebra Expressions | 1988 | PODS | 0.00040746149 |
| 75 | Sampling-Based Estimation of the Number of Distinct Values of an Attribute | 1995 | VLDB | 0.00037277061 |
| 156 | An Efficient, Cost-Driven Index Selection Tool for Microsoft SQL Server | 1997 | VLDB | 0.00028636811 |
| 178 | Processing Aggregate Relational Queries with Hard Time Constraints | 1989 | SIGMOD | 0.00026881845 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,215 | Every Row Counts: Combining Sketches and Sampling for Accurate Group-By Result Estimates | 2019 | CIDR |
| 2 | 11,897 | Distinct Sampling on Streaming Data with Near-Duplicates | 2018 | PODS |
| 3 | 8,998 | Histograms Reloaded: The Merits of Bucket Diversity | 2010 | SIGMOD |
| 4 | 12,362 | Get the Most out of Your Sample: Optimal Unbiased Estimators using Partial Information | 2011 | PODS |
| 5 | 5,998 | Approximate Distinct Counts for Billions of Datasets | 2019 | SIGMOD |
| 6 | 689 | On Synopses for Distinct-Value Estimation Under Multiset Operations | 2007 | SIGMOD |
| 7 | 255 | Distinct Sampling for Highly-Accurate Answers to Distinct Values Queries and Event Reports | 2001 | VLDB |
| 8 | 7,290 | Learning to be a Statistician: Learned Estimator for Number of Distinct Values | 2022 | VLDB |
| 9 | 1,806 | Effective Use of Block-Level Sampling in Statistics Estimation | 2004 | SIGMOD |
| 10 | 75 | Sampling-Based Estimation of the Number of Distinct Values of an Attribute | 1995 | VLDB |