DBScholar

Back to papers

Sampling-Based Estimation of the Number of Distinct Values of an Attribute

Summary: Introduces sampling estimators for attribute cardinality, including a skew-aware hybrid combining smoothed jackknife and Shlosser methods. First broad empirical comparison on highly skewed real database distributions; hybrid achieves best average precision. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
8469
Venue
VLDB
Year
1995
Pagerank
0.00037277061
Overall Rank
75 | 99.49%
DOI
-

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{haas_vldb95,
        title = {{Sampling-Based Estimation of the Number of Distinct Values of an Attribute}},
        author = {Haas, Peter J. and Naughton, Jeffrey F. and Seshadri, S. and Stokes, Lynne},
        journal = {PVLDB},
        series = {{VLDB} '95},
        pages = {311},
        year = {1995}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 59 citing papers.

Rank Citing Paper Year Venue Pagerank
11 Implementing Data Cubes Efficiently 1996 SIGMOD 0.00071822821
26 Models and Issues in Data Stream Systems 2002 PODS 0.00052982574
35 Improved Histograms for Selectivity Estimation of Range Predicates 1996 SIGMOD 0.00048481081
136 Join Synopses for Approximate Query Answering 1999 SIGMOD 0.00030123303
149 New Sampling-Based Summary Statistics for Improving Approximate Query Answers 1998 SIGMOD 0.00029226907
186 The Vertica Analytic Database: C-Store 7 Years Later 2012 VLDB 0.00026182534
207 On the Computation of Multidimensional Aggregates 1996 VLDB 0.00025088003
255 Distinct Sampling for Highly-Accurate Answers to Distinct Values Queries and Event Reports 2001 VLDB 0.00023174541
288 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00022296371
290 An Overview of Query Optimization in Relational Systems 1998 PODS 0.0002227038
363 Approximate Query Processing: Taming the TeraBytes! A Tutorial 2001 VLDB 0.0002005475
418 Tracking Join and Self-Join Sizes in Limited Storage 1999 PODS 0.00018812821
508 Random Sampling for Histogram Construction: How much is enough? 1998 SIGMOD 0.00017275873
566 Towards a Robust Query Optimizer: A Principled and Practical Approach 2005 SIGMOD 0.00016436005
654 Storage Estimation for Multidimensional Aggregates in the Presence of Hierarchies 1996 VLDB 0.0001527187
723 Dynamic Multidimensional Histograms 2002 SIGMOD 0.00014620977
730 Bifocal Sampling for Skew-Resistant Join Size Estimation 1996 SIGMOD 0.00014539362
801 Query Execution Techniques for Caching Expensive Methods 1996 SIGMOD 0.00013909408
1,053 Multi-dimensional Selectivity Estimation Using Compressed Histogram Information 1999 SIGMOD 0.00012401532
1,266 Compressing SQL Workloads 2002 SIGMOD 0.00011412078
1,516 Cardinality Estimation: An Experimental Survey 2018 VLDB 0.00010520885
1,736 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.8984415e-05
1,806 Effective Use of Block-Level Sampling in Statistics Estimation 2004 SIGMOD 9.7112151e-05
1,994 RainForest - A Framework for Fast Decision Tree Construction of Large Datasets 1998 VLDB 9.3409624e-05
2,127 Selectivity Estimation in Spatial Databases 1999 SIGMOD 9.127762e-05
2,281 Selectivity Estimation in Extensible Databases - A Neural Network Approach 1998 VLDB 8.8129279e-05
2,404 Cardinality Estimation Using Sample Views with Quality Assurance 2007 SIGMOD 8.6225576e-05
2,633 Relational Confidence Bounds Are Easy With The Bootstrap* 2005 SIGMOD 8.3224527e-05
2,898 Approximate Selection with Guarantees using Proxies 2020 VLDB 7.978725e-05
2,997 Adapting to Source Properties in Processing Data Integration Queries 2004 SIGMOD 7.8745158e-05
3,150 Comparing Data Streams Using Hamming Norms (How to Zero In) 2002 VLDB 7.7055991e-05
3,157 Turbo-Charging Estimate Convergence in DBO 2009 VLDB 7.6911286e-05
3,160 Graph Cube: On Warehousing and OLAP Multidimensional Networks 2011 SIGMOD 7.6823631e-05
3,197 Processing Set Expressions over Continuous Update Streams 2003 SIGMOD 7.6439415e-05
3,215 Every Row Counts: Combining Sketches and Sampling for Accurate Group-By Result Estimates 2019 CIDR 7.6324234e-05
3,228 Correlation Sketches for Approximate Join-Correlation Queries 2021 SIGMOD 7.6215176e-05
4,222 Density Biased Sampling: An Improved Method for Data Mining and Clustering 2000 SIGMOD 6.8227421e-05
4,409 MNC: Structure-Exploiting Sparsity Estimation for Matrix Expressions 2019 SIGMOD 6.7178579e-05
4,428 Arnold: Declarative Crowd-Machine Data Integration 2013 CIDR 6.7094035e-05
4,817 Efficiently Approximating Query Optimizer Plan Diagrams 2008 VLDB 6.4944225e-05
5,193 Sampling Algorithms in a Stream Operator 2005 SIGMOD 6.3238562e-05
5,470 Efficient Computation of Multiple Group By Queries 2005 SIGMOD 6.2070458e-05
5,537 Uncertainty Aware Query Execution Time Prediction 2014 VLDB 6.1820087e-05
5,604 Approximate Quantiles and the Order of the Stream 2006 PODS 6.1533846e-05
6,053 Modeling skewed distributions using multifractals and the '80-20 law' 1996 VLDB 5.9939282e-05
6,434 Yannakakis+: Practical Acyclic Query Evaluation with Theoretical Guarantees 2025 SIGMOD 5.8799421e-05
6,819 Estimating the Impact of Unknown Unknowns on Aggregate Query Results 2016 SIGMOD 5.7635226e-05
7,290 Learning to be a Statistician: Learned Estimator for Number of Distinct Values 2022 VLDB 5.6540503e-05
7,377 Efficient and Scalable Statistics Gathering for Large Databases in Oracle 11g 2008 SIGMOD 5.6297042e-05
7,770 Automated design of multidimensional clustering tables for relational databases 2004 VLDB 5.5474421e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 4 of 4 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers