Back to papers
Learning to be a Statistician: Learned Estimator for Number of Distinct Values
Summary: Proposes a supervised-learning NDV estimator, replacing heuristic sample methods with a data-driven model. Trains on synthetic data for workload-agnostic deployment as a UDF, outperforming existing estimators on nine real datasets; code available.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 12760
- Venue
- VLDB
- Year
- 2022
- Pagerank
- 4.6920008e-05
- Overall Rank
- 7,611 | 47.11%
- DOI
-
10.14778/3489496.3489508
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 60 |
Sampling-Based Estimation of the Number of Distinct Values of an Attribute |
1995 |
VLDB |
0.00064450997 |
| 203 |
Learned Cardinalities: Estimating Correlated Joins with Deep Learning |
2019 |
CIDR |
0.00034868567 |
| 380 |
Towards Estimation Error Guarantees for Distinct Values |
2000 |
PODS |
0.00024943236 |
| 606 |
DeepDB: Learn from Data, not from Queries! |
2020 |
VLDB |
0.00019251186 |
| 1,239 |
Selectivity Estimation for Range Predicates using Lightweight Models |
2019 |
VLDB |
0.00013091459 |
| 1,574 |
Approximate Query Processing: No Silver Bullet |
2017 |
SIGMOD |
0.00011289028 |
| 1,683 |
Cardinality Estimation: An Experimental Survey |
2018 |
VLDB |
0.0001091276 |
| 1,699 |
Are We Ready For Learned Cardinality Estimation? |
2021 |
VLDB |
0.00010848882 |
| 2,769 |
FLAT: Fast, Lightweight and Accurate Method for Cardinality Estimation |
2021 |
VLDB |
8.1512848e-05 |
| 2,844 |
Selectivity Estimation in Extensible Databases - A Neural Network Approach |
1998 |
VLDB |
8.0308994e-05 |
| 2,971 |
Estimating Join Selectivities using Bandwidth-Optimized Kernel Density Models |
2017 |
VLDB |
7.7935535e-05 |
| 3,955 |
Efficiently Approximating Selectivity Functions using Low Overhead Regression Models |
2020 |
VLDB |
6.5895015e-05 |
| 6,243 |
Approximate Distinct Counts for Billions of Datasets |
2019 |
SIGMOD |
5.1348218e-05 |
Semantically Similar Papers