Back to papers
PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models
Summary: Leverages semantic schema via pre-trained language models to estimate the number of distinct values (NDV) with reduced data access. PLM4NDV fuses target-column and table semantics to lower access costs, can operate with no data access, and outperforms baselines on large real-world datasets.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
7317
Venue
SIGMOD
Year
2025
Pagerank
5.093636e-05
Overall Rank
10,774 | 26.09%
DOI
10.1145/3725336
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
BibTeX Citation
Copy BibTeX
@inproceedings{xu_sigmod25,
title = {{PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models}},
author = {Xu, Xianghong and He, Xiao and Zhang, Tieying and Zhang, Lei and Shi, Rui and Chen, Jianjun},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3725336},
url = {https://dl.acm.org/doi/10.1145/3725336},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
Outgoing Citations (Sorted by Pagerank)
Showing 26 of 26 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
1
Access Path Selection in a Relational Database Management System
1979
SIGMOD
0.0024089429
75
Sampling-Based Estimation of the Number of Distinct Values of an Attribute
1995
VLDB
0.00037277061
141
Deep Entity Matching with Pre-Trained Language Models
2021
VLDB
0.0002964847
279
Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation
2024
VLDB
0.00022468369
288
Towards Estimation Error Guarantees for Distinct Values
2000
PODS
0.00022296371
388
Preventing Bad Plans by Bounding the Impact of Cardinality Estimation Errors
2009
VLDB
0.00019410042
397
TURL: Table Understanding through Representation Learning
2021
VLDB
0.00019278189
508
Random Sampling for Histogram Construction: How much is enough?
1998
SIGMOD
0.00017275873
689
On Synopses for Distinct-Value Estimation Under Multiset Operations
2007
SIGMOD
0.00014940023
1,516
Cardinality Estimation: An Experimental Survey
2018
VLDB
0.00010520885
1,806
Effective Use of Block-Level Sampling in Statistics Estimation
2004
SIGMOD
9.7112151e-05
1,920
D-Bot: Database Diagnosis System using Large Language Models
2024
VLDB
9.4846185e-05
1,923
Annotating Columns with Pre-trained Language Models
2022
SIGMOD
9.4789109e-05
2,748
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning
2023
SIGMOD
8.1707811e-05
2,790
GitTables: A Large-Scale Corpus of Relational Tables
2023
SIGMOD
8.1200509e-05
4,349
ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic Workloads
2024
VLDB
6.7504619e-05
4,363
GenRewrite: Query Rewriting via Large Language Models
2026
SIGMOD
6.7423909e-05
4,643
Learned Index Benefits: Machine Learning Based Index Performance Estimation
2022
VLDB
6.5907466e-05
6,920
UltraLogLog: A Practical and More Space-Efficient Alternative to HyperLogLog for Approximate Distinct Counting
2024
VLDB
5.7390922e-05
7,076
Refactoring Index Tuning Process with Benefit Estimation
2024
VLDB
5.7098893e-05
7,290
Learning to be a Statistician: Learned Estimator for Number of Distinct Values
2022
VLDB
5.6540503e-05
8,161
LAQy: Efficient and Reusable Query Approximations via Lazy Sampling
2023
SIGMOD
5.4752972e-05
8,692
FormaT5: Abstention and Examples for Conditional Table Formatting with Natural Language
2024
VLDB
5.3838716e-05
8,849
ByteCard: Enhancing ByteDance’s Data Warehouse with Learned Cardinality Estimation
2024
SIGMOD
5.3577504e-05
8,850
Learning-based Property Estimation with Polynomials
2024
SIGMOD
5.3577504e-05
10,806
AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse Estimators
2025
VLDB
5.093636e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
10,310
A Fast, Mergeable, and LDP Compatible Sketch for Counting the Number of Distinct Values in Fully Dynamic Tables
2026
SIGMOD
2
11,186
Unstructured Data Fusion for Schema and Data Extraction
2024
SIGMOD
3
5,792
Pre-training Summarization Models of Structured Datasets for Cardinality Estimation
2022
VLDB
4
1,573
Deep Learning Models for Selectivity Estimation of Multi-Attribute Queries
2020
SIGMOD
5
6,760
LPLM: A Neural Language Model for Cardinality Estimation of LIKE-Queries
2024
SIGMOD
6
8,850
Learning-based Property Estimation with Polynomials
2024
SIGMOD
7
3,073
DeepJoin: Joinable Table Discovery with Pre-trained Language Models
2023
VLDB
8
10,206
Bridging the Gap: Cardinality Estimation for Semantic Queries on Unstructured Data
2026
SIGMOD
9
7,290
Learning to be a Statistician: Learned Estimator for Number of Distinct Values
2022
VLDB
10
10,806
AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse Estimators
2025
VLDB