Database Paper Browser

Back to papers

PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models

Summary: Leverages semantic schema via pre-trained language models to estimate the number of distinct values (NDV) with reduced data access. PLM4NDV fuses target-column and table semantics to lower access costs, can operate with no data access, and outperforms baselines on large real-world datasets. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
7256
Venue
SIGMOD
Year
2025
Pagerank
4.1905499e-05
Overall Rank
10,507 | 26.98%
DOI
10.1145/3725336

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 26 of 26 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1 Access Path Selection in a Relational Database Management System 1979 SIGMOD 0.0040465394
60 Sampling-Based Estimation of the Number of Distinct Values of an Attribute 1995 VLDB 0.00064450997
219 Deep Entity Matching with Pre-Trained Language Models 2021 VLDB 0.00033354456
366 Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation 2024 VLDB 0.00025580097
380 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00024943236
514 TURL: Table Understanding through Representation Learning 2021 VLDB 0.00021280726
531 Random Sampling for Histogram Construction: How much is enough? 1998 SIGMOD 0.0002079072
627 Preventing Bad Plans by Bounding the Impact of Cardinality Estimation Errors 2009 VLDB 0.00018959896
737 On Synopses for Distinct-Value Estimation Under Multiset Operations 2007 SIGMOD 0.00017377393
1,683 Cardinality Estimation: An Experimental Survey 2018 VLDB 0.0001091276
1,790 Effective Use of Block-Level Sampling in Statistics Estimation 2004 SIGMOD 0.00010529479
1,953 D-Bot: Database Diagnosis System using Large Language Models 2024 VLDB 9.9701097e-05
2,513 Annotating Columns with Pre-trained Language Models 2022 SIGMOD 8.6155767e-05
2,948 Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning 2023 SIGMOD 7.8301799e-05
3,520 GitTables: A Large-Scale Corpus of Relational Tables 2023 SIGMOD 7.0136102e-05
5,001 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 5.7634197e-05
5,343 Learned Index Benefits: Machine Learning Based Index Performance Estimation 2022 VLDB 5.5582234e-05
5,405 ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic Workloads 2024 VLDB 5.5243727e-05
7,332 Refactoring Index Tuning Process with Benefit Estimation 2024 VLDB 4.7553758e-05
7,611 Learning to be a Statistician: Learned Estimator for Number of Distinct Values 2022 VLDB 4.6920008e-05
7,708 UltraLogLog: A Practical and More Space-Efficient Alternative to HyperLogLog for Approximate Distinct Counting 2024 VLDB 4.6675856e-05
8,370 LAQy: Efficient and Reusable Query Approximations via Lazy Sampling 2023 SIGMOD 4.5287754e-05
8,679 FormaT5: Abstention and Examples for Conditional Table Formatting with Natural Language 2024 VLDB 4.4644049e-05
8,834 ByteCard: Enhancing ByteDance’s Data Warehouse with Learned Cardinality Estimation 2024 SIGMOD 4.4351469e-05
8,835 Learning-based Property Estimation with Polynomials 2024 SIGMOD 4.4351469e-05
10,543 AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse Estimators 2025 VLDB 4.1905499e-05
Previous Page 1 / 1 Next

Semantically Similar Papers