DBScholar

Back to papers

Pre-training Summarization Models of Structured Datasets for Cardinality Estimation

Summary: Pre-trained models convert structured datasets into compact summaries for cardinality estimation, eliminating per-dataset training and speeding up summary construction up to 100x. Uses multiple summaries per dataset and learned summaries for hard columnsets; shows shared frequency and correlation patterns learned from a diverse corpus with incremental updates. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
13106
Venue
VLDB
Year
2022
Pagerank
6.0871213e-05
Overall Rank
5,792 | 60.27%
DOI
10.14778/3494124.3494127

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{lu_vldb22,
        title = {{Pre-training Summarization Models of Structured Datasets for Cardinality Estimation}},
        author = {Lu, Yao and Kandula, Srikanth and König, Arnd Christian and Chaudhuri, Surajit},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {3},
        pages = {414--426},
        doi = {10.14778/3494124.3494127},
        url = {https://doi.org/10.14778/3494124.3494127},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 10 of 10 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 21 of 21 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
18 How Good Are Query Optimizers, Really? 2016 VLDB 0.00059284255
43 The Case for Learned Index Structures 2018 SIGMOD 0.00046060254
84 Learned Cardinalities: Estimating Correlated Joins with Deep Learning 2019 CIDR 0.00035838391
122 Approximate Frequency Counts over Data Streams 2002 VLDB 0.00031260115
154 Neo: A Learned Query Optimizer 2019 VLDB 0.00028726181
159 CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies 2004 SIGMOD 0.00028129426
168 Wavelet-Based Histograms for Selectivity Estimation 1998 SIGMOD 0.00027541029
257 The History of Histograms (abridged) 2003 VLDB 0.00023154793
280 Selectivity Estimation using Probabilistic Models 2001 SIGMOD 0.00022454217
288 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00022296371
307 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021792475
323 DeepDB: Learn from Data, not from Queries! 2020 VLDB 0.00021264788
365 STHoles: A Multidimensional Workload-Aware Histogram 2001 SIGMOD 0.00020041735
401 Deep Unsupervised Cardinality Estimation 2020 VLDB 0.00019092557
513 NeuroCard: One Cardinality Estimator for All Tables 2021 VLDB 0.00017190574
689 On Synopses for Distinct-Value Estimation Under Multiset Operations 2007 SIGMOD 0.00014940023
692 Independence is Good: Dependency-Based Histogram Synopses for High-Dimensional Data 2001 SIGMOD 0.00014919816
697 Selectivity Estimation for Range Predicates using Lightweight Models 2019 VLDB 0.00014888851
1,071 Lightweight Graphical Models for Selectivity Estimation Without Independence Assumptions 2011 VLDB 0.00012322342
1,156 Wavelet Synopses with Error Guarantees 2002 SIGMOD 0.00011929041
1,503 Self-Tuning, GPU-Accelerated Kernel Density Models for Multidimensional Selectivity Estimation 2015 SIGMOD 0.000105564
Previous Page 1 / 1 Next

Semantically Similar Papers