DBScholar

Back to papers

Pre-training Summarization Models of Structured Datasets for Cardinality Estimation

Summary: Pre-trained models convert structured datasets into compact summaries for cardinality estimation, eliminating per-dataset training and speeding up summary construction up to 100x. Uses multiple summaries per dataset and learned summaries for hard columnsets; shows shared frequency and correlation patterns learned from a diverse corpus with incremental updates. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h865fbf8594335c4f
Venue
VLDB
Year
2022
Pagerank
5.9737602e-05
Overall Rank
5,835 | 60.79%
DOI
10.14778/3494124.3494127
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{lu_vldb22,
        title = {{Pre-training Summarization Models of Structured Datasets for Cardinality Estimation}},
        author = {Lu, Yao and Kandula, Srikanth and König, Arnd Christian and Chaudhuri, Surajit},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {3},
        pages = {414--426},
        doi = {10.14778/3494124.3494127},
        url = {https://doi.org/10.14778/3494124.3494127},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 10 of 10 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 21 of 21 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
15 How Good Are Query Optimizers, Really? 2016 VLDB 0.00061067652
40 The Case for Learned Index Structures 2018 SIGMOD 0.00046363107
85 Learned Cardinalities: Estimating Correlated Joins with Deep Learning 2019 CIDR 0.00035876108
124 Approximate Frequency Counts over Data Streams 2002 VLDB 0.00030586757
144 Neo: A Learned Query Optimizer 2019 VLDB 0.00029090793
160 CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies 2004 SIGMOD 0.00027827605
169 Wavelet-Based Histograms for Selectivity Estimation 1998 SIGMOD 0.00027126333
255 The History of Histograms (abridged) 2003 VLDB 0.00022974524
286 Selectivity Estimation using Probabilistic Models 2001 SIGMOD 0.00022112534
295 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00021908194
309 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021376597
318 DeepDB: Learn from Data, not from Queries! 2020 VLDB 0.00021166957
371 STHoles: A Multidimensional Workload-Aware Histogram 2001 SIGMOD 0.00019822444
406 Deep Unsupervised Cardinality Estimation 2020 VLDB 0.00019050182
510 NeuroCard: One Cardinality Estimator for All Tables 2021 VLDB 0.00017059914
691 Selectivity Estimation for Range Predicates using Lightweight Models 2019 VLDB 0.00014737455
700 Independence is Good: Dependency-Based Histogram Synopses for High-Dimensional Data 2001 SIGMOD 0.00014675903
707 On Synopses for Distinct-Value Estimation Under Multiset Operations 2007 SIGMOD 0.00014633741
1,058 Lightweight Graphical Models for Selectivity Estimation Without Independence Assumptions 2011 VLDB 0.00012224038
1,173 Wavelet Synopses with Error Guarantees 2002 SIGMOD 0.0001168187
1,509 Self-Tuning, GPU-Accelerated Kernel Density Models for Multidimensional Selectivity Estimation 2015 SIGMOD 0.00010436933
Previous Page 1 / 1 Next

Semantically Similar Papers