DBScholar

Back to papers

Can Large Language Models Predict Data Correlations from Column Names?

Summary: Introduces a Kaggle-derived benchmark for predicting column correlations from names. Evaluates language models across correlation and accuracy metrics, identifying name length, word ratio, and column types as key factors. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
13486
Venue
VLDB
Year
2023
Pagerank
6.2515841e-05
Overall Rank
5,354 | 63.27%
DOI
10.14778/3625054.3625066

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{trummer_vldb23,
        title = {{Can Large Language Models Predict Data Correlations from Column Names?}},
        author = {Trummer, Immanuel},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {13},
        pages = {4310--4323},
        doi = {10.14778/3625054.3625066},
        url = {https://doi.org/10.14778/3625054.3625066},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 26 of 26 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1 Access Path Selection in a Relational Database Management System 1979 SIGMOD 0.0024089429
18 How Good Are Query Optimizers, Really? 2016 VLDB 0.00059284255
154 Neo: A Learned Query Optimizer 2019 VLDB 0.00028726181
159 CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies 2004 SIGMOD 0.00028129426
420 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00018789852
459 ATHENA: An Ontology-Driven System for Natural Language Querying over Relational Data Stores 2016 VLDB 0.00018101318
555 NaLIR: An Interactive Natural Language Interface for Querying Relational Databases 2014 SIGMOD 0.00016568053
713 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00014672521
1,071 Lightweight Graphical Models for Selectivity Estimation Without Independence Assumptions 2011 VLDB 0.00012322342
1,170 QuickSel: Quick Selectivity Learning with Mixture Models 2020 SIGMOD 0.00011827259
1,337 DB-BERT: A Database Tuning Tool that "Reads the Manual" 2022 SIGMOD 0.00011117488
1,374 Data Profiling with Metanome 2015 VLDB 0.00010986078
1,648 BHUNT: Automatic Discovery of Fuzzy Algebraic Constraints in Relational Data 2003 VLDB 0.00010120668
1,712 SkinnerDB: Regret-Bounded Query Evaluation via Reinforcement Learning 2019 SIGMOD 9.9492299e-05
2,019 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2983994e-05
2,175 From Natural Language Processing to Neural Databases 2021 VLDB 9.0215773e-05
2,242 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 8.8823802e-05
2,521 CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 Codex 2022 VLDB 8.4729505e-05
3,309 Conditional Selectivity for Statistics on Query Expressions 2004 SIGMOD 7.5368417e-05
3,926 UDO: Universal Database Optimization using Reinforcement Learning 2021 VLDB 7.0128068e-05
4,386 Divide & Conquer-based Inclusion Dependency Discovery 2015 VLDB 6.7309577e-05
4,449 From BERT to GPT-3 Codex: Harnessing the Potential of Very Large Language Models for Data Management 2022 VLDB 6.697553e-05
4,963 Scrutinizer: Fact Checking Statistical Claims 2020 VLDB 6.4260475e-05
5,620 DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python 2021 SIGMOD 6.1482864e-05
7,038 Towards NLP-Enhanced Data Profiling Tools 2022 CIDR 5.7204576e-05
8,127 The Case for NLP-Enhanced Database Tuning: Towards Tuning Tools that “Read the Manual” 2021 VLDB 5.4826853e-05
Previous Page 1 / 1 Next

Semantically Similar Papers