DBScholar

Back to papers

Can Large Language Models Predict Data Correlations from Column Names?

Summary: Introduces a Kaggle-derived benchmark for predicting column correlations from names. Evaluates language models across correlation and accuracy metrics, identifying name length, word ratio, and column types as key factors. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hf07232afa1dd5cfb
Venue
VLDB
Year
2023
Pagerank
6.1592249e-05
Overall Rank
5,362 | 63.97%
DOI
10.14778/3625054.3625066
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{trummer_vldb23,
        title = {{Can Large Language Models Predict Data Correlations from Column Names?}},
        author = {Trummer, Immanuel},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {13},
        pages = {4310--4323},
        doi = {10.14778/3625054.3625066},
        url = {https://doi.org/10.14778/3625054.3625066},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 26 of 26 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1 Access Path Selection in a Relational Database Management System 1979 SIGMOD 0.0023943337
15 How Good Are Query Optimizers, Really? 2016 VLDB 0.00061067652
144 Neo: A Learned Query Optimizer 2019 VLDB 0.00029090793
160 CORDS: Automatic Discovery of Correlations and Soft Functional Dependencies 2004 SIGMOD 0.00027827605
329 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00020867521
439 ATHENA: An Ontology-Driven System for Natural Language Querying over Relational Data Stores 2016 VLDB 0.00018253425
496 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017318538
536 NaLIR: An Interactive Natural Language Interface for Querying Relational Databases 2014 SIGMOD 0.00016774383
1,058 Lightweight Graphical Models for Selectivity Estimation Without Independence Assumptions 2011 VLDB 0.00012224038
1,155 QuickSel: Quick Selectivity Learning with Mixture Models 2020 SIGMOD 0.00011777046
1,245 DB-BERT: A Database Tuning Tool that "Reads the Manual" 2022 SIGMOD 0.0001136308
1,394 Data Profiling with Metanome 2015 VLDB 0.00010791559
1,605 SkinnerDB: Regret-Bounded Query Evaluation via Reinforcement Learning 2019 SIGMOD 0.00010095581
1,663 BHUNT: Automatic Discovery of Fuzzy Algebraic Constraints in Relational Data 2003 VLDB 9.942673e-05
1,840 From Natural Language Processing to Neural Databases 2021 VLDB 9.5271706e-05
1,929 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 9.3551286e-05
1,975 CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 Codex 2022 VLDB 9.2801545e-05
1,993 RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 2021 VLDB 9.2348951e-05
3,361 Conditional Selectivity for Statistics on Query Expressions 2004 SIGMOD 7.3726415e-05
3,648 UDO: Universal Database Optimization using Reinforcement Learning 2021 VLDB 7.1366536e-05
4,247 Divide & Conquer-based Inclusion Dependency Discovery 2015 VLDB 6.703008e-05
4,444 From BERT to GPT-3 Codex: Harnessing the Potential of Very Large Language Models for Data Management 2022 VLDB 6.5957098e-05
5,069 Scrutinizer: Fact Checking Statistical Claims 2020 VLDB 6.28697e-05
5,642 DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python 2021 SIGMOD 6.0519777e-05
7,063 Towards NLP-Enhanced Data Profiling Tools 2022 CIDR 5.6078438e-05
8,264 The Case for NLP-Enhanced Database Tuning: Towards Tuning Tools that “Read the Manual” 2021 VLDB 5.3648071e-05
Previous Page 1 / 1 Next

Semantically Similar Papers