DBScholar

Back to papers

How Large Language Models Will Disrupt Data Management

Summary: LLMs provide semantic grounding of tuples, schemas, and queries, enabling automation breakthroughs in tasks that stalled (entity resolution, schema matching, data discovery, query synthesis). They also blur predictive models and IR, prompting new DB/architecture designs. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
he832304e7c395036
Venue
VLDB
Year
2023
Pagerank
7.4996147e-05
Overall Rank
3,236 | 78.26%
DOI
10.14778/3611479.3611527
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{fernandez_vldb23,
        title = {{How Large Language Models Will Disrupt Data Management}},
        author = {Fernandez, Raul Castro and Elmore, Aaron J. and Franklin, Michael J. and Krishnan, Sanjay and Tan, Chenhao},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {11},
        pages = {3302--3309},
        doi = {10.14778/3611479.3611527},
        url = {https://doi.org/10.14778/3611479.3611527},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 13 of 13 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 22 of 22 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
17 Provenance Semirings 2007 PODS 0.00059813669
134 Deep Entity Matching with Pre-Trained Language Models 2021 VLDB 0.00030042569
329 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00020867521
377 TURL: Table Understanding through Representation Learning 2021 VLDB 0.00019564011
484 Data Integration for the Relational Web 2009 VLDB 0.00017541069
536 NaLIR: An Interactive Natural Language Interface for Querying Relational Databases 2014 SIGMOD 0.00016774383
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016083582
824 Data Integration: The Teenage Years 2006 VLDB 0.00013641217
992 Web-scale Data Integration: You can only afford to Pay As You Go 2007 CIDR 0.00012641942
1,245 DB-BERT: A Database Tuning Tool that "Reads the Manual" 2022 SIGMOD 0.0001136308
1,692 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis 2018 SIGMOD 9.8526476e-05
1,975 CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 Codex 2022 VLDB 9.2801545e-05
2,093 Sato: Contextual Semantic Type Detection in Tables 2020 VLDB 9.0593586e-05
2,247 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel 2023 VLDB 8.7597965e-05
3,420 Ava: From Data to Insights Through Conversation 2017 CIDR 7.3130183e-05
3,450 MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud 2023 VLDB 7.2889709e-05
3,617 Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation 2022 SIGMOD 7.1540505e-05
3,812 FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline 2023 VLDB 7.0056415e-05
4,674 Knowledge Graphs 2021: A Data Odyssey 2021 VLDB 6.4731383e-05
5,011 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 6.3131083e-05
6,979 Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach 2023 SIGMOD 5.6272426e-05
8,264 The Case for NLP-Enhanced Database Tuning: Towards Tuning Tools that “Read the Manual” 2021 VLDB 5.3648071e-05
Previous Page 1 / 1 Next

Semantically Similar Papers