DBScholar

Back to papers

Can Foundation Models Wrangle Your Data?

Summary: Casts five data-cleaning and integration tasks as prompts, showing large foundation models achieve state-of-the-art performance without task-specific fine-tuning. Highlights privacy/domain adaptation challenges and accessibility opportunities for data management. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h551ca782cf25abbd
Venue
VLDB
Year
2023
Pagerank
0.00020867521
Overall Rank
329 | 97.80%
DOI
10.14778/3574245.3574258
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{narayan_vldb23,
        title = {{Can Foundation Models Wrangle Your Data?}},
        author = {Narayan, Avanika and Chami, Ines and Orr, Laurel and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {4},
        pages = {738--746},
        doi = {10.14778/3574245.3574258},
        url = {https://doi.org/10.14778/3574245.3574258},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 64 citing papers.

Rank Citing Paper Year Venue Pagerank
496 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017318538
1,387 Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes 2023 CIDR 0.00010829741
1,929 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 9.3551286e-05
1,977 Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks 2024 SIGMOD 9.2760522e-05
2,227 GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian Optimization 2024 VLDB 8.8007923e-05
2,273 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 8.7093949e-05
2,378 DeepJoin: Joinable Table Discovery with Pre-trained Language Models 2023 VLDB 8.5495442e-05
3,236 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.4996147e-05
3,321 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 7.427185e-05
3,473 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.2697311e-05
3,668 Revisiting Prompt Engineering via Declarative Crowdsourcing 2024 CIDR 7.1144218e-05
3,900 SQLStorm: Taking Database Benchmarking into the LLM Era 2025 VLDB 6.9320915e-05
3,967 RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes 2024 VLDB 6.887577e-05
4,213 Magneto: Combining Small and Large Language Models for Schema Matching 2025 VLDB 6.7281505e-05
4,352 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 6.6396584e-05
4,591 ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models 2024 VLDB 6.5113872e-05
4,678 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.4720352e-05
4,789 Hybrid Querying Over Relational Databases and Large Language Models 2025 CIDR 6.4117376e-05
5,208 ELEET: Efficient Learned Query Execution over Text and Tables 2024 VLDB 6.2254341e-05
5,322 Observatory: Characterizing Embeddings of Relational Tables 2024 VLDB 6.1809951e-05
5,362 Can Large Language Models Predict Data Correlations from Column Names? 2023 VLDB 6.1592249e-05
5,393 SchemaPile: A Large Collection of Relational Database Schemas 2024 SIGMOD 6.1494131e-05
6,232 The Fast and the Private: Task-based Dataset Search 2024 CIDR 5.8417106e-05
7,070 How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses 2024 VLDB 5.6056639e-05
7,304 Mind the Data Gap: Bridging LLMs to Enterprise Data Integration 2025 CIDR 5.5576403e-05
7,386 Towards Foundation Database Models 2025 CIDR 5.5360368e-05
7,531 Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence 2025 VLDB 5.4994676e-05
7,544 Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines 2023 SIGMOD 5.496984e-05
7,589 Data Imputation with Limited Data Redundancy Using Data Lakes 2025 VLDB 5.4880216e-05
7,788 LLM-Matcher: A Name-Based Schema Matching Tool using Large Language Models 2025 SIGMOD 5.4509905e-05
8,021 Generating Succinct Descriptions of Database Schemata for Cost-Efficient Prompting of Large Language Models 2024 VLDB 5.4041713e-05
8,236 Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents 2025 SIGMOD 5.3698248e-05
8,329 SMARTFEAT: Efficient Feature Construction through Feature-Level Foundation Model Interactions 2024 CIDR 5.3520477e-05
8,429 On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing 2025 VLDB 5.3324907e-05
8,526 Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees 2026 SIGMOD 5.3215522e-05
8,869 FormaT5: Abstention and Examples for Conditional Table Formatting with Natural Language 2024 VLDB 5.2605805e-05
8,897 Optimized Batch Prompting for Cost-effective LLMs 2025 VLDB 5.2534908e-05
9,078 Unveiling Challenges for LLMs in Enterprise Data Engineering 2026 VLDB 5.2258409e-05
9,239 Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables 2025 SIGMOD 5.2032182e-05
9,573 Deduplicated Sampling On-Demand 2025 VLDB 5.154741e-05
9,623 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.1486163e-05
9,727 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.1325223e-05
9,815 Lingua Manga: A Generic Large Language Model Centric System for Data Curation 2023 VLDB 5.1233734e-05
9,838 Automating the Enterprise with Foundation Models 2024 VLDB 5.1221535e-05
10,169 Cents: A Flexible and Cost-Effective Framework for LLM-Based Table Understanding 2025 VLDB 5.0658661e-05
10,199 Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees 2026 VLDB 5.0599411e-05
10,209 Blocker and Matcher Can Mutually Benefit: A Co-Learning Framework for Low-Resource Entity Resolution 2024 VLDB 5.0586032e-05
10,414 Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees 2026 SIGMOD 4.9769913e-05
10,538 In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration 2026 SIGMOD 4.9769913e-05
10,767 Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs 2026 VLDB 4.9769913e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 15 of 15 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers