DBScholar

Back to papers

Can Foundation Models Wrangle Your Data?

Summary: Casts five data-cleaning and integration tasks as prompts, showing large foundation models achieve state-of-the-art performance without task-specific fine-tuning. Highlights privacy/domain adaptation challenges and accessibility opportunities for data management. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
13514
Venue
VLDB
Year
2023
Pagerank
0.00018789852
Overall Rank
420 | 97.13%
DOI
10.14778/3574245.3574258

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{narayan_vldb23,
        title = {{Can Foundation Models Wrangle Your Data?}},
        author = {Narayan, Avanika and Chami, Ines and Orr, Laurel and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {4},
        pages = {738--746},
        doi = {10.14778/3574245.3574258},
        url = {https://doi.org/10.14778/3574245.3574258},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 56 citing papers.

Rank Citing Paper Year Venue Pagerank
713 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00014672521
1,550 Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes 2023 CIDR 0.00010385904
2,099 Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks 2024 SIGMOD 9.1682353e-05
2,242 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 8.8823802e-05
2,298 GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian Optimization 2024 VLDB 8.7886538e-05
2,956 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 7.9203461e-05
3,073 DeepJoin: Joinable Table Discovery with Pre-trained Language Models 2023 VLDB 7.785842e-05
3,436 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.4157897e-05
3,536 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.3297343e-05
4,050 Revisiting Prompt Engineering via Declarative Crowdsourcing 2024 CIDR 6.9368666e-05
4,363 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 6.7423909e-05
4,515 ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models 2024 VLDB 6.6492389e-05
4,621 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.6024412e-05
4,944 Hybrid Querying Over Relational Databases and Large Language Models 2025 CIDR 6.432467e-05
5,293 Magneto: Combining Small and Large Language Models for Schema Matching 2025 VLDB 6.279939e-05
5,354 Can Large Language Models Predict Data Correlations from Column Names? 2023 VLDB 6.2515841e-05
5,466 SchemaPile: A Large Collection of Relational Database Schemas 2024 SIGMOD 6.2075052e-05
5,682 RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes 2024 VLDB 6.123486e-05
5,744 SQLStorm: Taking Database Benchmarking into the LLM Era 2025 VLDB 6.1019672e-05
5,816 Observatory: Characterizing Embeddings of Relational Tables 2024 VLDB 6.0776454e-05
6,094 The Fast and the Private: Task-based Dataset Search 2024 CIDR 5.9786215e-05
6,118 ELEET: Efficient Learned Query Execution over Text and Tables 2024 VLDB 5.9698795e-05
6,928 How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses 2024 VLDB 5.7370426e-05
7,075 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 5.7099047e-05
7,228 Mind the Data Gap: Bridging LLMs to Enterprise Data Integration 2025 CIDR 5.66667e-05
7,395 Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines 2023 SIGMOD 5.6257796e-05
7,865 Generating Succinct Descriptions of Database Schemata for Cost-Efficient Prompting of Large Language Models 2024 VLDB 5.5282752e-05
8,197 SMARTFEAT: Efficient Feature Construction through Feature-Level Foundation Model Interactions 2024 CIDR 5.4691464e-05
8,692 FormaT5: Abstention and Examples for Conditional Table Formatting with Natural Language 2024 VLDB 5.3838716e-05
8,829 Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees 2026 SIGMOD 5.3613784e-05
8,854 Towards Foundation Database Models 2025 CIDR 5.357608e-05
8,906 Unveiling Challenges for LLMs in Enterprise Data Engineering 2026 VLDB 5.3483178e-05
9,155 Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents 2025 SIGMOD 5.3122817e-05
9,380 Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence 2025 VLDB 5.2755515e-05
9,381 Deduplicated Sampling On-Demand 2025 VLDB 5.2755515e-05
9,437 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.2687567e-05
9,539 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.2528121e-05
9,550 Data Imputation with Limited Data Redundancy Using Data Lakes 2025 VLDB 5.2528121e-05
9,631 Lingua Manga: A Generic Large Language Model Centric System for Data Curation 2023 VLDB 5.2434488e-05
9,654 Automating the Enterprise with Foundation Models 2024 VLDB 5.2422003e-05
9,972 Cents: A Flexible and Cost-Effective Framework for LLM-Based Table Understanding 2025 VLDB 5.1845938e-05
10,186 Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees 2026 SIGMOD 5.093636e-05
10,318 In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration 2026 SIGMOD 5.093636e-05
10,724 LLM-Matcher: A Name-Based Schema Matching Tool using Large Language Models 2025 SIGMOD 5.093636e-05
10,746 A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces 2025 SIGMOD 5.093636e-05
10,785 Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables 2025 SIGMOD 5.093636e-05
10,856 Optimized Batch Prompting for Cost-effective LLMs 2025 VLDB 5.093636e-05
10,867 Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation 2025 VLDB 5.093636e-05
10,882 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 5.093636e-05
10,924 On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing 2025 VLDB 5.093636e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 15 of 15 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers