DBScholar

Back to papers

Can Foundation Models Wrangle Your Data?

Summary: Casts five data-cleaning and integration tasks as prompts, showing large foundation models achieve state-of-the-art performance without task-specific fine-tuning. Highlights privacy/domain adaptation challenges and accessibility opportunities for data management. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h551ca782cf25abbd
Venue
VLDB
Year
2023
Pagerank
0.00020858443
Overall Rank
329 | 97.79%
DOI
10.14778/3574245.3574258

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{narayan_vldb23,
        title = {{Can Foundation Models Wrangle Your Data?}},
        author = {Narayan, Avanika and Chami, Ines and Orr, Laurel and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {4},
        pages = {738--746},
        doi = {10.14778/3574245.3574258},
        url = {https://doi.org/10.14778/3574245.3574258},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 64 citing papers.

Rank Citing Paper Year Venue Pagerank
501 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017267905
1,392 Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes 2023 CIDR 0.00010807936
1,932 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 9.34643e-05
1,978 Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks 2024 SIGMOD 9.2730152e-05
2,231 GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian Optimization 2024 VLDB 8.7982985e-05
2,341 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 8.6066183e-05
2,382 DeepJoin: Joinable Table Discovery with Pre-trained Language Models 2023 VLDB 8.5458532e-05
3,250 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.4938546e-05
3,336 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 7.4137763e-05
3,473 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.2728706e-05
3,695 Revisiting Prompt Engineering via Declarative Crowdsourcing 2024 CIDR 7.092445e-05
3,973 RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes 2024 VLDB 6.8876964e-05
4,132 SQLStorm: Taking Database Benchmarking into the LLM Era 2025 VLDB 6.7885553e-05
4,218 Magneto: Combining Small and Large Language Models for Schema Matching 2025 VLDB 6.7270169e-05
4,351 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 6.642803e-05
4,589 ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models 2024 VLDB 6.5144711e-05
4,675 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.474947e-05
4,787 Hybrid Querying Over Relational Databases and Large Language Models 2025 CIDR 6.4143303e-05
5,332 Observatory: Characterizing Embeddings of Relational Tables 2024 VLDB 6.1766186e-05
5,357 Can Large Language Models Predict Data Correlations from Column Names? 2023 VLDB 6.1619918e-05
5,385 SchemaPile: A Large Collection of Relational Database Schemas 2024 SIGMOD 6.1523255e-05
5,410 ELEET: Efficient Learned Query Execution over Text and Tables 2024 VLDB 6.1408676e-05
6,229 The Fast and the Private: Task-based Dataset Search 2024 CIDR 5.8444773e-05
7,068 How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses 2024 VLDB 5.6083188e-05
7,301 Mind the Data Gap: Bridging LLMs to Enterprise Data Integration 2025 CIDR 5.5602724e-05
7,385 Towards Foundation Database Models 2025 CIDR 5.5385053e-05
7,526 Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence 2025 VLDB 5.5020723e-05
7,538 Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines 2023 SIGMOD 5.4995874e-05
7,583 Data Imputation with Limited Data Redundancy Using Data Lakes 2025 VLDB 5.4906208e-05
7,781 LLM-Matcher: A Name-Based Schema Matching Tool using Large Language Models 2025 SIGMOD 5.4535721e-05
8,016 Generating Succinct Descriptions of Database Schemata for Cost-Efficient Prompting of Large Language Models 2024 VLDB 5.4065977e-05
8,230 Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents 2025 SIGMOD 5.372368e-05
8,322 SMARTFEAT: Efficient Feature Construction through Feature-Level Foundation Model Interactions 2024 CIDR 5.3545825e-05
8,420 On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing 2025 VLDB 5.3350162e-05
8,860 FormaT5: Abstention and Examples for Conditional Table Formatting with Natural Language 2024 VLDB 5.263072e-05
8,888 Optimized Batch Prompting for Cost-effective LLMs 2025 VLDB 5.2559789e-05
8,997 Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees 2026 SIGMOD 5.2410834e-05
9,069 Unveiling Challenges for LLMs in Enterprise Data Engineering 2026 VLDB 5.2283159e-05
9,229 Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables 2025 SIGMOD 5.2056825e-05
9,565 Deduplicated Sampling On-Demand 2025 VLDB 5.1571823e-05
9,616 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.1510548e-05
9,722 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.1349531e-05
9,808 Lingua Manga: A Generic Large Language Model Centric System for Data Curation 2023 VLDB 5.1257999e-05
9,831 Automating the Enterprise with Foundation Models 2024 VLDB 5.1245795e-05
10,165 Cents: A Flexible and Cost-Effective Framework for LLM-Based Table Understanding 2025 VLDB 5.0682654e-05
10,210 Blocker and Matcher Can Mutually Benefit: A Co-Learning Framework for Low-Resource Entity Resolution 2024 VLDB 5.0596605e-05
10,402 Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees 2026 SIGMOD 4.9793485e-05
10,527 In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration 2026 SIGMOD 4.9793485e-05
10,757 Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs 2026 VLDB 4.9793485e-05
10,791 BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents 2026 VLDB 4.9793485e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 15 of 15 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers