Back to papers
Data Imputation with Limited Data Redundancy Using Data Lakes
Summary: LakeFill leverages LLMs and data lakes for tuple-level retrieval and encoding of incomplete tuples to find cross-table candidates when intra-table redundancy is low. It uses checklist-based reranking and a two-stage confidence-aware reasoner, beating prior methods.
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 13967
- Venue
- VLDB
- Year
- 2025
- Pagerank
- 4.3300131e-05
- Overall Rank
- 9,481 | 34.11%
- DOI
-
10.14778/3748191.3748200
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 192 |
HoloClean: Holistic Data Repairs with Probabilistic Inference |
2017 |
VLDB |
0.00035692958 |
| 219 |
Deep Entity Matching with Pre-Trained Language Models |
2021 |
VLDB |
0.00033354456 |
| 514 |
TURL: Table Understanding through Representation Learning |
2021 |
VLDB |
0.00021280726 |
| 516 |
Can Foundation Models Wrangle Your Data? |
2023 |
VLDB |
0.00021194444 |
| 1,160 |
Towards Certain Fixes with Editing Rules and Master Data |
2010 |
VLDB |
0.0001358129 |
| 1,403 |
Detecting Data Errors: Where are we and what needs to be done? |
2016 |
VLDB |
0.00012180046 |
| 1,544 |
KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing |
2015 |
SIGMOD |
0.00011438274 |
| 1,895 |
Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning |
2020 |
VLDB |
0.00010174634 |
| 2,348 |
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation |
2021 |
VLDB |
8.9903659e-05 |
| 2,585 |
Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks |
2024 |
SIGMOD |
8.4909917e-05 |
| 3,666 |
The Dawn of Natural Language to SQL: Are We Fully Ready? |
2024 |
VLDB |
6.8606092e-05 |
| 3,973 |
HAIChart: Human and AI Paired Visualization System |
2024 |
VLDB |
6.5721521e-05 |
| 8,263 |
Learned Data-aware Image Representations of Line Charts for Similarity Search |
2023 |
SIGMOD |
4.5414364e-05 |
| 9,073 |
VerifAI: Verified Generative AI |
2024 |
CIDR |
4.396857e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 6,798 |
DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models |
2024 |
SIGMOD |
4.9186164e-05 |
| 1,088 |
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes |
2024 |
VLDB |
0.00014158762 |
| 5,470 |
RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes |
2024 |
VLDB |
5.4894925e-05 |
| 10,817 |
TARImpute: Task-Aware auto-Recommender System for Missing Value Imputation Algorithms with Clustering Case Studies |
2025 |
VLDB |
4.1905499e-05 |
| 10,956 |
Certain and Approximately Certain Models for Statistical Learning |
2024 |
SIGMOD |
4.1905499e-05 |
| 3,313 |
Efficient and Effective Data Imputation with Influence Functions |
2022 |
VLDB |
7.2336734e-05 |
| 2,575 |
Query Optimization for Dynamic Imputation |
2017 |
VLDB |
8.5100213e-05 |
| 5,262 |
Enriching Data Imputation with Extensive Similarity Neighbors |
2015 |
VLDB |
5.5961016e-05 |
| 9,855 |
In-Database Data Imputation |
2024 |
SIGMOD |
4.2652623e-05 |
| 10,683 |
On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing |
2025 |
VLDB |
4.1905499e-05 |