Data Imputation with Limited Data Redundancy Using Data Lakes
Summary: LakeFill leverages LLMs and data lakes for tuple-level retrieval and encoding of incomplete tuples to find cross-table candidates when intra-table redundancy is low. It uses checklist-based reranking and a two-stage confidence-aware reasoner, beating prior methods. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Chenyu Yang (Hong Kong University of Science and Technology)
- 2. Yuyu Luo (Hong Kong University of Science and Technology)
- 3. Chuanxuan Cui (Renmin University of China)
- 4. Ju Fan (Renmin University of China)
- 5. Chengliang Chai (Beijing Institute of Technology)
- 6. Nan Tang (Hong Kong University of Science and Technology)
BibTeX Citation
@article{yang_vldb25,
title = {{Data Imputation with Limited Data Redundancy Using Data Lakes}},
author = {Yang, Chenyu and Luo, Yuyu and Cui, Chuanxuan and Fan, Ju and Chai, Chengliang and Tang, Nan},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {10},
pages = {3354--3367},
doi = {10.14778/3748191.3748200},
url = {https://doi.org/10.14778/3748191.3748200},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,090 | LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning | 2026 | VLDB | 5.601767e-05 |
| 10,352 | Towards Scalable Visual Data Wrangling via Direct Manipulation | 2026 | CIDR | 4.9793485e-05 |
| 10,853 | Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models | 2026 | VLDB | 4.9793485e-05 |
| 11,034 | Data Agents: Rethinking Data Systems in the AI Agent Era | 2026 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 501 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB |
| 2 | 3,973 | RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes | 2024 | VLDB |
| 3 | 11,403 | TARImpute: Task-Aware auto-Recommender System for Missing Value Imputation Algorithms with Clustering Case Studies | 2025 | VLDB |
| 4 | 11,514 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |
| 5 | 3,441 | Efficient and Effective Data Imputation with Influence Functions | 2022 | VLDB |
| 6 | 2,239 | Query Optimization for Dynamic Imputation | 2017 | VLDB |
| 7 | 5,278 | Enriching Data Imputation with Extensive Similarity Neighbors | 2015 | VLDB |
| 8 | 10,181 | In-Database Data Imputation | 2024 | SIGMOD |
| 9 | 10,853 | Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models | 2026 | VLDB |
| 10 | 8,420 | On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing | 2025 | VLDB |