Data Imputation with Limited Data Redundancy Using Data Lakes
Summary: LakeFill leverages LLMs and data lakes for tuple-level retrieval and encoding of incomplete tuples to find cross-table candidates when intra-table redundancy is low. It uses checklist-based reranking and a two-stage confidence-aware reasoner, beating prior methods. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Chenyu Yang (Hong Kong University of Science and Technology)
- 2. Yuyu Luo (Hong Kong University of Science and Technology)
- 3. Chuanxuan Cui (Renmin University of China)
- 4. Ju Fan (Renmin University of China)
- 5. Chengliang Chai (Beijing Institute of Technology)
- 6. Nan Tang (Hong Kong University of Science and Technology)
BibTeX Citation
@article{yang_vldb25,
title = {{Data Imputation with Limited Data Redundancy Using Data Lakes}},
author = {Yang, Chenyu and Luo, Yuyu and Cui, Chuanxuan and Fan, Ju and Chai, Chengliang and Tang, Nan},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {10},
pages = {3354--3367},
doi = {10.14778/3748191.3748200},
url = {https://doi.org/10.14778/3748191.3748200},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,131 | Towards Scalable Visual Data Wrangling via Direct Manipulation | 2026 | CIDR | 5.093636e-05 |
| 10,587 | LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning | 2026 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,545 | DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models | 2024 | SIGMOD |
| 2 | 713 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB |
| 3 | 5,682 | RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes | 2024 | VLDB |
| 4 | 11,039 | TARImpute: Task-Aware auto-Recommender System for Missing Value Imputation Algorithms with Clustering Case Studies | 2025 | VLDB |
| 5 | 11,169 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |
| 6 | 3,386 | Efficient and Effective Data Imputation with Influence Functions | 2022 | VLDB |
| 7 | 5,167 | Enriching Data Imputation with Extensive Similarity Neighbors | 2015 | VLDB |
| 8 | 2,208 | Query Optimization for Dynamic Imputation | 2017 | VLDB |
| 9 | 9,993 | In-Database Data Imputation | 2024 | SIGMOD |
| 10 | 10,924 | On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing | 2025 | VLDB |