Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
Summary: Auto-Fill post-trains three small language-model specialists for world knowledge, text reasoning, and code reasoning in missing-cell prediction. A calibrated ensemble selects or abstains, outperforming frontier models on 2,200 tables at under 1% of their cost. (summarized by gpt-5.6-luna on Aug 28 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Yurong Liu (New York University)
- 2. Yeye He (Microsoft)
- 3. Haoyu Dong (Microsoft)
- 4. Junjie Xing (Microsoft)
- 5. Shi Han (Microsoft)
- 6. Dongmei Zhang (Microsoft)
- 7. Surajit Chaudhuri (Microsoft)
BibTeX Citation
@article{liu_vldb26,
title = {{Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models}},
author = {Liu, Yurong and He, Yeye and Dong, Haoyu and Xing, Junjie and Han, Shi and Zhang, Dongmei and Chaudhuri, Surajit},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {11},
pages = {3160--3173},
doi = {10.14778/3836663.3836680},
url = {https://doi.org/10.14778/3836663.3836680},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,043 | Data Cleaning: Overview and Emerging Challenges | 2016 | SIGMOD | 0.00012335114 |
| 1,342 | Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning | 2020 | VLDB | 0.00010968223 |
| 1,978 | Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks | 2024 | SIGMOD | 9.2730152e-05 |
| 3,263 | Automatic Data Repair: Are We Ready to Deploy? | 2024 | VLDB | 7.4809771e-05 |
| 7,583 | Data Imputation with Limited Data Redundancy Using Data Lakes | 2025 | VLDB | 5.4906208e-05 |
| 7,738 | Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes | 2021 | SIGMOD | 5.4623434e-05 |
| 8,420 | On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing | 2025 | VLDB | 5.3350162e-05 |
| 8,979 | Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph | 2023 | VLDB | 5.2443558e-05 |
| 9,229 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD | 5.2056825e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,978 | Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks | 2024 | SIGMOD |
| 2 | 2,960 | Auto-Detect: Data-Driven Error Detection in Tables | 2018 | SIGMOD |
| 3 | 1,780 | Annotating Columns with Pre-trained Language Models | 2022 | SIGMOD |
| 4 | 10,869 | DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation | 2026 | VLDB |
| 5 | 7,270 | AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework | 2025 | VLDB |
| 6 | 11,514 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |
| 7 | 8,420 | On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing | 2025 | VLDB |
| 8 | 5,687 | DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models | 2024 | SIGMOD |
| 9 | 9,229 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD |
| 10 | 7,583 | Data Imputation with Limited Data Redundancy Using Data Lakes | 2025 | VLDB |