On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing
Summary: UnIMP: unified LLM-enhanced imputation for mixed-type (numeric/categorical/text) tables using a cell-oriented hypergraph and BiHMP — a bidirectional high-order message-passing network that captures inter-column heterogeneity and intra-column homogeneity. Xfusion adapters align BiHMP with LLMs and a pretrain+fine-tune pipeline with chunking and progressive masking yields theoretical guarantees and superior empirical results on 10 real-world datasets. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Jianwei Wang (University of New South Wales)
- 2. Kai Wang (Shanghai Jiao Tong University)
- 3. Ying Zhang (Zhejiang Gongshang University)
- 4. Wenjie Zhang (University of New South Wales)
- 5. Xiwei Xu (Commonwealth Scientific and Industrial Research Organisation)
- 6. Xuemin Lin (Shanghai Jiao Tong University)
BibTeX Citation
@article{wang_vldb25,
title = {{On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing}},
author = {Wang, Jianwei and Wang, Kai and Zhang, Ying and Zhang, Wenjie and Xu, Xiwei and Lin, Xuemin},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {10},
pages = {3421--3434},
doi = {10.14778/3748191.3748205},
url = {https://doi.org/10.14778/3748191.3748205},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,073 | Machine Learning for Graph Data Management and Query Processing | 2025 | VLDB | 5.2283159e-05 |
| 10,853 | Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models | 2026 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 329 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00020858443 |
| 1,978 | Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks | 2024 | SIGMOD | 9.2730152e-05 |
| 5,998 | Neural Attributed Community Search at Billion Scale | 2023 | SIGMOD | 5.9198696e-05 |
| 6,775 | Missing Data Imputation with Uncertainty-Driven Network | 2024 | SIGMOD | 5.685295e-05 |
| 8,990 | Efficient Unsupervised Community Search with Pre-trained Graph Transformer | 2024 | VLDB | 5.2421474e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,928 | TabulaX: Leveraging Large Language Models for Multi-Class Table Transformations | 2025 | VLDB |
| 2 | 10,869 | DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation | 2026 | VLDB |
| 3 | 8,003 | Missing Value Imputation for Multi-attribute Sensor Data Streams via Message Propagation | 2024 | VLDB |
| 4 | 5,149 | LLM for Data Management | 2024 | VLDB |
| 5 | 6,775 | Missing Data Imputation with Uncertainty-Driven Network | 2024 | SIGMOD |
| 6 | 5,687 | DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models | 2024 | SIGMOD |
| 7 | 3,441 | Efficient and Effective Data Imputation with Influence Functions | 2022 | VLDB |
| 8 | 11,529 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 9 | 10,853 | Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models | 2026 | VLDB |
| 10 | 7,583 | Data Imputation with Limited Data Redundancy Using Data Lakes | 2025 | VLDB |