On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing
Summary: UnIMP: unified LLM-enhanced imputation for mixed-type (numeric/categorical/text) tables using a cell-oriented hypergraph and BiHMP — a bidirectional high-order message-passing network that captures inter-column heterogeneity and intra-column homogeneity. Xfusion adapters align BiHMP with LLMs and a pretrain+fine-tune pipeline with chunking and progressive masking yields theoretical guarantees and superior empirical results on 10 real-world datasets. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Jianwei Wang (University of New South Wales)
- 2. Kai Wang (Shanghai Jiao Tong University)
- 3. Ying Zhang (Zhejiang Gongshang University)
- 4. Wenjie Zhang (University of New South Wales)
- 5. Xiwei Xu (Commonwealth Scientific and Industrial Research Organisation)
- 6. Xuemin Lin (Shanghai Jiao Tong University)
BibTeX Citation
@article{wang_vldb25,
title = {{On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing}},
author = {Wang, Jianwei and Wang, Kai and Zhang, Ying and Zhang, Wenjie and Xu, Xiwei and Lin, Xuemin},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {10},
pages = {3421--3434},
doi = {10.14778/3748191.3748205},
url = {https://doi.org/10.14778/3748191.3748205},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,082 | Machine Learning for Graph Data Management and Query Processing | 2025 | VLDB | 5.2258409e-05 |
| 10,862 | Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models | 2026 | VLDB | 4.9769913e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 329 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00020867521 |
| 1,977 | Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks | 2024 | SIGMOD | 9.2760522e-05 |
| 6,000 | Neural Attributed Community Search at Billion Scale | 2023 | SIGMOD | 5.9170672e-05 |
| 6,780 | Missing Data Imputation with Uncertainty-Driven Network | 2024 | SIGMOD | 5.6826037e-05 |
| 8,999 | Efficient Unsupervised Community Search with Pre-trained Graph Transformer | 2024 | VLDB | 5.2396658e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,933 | TabulaX: Leveraging Large Language Models for Multi-Class Table Transformations | 2025 | VLDB |
| 2 | 10,878 | DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation | 2026 | VLDB |
| 3 | 8,008 | Missing Value Imputation for Multi-attribute Sensor Data Streams via Message Propagation | 2024 | VLDB |
| 4 | 6,780 | Missing Data Imputation with Uncertainty-Driven Network | 2024 | SIGMOD |
| 5 | 4,727 | LLM for Data Management | 2024 | VLDB |
| 6 | 5,686 | DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models | 2024 | SIGMOD |
| 7 | 3,441 | Efficient and Effective Data Imputation with Influence Functions | 2022 | VLDB |
| 8 | 11,535 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 9 | 10,862 | Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models | 2026 | VLDB |
| 10 | 7,589 | Data Imputation with Limited Data Redundancy Using Data Lakes | 2025 | VLDB |