On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing
Summary: UnIMP: unified LLM-enhanced imputation for mixed-type (numeric/categorical/text) tables using a cell-oriented hypergraph and BiHMP — a bidirectional high-order message-passing network that captures inter-column heterogeneity and intra-column homogeneity. Xfusion adapters align BiHMP with LLMs and a pretrain+fine-tune pipeline with chunking and progressive masking yields theoretical guarantees and superior empirical results on 10 real-world datasets. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Jianwei Wang (University of New South Wales)
- 2. Kai Wang (Shanghai Jiao Tong University)
- 3. Ying Zhang (Zhejiang Gongshang University)
- 4. Wenjie Zhang (University of New South Wales)
- 5. Xiwei Xu (Commonwealth Scientific and Industrial Research Organisation)
- 6. Xuemin Lin (Shanghai Jiao Tong University)
BibTeX Citation
@article{wang_vldb25,
title = {{On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message Passing}},
author = {Wang, Jianwei and Wang, Kai and Zhang, Ying and Zhang, Wenjie and Xu, Xiwei and Lin, Xuemin},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {10},
pages = {3421--3434},
doi = {10.14778/3748191.3748205},
url = {https://doi.org/10.14778/3748191.3748205},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,067 | Machine Learning for Graph Data Management and Query Processing | 2025 | VLDB | 5.093636e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 420 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00018789852 |
| 2,099 | Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks | 2024 | SIGMOD | 9.1682353e-05 |
| 5,877 | Neural Attributed Community Search at Billion Scale | 2023 | SIGMOD | 6.0551011e-05 |
| 6,642 | Missing Data Imputation with Uncertainty-Driven Network | 2024 | SIGMOD | 5.8157857e-05 |
| 8,821 | Efficient Unsupervised Community Search with Pre-trained Graph Transformer | 2024 | VLDB | 5.3624668e-05 |
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,169 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |
| 2 | 10,614 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning | 2026 | VLDB |
| 3 | 9,505 | TabulaX: Leveraging Large Language Models for Multi-Class Table Transformations | 2025 | VLDB |
| 4 | 7,844 | Missing Value Imputation for Multi-attribute Sensor Data Streams via Message Propagation | 2024 | VLDB |
| 5 | 6,101 | LLM for Data Management | 2024 | VLDB |
| 6 | 6,642 | Missing Data Imputation with Uncertainty-Driven Network | 2024 | SIGMOD |
| 7 | 6,545 | DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language Models | 2024 | SIGMOD |
| 8 | 3,386 | Efficient and Effective Data Imputation with Influence Functions | 2022 | VLDB |
| 9 | 11,186 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 10 | 9,550 | Data Imputation with Limited Data Redundancy Using Data Lakes | 2025 | VLDB |