Win-Win: On Simultaneous Clustering and Imputing over Incomplete Data
Summary: Joint clustering-and-imputation formulation for incomplete data (NP-hard), showing simultaneous optimization yields mutually reinforcing gains versus impute-then-cluster. Exact ILP and practical LP-relaxation + local-neighbor (LN) approximations with guarantees; empirical wins on real datasets. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Yu Sun (Nankai University)
- 2. Jingyu Zhu (Nankai University)
- 3. Xiao Xu (Ant Financial)
- 4. Xian Xu (Ant Financial)
- 5. Yuyao Sun (Ant Financial)
- 6. Shaoxu Song (Tsinghua University)
- 7. Xiang Li (Ping An Health Technology)
- 8. Xiaojie Yuan (Nankai University)
BibTeX Citation
@article{sun_vldb24,
title = {{Win-Win: On Simultaneous Clustering and Imputing over Incomplete Data}},
author = {Sun, Yu and Zhu, Jingyu and Xu, Xiao and Xu, Xian and Sun, Yuyao and Song, Shaoxu and Li, Xiang and Yuan, Xiaojie},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {11},
pages = {3045--3057},
doi = {10.14778/3681954.3681982},
url = {https://doi.org/10.14778/3681954.3681982},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 112 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00032801121 |
| 549 | ERACER: A Database Approach for Statistical Inference and Data Cleaning | 2010 | SIGMOD | 0.00016692839 |
| 661 | Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes | 2013 | SIGMOD | 0.0001519162 |
| 2,147 | Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions | 2021 | VLDB | 9.0831495e-05 |
| 2,369 | Data Generation using Declarative Constraints | 2011 | SIGMOD | 8.682429e-05 |
| 8,540 | SPARSI: Partitioning Sensitive Data amongst Multiple Adversaries | 2013 | VLDB | 5.4119882e-05 |
| 10,065 | On Saving Outliers for Better Clustering over Noisy Data | 2021 | SIGMOD | 5.1648805e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,402 | Approximation Algorithms for Co-Clustering | 2008 | PODS |
| 2 | 9,550 | Data Imputation with Limited Data Redundancy Using Data Lakes | 2025 | VLDB |
| 3 | 11,039 | TARImpute: Task-Aware auto-Recommender System for Missing Value Imputation Algorithms with Clustering Case Studies | 2025 | VLDB |
| 4 | 9,993 | In-Database Data Imputation | 2024 | SIGMOD |
| 5 | 3,886 | GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data | 2023 | SIGMOD |
| 6 | 3,386 | Efficient and Effective Data Imputation with Influence Functions | 2022 | VLDB |
| 7 | 8,002 | Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints | 2020 | SIGMOD |
| 8 | 2,208 | Query Optimization for Dynamic Imputation | 2017 | VLDB |
| 9 | 11,169 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |
| 10 | 5,167 | Enriching Data Imputation with Extensive Similarity Neighbors | 2015 | VLDB |