Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions
Summary: Introduces Certain Predictions (CP): a test point is certain when classifiers trained over all possible worlds of incomplete data agree on its label. For nearest-neighbor models, CP checking/counting is tractable despite exponentially many worlds, powering CPClean. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Bojan Karlaš (ETH Zurich)
- 2. Peng Li (Georgia Institute of Technology)
- 3. Renzhi Wu (Georgia Institute of Technology)
- 4. Nezihe Merve Gürel (ETH Zurich)
- 5. Xu Chu (Georgia Institute of Technology)
- 6. Wentao Wu (Microsoft)
- 7. Ce Zhang (ETH Zurich)
BibTeX Citation
@article{karlas_vldb21,
title = {{Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions}},
author = {Karlaš, Bojan and Li, Peng and Wu, Renzhi and Gürel, Nezihe Merve and Chu, Xu and Wu, Wentao and Zhang, Ce},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {3},
pages = {255--267},
doi = {10.14778/3430915.3430917},
url = {https://doi.org/10.14778/3430915.3430917},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 21 of 21 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 33 | Consistent Query Answers in Inconsistent Databases | 1999 | PODS | 0.00049349003 |
| 104 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00033690989 |
| 483 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB | 0.00017590977 |
| 894 | Efficient Task-Specific Data Valuation for Nearest Neighbor Algorithms | 2019 | VLDB | 0.00013212578 |
| 1,099 | KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing | 2015 | SIGMOD | 0.00012037058 |
| 1,720 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD | 9.7965659e-05 |
| 4,835 | Nearest-Neighbor Searching Under Uncertainty | 2012 | PODS | 6.3877419e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,835 | Nearest-Neighbor Searching Under Uncertainty | 2012 | PODS |
| 2 | 10,975 | Minimal Data Cleaning for Model Training by MinPrep | 2026 | VLDB |
| 3 | 4,304 | Data Series Progressive Similarity Search with Probabilistic Quality Guarantees | 2020 | SIGMOD |
| 4 | 1,635 | Efficient Search for the Top-k Probable Nearest Neighbors in Uncertain Databases | 2008 | VLDB |
| 5 | 5,911 | On Efficient Approximate Queries over Machine Learning Models | 2023 | VLDB |
| 6 | 12,092 | Minimization of Classifier Construction Cost for Search Queries | 2020 | SIGMOD |
| 7 | 3,941 | GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data | 2023 | SIGMOD |
| 8 | 10,101 | Explaining k-Nearest Neighbors: Abductive and Counterfactual Explanations | 2025 | PODS |
| 9 | 5,278 | Enriching Data Imputation with Extensive Similarity Neighbors | 2015 | VLDB |
| 10 | 11,514 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |