Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions
Summary: Introduces Certain Predictions (CP): a test point is certain when classifiers trained over all possible worlds of incomplete data agree on its label. For nearest-neighbor models, CP checking/counting is tractable despite exponentially many worlds, powering CPClean. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Bojan Karlaš (ETH Zurich)
- 2. Peng Li (Georgia Institute of Technology)
- 3. Renzhi Wu (Georgia Institute of Technology)
- 4. Nezihe Merve Gürel (ETH Zurich)
- 5. Xu Chu (Georgia Institute of Technology)
- 6. Wentao Wu (Microsoft)
- 7. Ce Zhang (ETH Zurich)
BibTeX Citation
@article{karlas_vldb21,
title = {{Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions}},
author = {Karlaš, Bojan and Li, Peng and Wu, Renzhi and Gürel, Nezihe Merve and Chu, Xu and Wu, Wentao and Zhang, Ce},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {3},
pages = {255--267},
doi = {10.14778/3430915.3430917},
url = {https://doi.org/10.14778/3430915.3430917},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 19 of 19 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 33 | Consistent Query Answers in Inconsistent Databases | 1999 | PODS | 0.00049907763 |
| 112 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00032801121 |
| 582 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB | 0.00016148948 |
| 1,066 | Efficient Task-Specific Data Valuation for Nearest Neighbor Algorithms | 2019 | VLDB | 0.00012333161 |
| 1,101 | KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing | 2015 | SIGMOD | 0.00012168934 |
| 1,736 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD | 9.8984415e-05 |
| 4,850 | Nearest-Neighbor Searching Under Uncertainty | 2012 | PODS | 6.480347e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,561 | Efficient Probabilistic Reverse Nearest Neighbor Query Processing on Uncertain Data | 2011 | VLDB |
| 2 | 4,850 | Nearest-Neighbor Searching Under Uncertainty | 2012 | PODS |
| 3 | 4,212 | Data Series Progressive Similarity Search with Probabilistic Quality Guarantees | 2020 | SIGMOD |
| 4 | 1,605 | Efficient Search for the Top-k Probable Nearest Neighbors in Uncertain Databases | 2008 | VLDB |
| 5 | 9,353 | On Efficient Approximate Queries over Machine Learning Models | 2023 | VLDB |
| 6 | 11,790 | Minimization of Classifier Construction Cost for Search Queries | 2020 | SIGMOD |
| 7 | 3,886 | GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data | 2023 | SIGMOD |
| 8 | 9,914 | Explaining k-Nearest Neighbors: Abductive and Counterfactual Explanations | 2025 | PODS |
| 9 | 5,167 | Enriching Data Imputation with Extensive Similarity Neighbors | 2015 | VLDB |
| 10 | 11,169 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |