Local Search Methods for k-Means with Outliers
Summary: Introduces a simple local-search algorithm for k-means with outliers, achieving a constant-factor approximation—apparently the first practical provable method for the standard objective. Combines with sketching for scale and outperforms recent heuristics empirically. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Shalmoli Gupta (University of Illinois Urbana-Champaign)
- 2. Ravi Kumar (Google)
- 3. Kefu Lu (Washington University)
- 4. Benjamin Moseley (Washington University)
- 5. Sergei Vassilvitskii (Google)
BibTeX Citation
@article{gupta_vldb17,
title = {{Local Search Methods for k-Means with Outliers}},
author = {Gupta, Shalmoli and Kumar, Ravi and Lu, Kefu and Moseley, Benjamin and Vassilvitskii, Sergei},
journal = {PVLDB},
series = {{VLDB} '17},
volume = {10},
number = {7},
pages = {757--768},
doi = {10.14778/3067421.3067426},
url = {https://doi.org/10.14778/3067421.3067426},
year = {2017}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,664 | Fast Density-Peaks Clustering: Multicore-based Parallelization Approach | 2021 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 142 | LOF: Identifying Density-Based Local Outliers | 2000 | SIGMOD | 0.0002962566 |
| 351 | CURE: An Efficient Clustering Algorithm for Large Databases | 1998 | SIGMOD | 0.00020424271 |
| 578 | Efficient Algorithms for Mining Outliers from Large Data Sets | 2000 | SIGMOD | 0.00016221871 |
| 693 | Algorithms for Mining Distance-Based Outliers in Large Datasets | 1998 | VLDB | 0.00014918477 |
| 2,002 | SQLEM: Fast Clustering in SQL using the EM Algorithm | 2000 | SIGMOD | 9.3312236e-05 |
| 2,085 | Scalable K-Means++ | 2012 | VLDB | 9.1943614e-05 |
| 2,660 | Finding Intensional Knowledge of Distance-Based Outliers | 1999 | VLDB | 8.2869029e-05 |
| 4,209 | Outlier Detection for High Dimensional Data | 2001 | SIGMOD | 6.8319192e-05 |
| 6,205 | Outlier-robust Clustering using Independent Components | 2008 | SIGMOD | 5.9443504e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,205 | Outlier-robust Clustering using Independent Components | 2008 | SIGMOD |
| 2 | 5,314 | Optimal Differentially Private Algorithms for k-Means Clustering | 2018 | PODS |
| 3 | 7,360 | On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection | 2021 | VLDB |
| 4 | 12,764 | k-Means Projective Clustering | 2004 | PODS |
| 5 | 693 | Algorithms for Mining Distance-Based Outliers in Large Datasets | 1998 | VLDB |
| 6 | 578 | Efficient Algorithms for Mining Outliers from Large Data Sets | 2000 | SIGMOD |
| 7 | 10,176 | Clustering with Set Outliers and Applications in Relational Clustering | 2026 | PODS |
| 8 | 9,952 | Distance-Based Outlier Detection: Consolidation and Renewed Bearing | 2010 | VLDB |
| 9 | 4,366 | Solving k-center Clustering (with Outliers) in MapReduce and Streaming, almost as Accurately as Sequentially | 2019 | VLDB |
| 10 | 10,065 | On Saving Outliers for Better Clustering over Noisy Data | 2021 | SIGMOD |