On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection
Summary: UniK unifies pruning-based accelerations for Lloyd's k-means into an evaluation framework with fine-grained performance breakdown. An optimized UniK-hybrid pruning strategy improves efficiency, with ML-based automatic selection of the best accelerator. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sheng Wang (New York University)
- 2. Yuan Sun (Royal Melbourne Institute of Technology University)
- 3. Zhifeng Bao (Royal Melbourne Institute of Technology University)
BibTeX Citation
@article{wang_vldb21,
title = {{On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection}},
author = {Wang, Sheng and Sun, Yuan and Bao, Zhifeng},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {2},
pages = {163--175},
doi = {10.14778/3425879.3425887},
url = {https://doi.org/10.14778/3425879.3425887},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,602 | Marigold: Efficient k-means Clustering in High Dimensions | 2023 | VLDB | 5.404485e-05 |
| 10,615 | Highly-Efficient Large-Scale k-means with Individual Fairness | 2026 | VLDB | 5.093636e-05 |
| 10,959 | Federated and Balanced Clustering for High-dimensional Data | 2025 | VLDB | 5.093636e-05 |
| 11,396 | Prerequisite-driven Fair Clustering on Heterogeneous Information Networks | 2023 | SIGMOD | 5.093636e-05 |
| 11,420 | F3 KM: Federated, Fair, and Fast k-means | 2023 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 56 | M-tree: An Efficient Access Method for Similarity Search in Metric Spaces | 1997 | VLDB | 0.00040719947 |
| 86 | Automatic Database Management System Tuning Through Large-scale Machine Learning | 2017 | SIGMOD | 0.00035316107 |
| 498 | QTune: A Query-Aware Database Tuning System with Deep Reinforcement Learning | 2019 | VLDB | 0.00017440583 |
| 871 | Framework for Evaluating Clustering Algorithms in Duplicate Detection | 2009 | VLDB | 0.00013496531 |
| 2,085 | Scalable K-Means++ | 2012 | VLDB | 9.1943614e-05 |
| 2,556 | NG-DBSCAN: Scalable Density-Based Clustering for Arbitrary Data | 2017 | VLDB | 8.4162172e-05 |
| 4,627 | Pivot-based Metric Indexing | 2017 | VLDB | 6.5963674e-05 |
| 5,112 | Fast Large-Scale Trajectory Clustering | 2020 | VLDB | 6.3596697e-05 |
| 8,778 | Evaluating Clustering in Subspace Projections of High Dimensional Data | 2009 | VLDB | 5.3753402e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,253 | Approximation Algorithms for Clustering Uncertain Data | 2008 | PODS |
| 2 | 8,778 | Evaluating Clustering in Subspace Projections of High Dimensional Data | 2009 | VLDB |
| 3 | 11,144 | Improved Approximation Algorithms for Relational Clustering | 2024 | PODS |
| 4 | 10,615 | Highly-Efficient Large-Scale k-means with Individual Fairness | 2026 | VLDB |
| 5 | 12,764 | k-Means Projective Clustering | 2004 | PODS |
| 6 | 11,184 | Settling Time vs. Accuracy Tradeoffs for Clustering Big Data | 2024 | SIGMOD |
| 7 | 11,253 | Ensemble Clustering based on Meta-Learning and Hyperparameter Optimization | 2024 | VLDB |
| 8 | 11,161 | Efficient Algorithm for K-Multiple-Means | 2024 | SIGMOD |
| 9 | 9,575 | Local Search Methods for k-Means with Outliers | 2017 | VLDB |
| 10 | 2,085 | Scalable K-Means++ | 2012 | VLDB |