On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection
Summary: UniK unifies pruning-based accelerations for Lloyd's k-means into an evaluation framework with fine-grained performance breakdown. An optimized UniK-hybrid pruning strategy improves efficiency, with ML-based automatic selection of the best accelerator. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sheng Wang
- 2. Yuan Sun
- 3. Zhifeng Bao
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,512 | Marigold: Efficient k-means Clustering in High Dimensions | 2023 | VLDB | 4.4904708e-05 |
| 10,329 | Highly-Efficient Large-Scale k-means with Individual Fairness | 2026 | VLDB | 4.1905499e-05 |
| 10,723 | Federated and Balanced Clustering for High-dimensional Data | 2025 | VLDB | 4.1905499e-05 |
| 11,195 | Prerequisite-driven Fair Clustering on Heterogeneous Information Networks | 2023 | SIGMOD | 4.1905499e-05 |
| 11,221 | F3 KM: Federated, Fair, and Fast k-means | 2023 | SIGMOD | 4.1905499e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 91 | M-tree: An Efficient Access Method for Similarity Search in Metric Spaces | 1997 | VLDB | 0.00051785122 |
| 183 | Automatic Database Management System Tuning Through Large-scale Machine Learning | 2017 | SIGMOD | 0.00036859633 |
| 779 | QTune: A Query-Aware Database Tuning System with Deep Reinforcement Learning | 2019 | VLDB | 0.00016719473 |
| 942 | Framework for Evaluating Clustering Algorithms in Duplicate Detection | 2009 | VLDB | 0.00015143877 |
| 2,146 | Scalable K-Means++ | 2012 | VLDB | 9.4341455e-05 |
| 2,669 | NG-DBSCAN: Scalable Density-Based Clustering for Arbitrary Data | 2017 | VLDB | 8.3429208e-05 |
| 5,086 | Pivot-based Metric Indexing | 2017 | VLDB | 5.7036022e-05 |
| 5,873 | Fast Large-Scale Trajectory Clustering | 2020 | VLDB | 5.2897929e-05 |
| 8,473 | Evaluating Clustering in Subspace Projections of High Dimensional Data | 2009 | VLDB | 4.4985288e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| Overall Rank | Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,863 | Approximation Algorithms for Clustering Uncertain Data | 2008 | PODS | 0.00010287379 |
| 8,473 | Evaluating Clustering in Subspace Projections of High Dimensional Data | 2009 | VLDB | 4.4985288e-05 |
| 10,928 | Improved Approximation Algorithms for Relational Clustering | 2024 | PODS | 4.1905499e-05 |
| 10,329 | Highly-Efficient Large-Scale k-means with Individual Fairness | 2026 | VLDB | 4.1905499e-05 |
| 12,580 | k-Means Projective Clustering | 2004 | PODS | 4.1905499e-05 |
| 10,974 | Settling Time vs. Accuracy Tradeoffs for Clustering Big Data | 2024 | SIGMOD | 4.1905499e-05 |
| 11,048 | Ensemble Clustering based on Meta-Learning and Hyperparameter Optimization | 2024 | VLDB | 4.1905499e-05 |
| 10,946 | Efficient Algorithm for K-Multiple-Means | 2024 | SIGMOD | 4.1905499e-05 |
| 9,426 | Local Search Methods for k-Means with Outliers | 2017 | VLDB | 4.3399748e-05 |
| 2,146 | Scalable K-Means++ | 2012 | VLDB | 9.4341455e-05 |