Scalable K-Means++
Summary: Introduces k-means||, a parallel k-means++ initialization that reduces k sequential data passes to logarithmic—and practically constant—passes. Provably near-optimal and empirically faster than k-means++ on large-scale data. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Bahman Bahmani (Stanford University)
- 2. Benjamin Moseley (University of Illinois Urbana-Champaign)
- 3. Andrea Vattani (University of California)
- 4. Ravi Kumar (Yahoo)
- 5. Sergei Vassilvitskii (Yahoo)
BibTeX Citation
@article{bahmani_vldb12,
title = {{Scalable K-Means++}},
author = {Bahmani, Bahman and Moseley, Benjamin and Vattani, Andrea and Kumar, Ravi and Vassilvitskii, Sergei},
journal = {PVLDB},
series = {{VLDB} '12},
volume = {5},
number = {7},
pages = {622--633},
doi = {10.14778/2180912.2180915},
url = {https://doi.org/10.14778/2180912.2180915},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 13 of 13 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 32 | BIRCH: An Efficient Data Clustering Method for Very Large Databases | 1996 | SIGMOD | 0.00049737458 |
| 363 | CURE: An Efficient Clustering Algorithm for Large Databases | 1998 | SIGMOD | 0.00019987463 |
| 561 | Densest Subgraph in Streaming and MapReduce | 2012 | VLDB | 0.00016396211 |
| 962 | Fast Personalized PageRank on MapReduce | 2011 | SIGMOD | 0.00012820478 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,417 | Theoretically-Efficient and Practical Parallel DBSCAN | 2020 | SIGMOD |
| 2 | 12,348 | K-means Split Revisited: Well-grounded Approach and Experimental Evaluation | 2016 | SIGMOD |
| 3 | 11,062 | Highly-Efficient Large-Scale k-means with Individual Fairness | 2026 | VLDB |
| 4 | 5,438 | Optimal Differentially Private Algorithms for k-Means Clustering | 2018 | PODS |
| 5 | 11,527 | Settling Time vs. Accuracy Tradeoffs for Clustering Big Data | 2024 | SIGMOD |
| 6 | 13,054 | k-Means Projective Clustering | 2004 | PODS |
| 7 | 9,751 | Local Search Methods for k-Means with Outliers | 2017 | VLDB |
| 8 | 4,460 | Solving k-center Clustering (with Outliers) in MapReduce and Streaming, almost as Accurately as Sequentially | 2019 | VLDB |
| 9 | 7,502 | On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection | 2021 | VLDB |
| 10 | 11,508 | Efficient Algorithm for K-Multiple-Means | 2024 | SIGMOD |