Scalable K-Means++
Summary: Introduces k-means||, a parallel k-means++ initialization that reduces k sequential data passes to logarithmic—and practically constant—passes. Provably near-optimal and empirically faster than k-means++ on large-scale data. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Bahman Bahmani (Stanford University)
- 2. Benjamin Moseley (University of Illinois Urbana-Champaign)
- 3. Andrea Vattani (University of California)
- 4. Ravi Kumar (Yahoo)
- 5. Sergei Vassilvitskii (Yahoo)
BibTeX Citation
@article{bahmani_vldb12,
title = {{Scalable K-Means++}},
author = {Bahmani, Bahman and Moseley, Benjamin and Vattani, Andrea and Kumar, Ravi and Vassilvitskii, Sergei},
journal = {PVLDB},
series = {{VLDB} '12},
volume = {5},
number = {7},
pages = {622--633},
doi = {10.14778/2180912.2180915},
url = {https://doi.org/10.14778/2180912.2180915},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 13 of 13 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 31 | BIRCH: An Efficient Data Clustering Method for Very Large Databases | 1996 | SIGMOD | 0.00050347119 |
| 351 | CURE: An Efficient Clustering Algorithm for Large Databases | 1998 | SIGMOD | 0.00020424271 |
| 564 | Densest Subgraph in Streaming and MapReduce | 2012 | VLDB | 0.00016485347 |
| 945 | Fast Personalized PageRank on MapReduce | 2011 | SIGMOD | 0.00013066956 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,294 | Theoretically-Efficient and Practical Parallel DBSCAN | 2020 | SIGMOD |
| 2 | 12,053 | K-means Split Revisited: Well-grounded Approach and Experimental Evaluation | 2016 | SIGMOD |
| 3 | 10,615 | Highly-Efficient Large-Scale k-means with Individual Fairness | 2026 | VLDB |
| 4 | 5,314 | Optimal Differentially Private Algorithms for k-Means Clustering | 2018 | PODS |
| 5 | 11,184 | Settling Time vs. Accuracy Tradeoffs for Clustering Big Data | 2024 | SIGMOD |
| 6 | 12,764 | k-Means Projective Clustering | 2004 | PODS |
| 7 | 9,575 | Local Search Methods for k-Means with Outliers | 2017 | VLDB |
| 8 | 4,366 | Solving k-center Clustering (with Outliers) in MapReduce and Streaming, almost as Accurately as Sequentially | 2019 | VLDB |
| 9 | 7,360 | On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection | 2021 | VLDB |
| 10 | 11,161 | Efficient Algorithm for K-Multiple-Means | 2024 | SIGMOD |