LOG-Means: Efficiently Estimating the Number of Clusters in Large Datasets
Summary: LOG-Means estimates the optimal number of clusters with sublinear dependence on the search space, enabling fast tuning on large datasets and Spark. In Apache Spark experiments, it outperforms 13 baselines in runtime and accuracy, delivering the most systematic large-space comparison to date. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Manuel Fritz (University of Stuttgart)
- 2. Michael Behringer (University of Stuttgart)
- 3. Holger Schwarz (University of Stuttgart)
BibTeX Citation
@article{fritz_vldb20,
title = {{LOG-Means: Efficiently Estimating the Number of Clusters in Large Datasets}},
author = {Fritz, Manuel and Behringer, Michael and Schwarz, Holger},
journal = {PVLDB},
series = {{VLDB} '20},
volume = {13},
number = {11},
pages = {2118--2131},
doi = {10.14778/3407790.3407813},
url = {https://doi.org/10.14778/3407790.3407813},
year = {2020}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,959 | Federated and Balanced Clustering for High-dimensional Data | 2025 | VLDB | 5.093636e-05 |
| 11,253 | Ensemble Clustering based on Meta-Learning and Hyperparameter Optimization | 2024 | VLDB | 5.093636e-05 |
| 13,387 | ML2DAC: Meta-learning to Democratize AutoML for Clustering Analyses | 2023 | SIGMOD | - |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,085 | Scalable K-Means++ | 2012 | VLDB | 9.1943614e-05 |
| 4,366 | Solving k-center Clustering (with Outliers) in MapReduce and Streaming, almost as Accurately as Sequentially | 2019 | VLDB | 6.7407122e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,161 | Efficient Algorithm for K-Multiple-Means | 2024 | SIGMOD |
| 2 | 11,664 | Fast Density-Peaks Clustering: Multicore-based Parallelization Approach | 2021 | SIGMOD |
| 3 | 5,294 | Theoretically-Efficient and Practical Parallel DBSCAN | 2020 | SIGMOD |
| 4 | 5,998 | Approximate Distinct Counts for Billions of Datasets | 2019 | SIGMOD |
| 5 | 11,253 | Ensemble Clustering based on Meta-Learning and Hyperparameter Optimization | 2024 | VLDB |
| 6 | 8,070 | Towards Metric DBSCAN: Exact, Approximate, and Streaming Algorithms | 2024 | SIGMOD |
| 7 | 10,794 | SBSC: A fast Self-tuned Bipartite proximity graph-based Spectral Clustering | 2025 | SIGMOD |
| 8 | 11,184 | Settling Time vs. Accuracy Tradeoffs for Clustering Big Data | 2024 | SIGMOD |
| 9 | 7,360 | On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection | 2021 | VLDB |
| 10 | 8,778 | Evaluating Clustering in Subspace Projections of High Dimensional Data | 2009 | VLDB |