Marigold: Efficient k-means Clustering in High Dimensions
Summary: Marigold accelerates k-means in high dimensions by aggressively pruning distance computations via a tight distance‑bounding scheme, stepwise multiresolution transforms, and triangle‑inequality exploitation. Novel combination yields near real‑time clustering (≈10× speedup on ARPES and other real-world datasets) without degrading k‑means accuracy. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Kasper Overgaard Mortensen (Aarhus University)
- 2. Fatemeh Zardbani (Aarhus University)
- 3. Mohammad Ahsanul Haque (Aarhus University)
- 4. Steinn Ymir Agustsson (Aarhus University)
- 5. Davide Mottin (Aarhus University)
- 6. Philip Hofmann (Aarhus University)
- 7. Panagiotis Karras (Aarhus University)
BibTeX Citation
@article{mortensen_vldb23,
title = {{Marigold: Efficient k-means Clustering in High Dimensions}},
author = {Mortensen, Kasper Overgaard and Zardbani, Fatemeh and Haque, Mohammad Ahsanul and Agustsson, Steinn Ymir and Mottin, Davide and Hofmann, Philip and Karras, Panagiotis},
journal = {PVLDB},
series = {{VLDB} '23},
volume = {16},
number = {7},
pages = {1740--1748},
doi = {10.14778/3587136.3587147},
url = {https://doi.org/10.14778/3587136.3587147},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,691 | A Flexible Framework for Query-oriented Interactive Community Search | 2025 | VLDB | 5.2351259e-05 |
| 10,959 | Federated and Balanced Clustering for High-dimensional Data | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,053 | Multi-dimensional Selectivity Estimation Using Compressed Histogram Information | 1999 | SIGMOD | 0.00012401532 |
| 2,085 | Scalable K-Means++ | 2012 | VLDB | 9.1943614e-05 |
| 7,360 | On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection | 2021 | VLDB | 5.6340845e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,070 | Towards Metric DBSCAN: Exact, Approximate, and Streaming Algorithms | 2024 | SIGMOD |
| 2 | 11,664 | Fast Density-Peaks Clustering: Multicore-based Parallelization Approach | 2021 | SIGMOD |
| 3 | 12,815 | A Shrinking-Based Approach for Multi-Dimensional Data Analysis | 2003 | VLDB |
| 4 | 1,833 | Finding Generalized Projected Clusters in High Dimensional Spaces | 2000 | SIGMOD |
| 5 | 9,575 | Local Search Methods for k-Means with Outliers | 2017 | VLDB |
| 6 | 8,778 | Evaluating Clustering in Subspace Projections of High Dimensional Data | 2009 | VLDB |
| 7 | 2,085 | Scalable K-Means++ | 2012 | VLDB |
| 8 | 7,360 | On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection | 2021 | VLDB |
| 9 | 11,161 | Efficient Algorithm for K-Multiple-Means | 2024 | SIGMOD |
| 10 | 12,764 | k-Means Projective Clustering | 2004 | PODS |