Data Bubbles for Non-Vector Data: Speeding-up Hierarchical Clustering in Arbitrary Metric Spaces
Summary: Data Bubbles, a distance-based summarization for non-vector data, speeds up hierarchical clustering in arbitrary metric spaces. By relying solely on pairwise distances and avoiding vector-space statistics, it yields compact representatives that preserve clustering structure with little quality loss and large runtime gains. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Jianjun Zhou (University of Alberta)
- 2. Jörg Sander (University of Alberta)
BibTeX Citation
@article{zhou_vldb03,
title = {{Data Bubbles for Non-Vector Data: Speeding-up Hierarchical Clustering in Arbitrary Metric Spaces}},
author = {Zhou, Jianjun and Sander, Jörg},
journal = {PVLDB},
series = {{VLDB} '03},
doi = {10.1016/B978-012722442-8/50047-1},
url = {https://doi.org/10.1016/B978-012722442-8/50047-1},
year = {2003}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 31 | BIRCH: An Efficient Data Clustering Method for Very Large Databases | 1996 | SIGMOD | 0.00050347119 |
| 291 | OPTICS: Ordering Points To Identify the Clustering Structure | 1999 | SIGMOD | 0.00022264197 |
| 483 | FastMap: A Fast Algorithm for Indexing, Data-Mining and Visualization of Traditional and Multimedia Datasets | 1995 | SIGMOD | 0.00017756569 |
| 9,013 | Data Bubbles: Quality Preserving Performance Boosting for Hierarchical Clustering | 2001 | SIGMOD | 5.3321448e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,657 | Finding Near Neighbors Through Cluster Pruning | 2007 | PODS |
| 2 | 8,778 | Evaluating Clustering in Subspace Projections of High Dimensional Data | 2009 | VLDB |
| 3 | 907 | A Framework for Clustering Evolving Data Streams | 2003 | VLDB |
| 4 | 5,963 | Hierarchical Subspace Sampling: A Unified Framework for High Dimensional Data Reduction, Selectivity Estimation and Nearest Neighbor Search | 2002 | SIGMOD |
| 5 | 10,153 | Faster Relational Algorithms Using Geometric Data Structures | 2026 | PODS |
| 6 | 11,647 | Data Summarization with Hierarchical Taxonomy | 2021 | SIGMOD |
| 7 | 8,070 | Towards Metric DBSCAN: Exact, Approximate, and Streaming Algorithms | 2024 | SIGMOD |
| 8 | 8,728 | Computing A Well-Representative Summary of Conjunctive Query Results | 2024 | PODS |
| 9 | 9,013 | Data Bubbles: Quality Preserving Performance Boosting for Hierarchical Clustering | 2001 | SIGMOD |
| 10 | 7,291 | Incremental and Effective Data Summarization for Dynamic Hierarchical Clustering | 2004 | SIGMOD |