DBScholar

Back to papers

BIRCH: An Efficient Data Clustering Method for Very Large Databases

Summary: BIRCH offers incremental, memory-efficient clustering for very large databases; first DB clustering method to effectively handle noise. Shows strong time/space efficiency and single-scan quality; outperforms CLARANS on large datasets. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
2937
Venue
SIGMOD
Year
1996
Pagerank
0.00050347119
Overall Rank
31 | 99.79%
DOI
10.1145/233269.233324

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{zhang_sigmod96,
        title = {{BIRCH: An Efficient Data Clustering Method for Very Large Databases}},
        author = {Zhang, Tian and Ramakrishnan, Raghu and Livny, Miron},
        series = {{SIGMOD} '96},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/233269.233324},
        url = {https://dl.acm.org/doi/10.1145/233269.233324},
        year = {1996}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 84 citing papers.

Rank Citing Paper Year Venue Pagerank
142 LOF: Identifying Density-Based Local Outliers 2000 SIGMOD 0.0002962566
291 OPTICS: Ordering Points To Identify the Clustering Structure 1999 SIGMOD 0.00022264197
304 Automatic Subspace Clustering of High Dimensional Data for Data Mining Applications 1998 SIGMOD 0.00021917388
318 Storing Semistructured Data with STORED 1999 SIGMOD 0.00021380325
351 CURE: An Efficient Clustering Algorithm for Large Databases 1998 SIGMOD 0.00020424271
578 Efficient Algorithms for Mining Outliers from Large Data Sets 2000 SIGMOD 0.00016221871
610 Efficiently Supporting Ad Hoc Queries in Large Datasets of Time Sequences 1997 SIGMOD 0.00015783259
693 Algorithms for Mining Distance-Based Outliers in Large Datasets 1998 VLDB 0.00014918477
907 A Framework for Clustering Evolving Data Streams 2003 VLDB 0.00013309819
918 Trajectory Clustering: A Partition-and-Group Framework 2007 SIGMOD 0.00013215285
1,044 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.0001244236
1,053 Multi-dimensional Selectivity Estimation Using Compressed Histogram Information 1999 SIGMOD 0.00012401532
1,252 Streaming Pattern Discovery in Multiple Time-Series 2005 VLDB 0.00011483752
1,285 STING: A Statistical Information Grid Approach to Spatial Data Mining 1997 VLDB 0.00011321748
1,602 Local Dimensionality Reduction: A New Approach to Indexing High Dimensional Spaces 2000 VLDB 0.00010239526
1,646 Fast Algorithms for Projected Clustering 1999 SIGMOD 0.00010128564
1,713 Clustering Categorical Data: An Approach Based on Dynamical Systems 1998 VLDB 9.94811e-05
1,721 Semantic Compression and Pattern Extraction with Fascicles 1999 VLDB 9.9258415e-05
1,771 Plan Selection based on Query Clustering 2002 VLDB 9.7942089e-05
1,893 CS2: A New Database Synopsis for Query Estimation 2013 SIGMOD 9.5269935e-05
1,907 Incremental Clustering for Mining in a Data Warehousing Environment 1998 VLDB 9.501636e-05
1,913 Information-Theoretic Tools for Mining Database Structure from Large Data Sets 2004 SIGMOD 9.4928569e-05
1,941 WaveCluster: A Multi-Resolution Clustering Approach for Very Large Spatial Databases 1998 VLDB 9.4453261e-05
2,002 SQLEM: Fast Clustering in SQL using the EM Algorithm 2000 SIGMOD 9.3312236e-05
2,085 Scalable K-Means++ 2012 VLDB 9.1943614e-05
2,204 Automatic Categorization of Query Results 2004 SIGMOD 8.959723e-05
2,218 Maintaining Variance and k–Medians over Data Stream Windows 2003 PODS 8.9331834e-05
2,253 Approximation Algorithms for Clustering Uncertain Data 2008 PODS 8.8664016e-05
2,632 DEVise: Integrated Querying and Visual Exploration of Large Datasets 1997 SIGMOD 8.3225941e-05
2,699 Approximate XML Joins 2002 SIGMOD 8.2433011e-05
2,979 Indexing the Distance: An Efficient Method to KNN Processing 2001 VLDB 7.8984588e-05
3,180 Approximate XML Query Answers 2004 SIGMOD 7.6599179e-05
3,264 WALRUS: A Similarity Retrieval Algorithm for Image Databases 1999 SIGMOD 7.5841142e-05
3,550 A Monte Carlo Algorithm for Fast Projective Clustering 2002 SIGMOD 7.3200789e-05
3,663 Optimal Grid-Clustering: Towards Breaking the Curse of Dimensionality in High-Dimensional Clustering 1999 VLDB 7.2151113e-05
3,720 VSS: A Storage System for Video Analytics 2021 SIGMOD 7.1739329e-05
3,788 Graph-Based Synopses for Relational Selectivity Estimation 2006 SIGMOD 7.1244416e-05
3,881 Using Trees to Depict a Forest 2009 VLDB 7.0490549e-05
3,912 Compressing Large Boolean Matrices Using Reordering Techniques 2004 VLDB 7.0231849e-05
4,209 Outlier Detection for High Dimensional Data 2001 SIGMOD 6.8319192e-05
4,222 Density Biased Sampling: An Improved Method for Data Mining and Clustering 2000 SIGMOD 6.8227421e-05
4,379 AutoPlait: Automatic Mining of Co-evolving Time Sequences 2014 SIGMOD 6.735265e-05
4,545 Association Rules over Interval Data 1997 SIGMOD 6.6382392e-05
4,651 Density-based Place Clustering in Geo-Social Networks 2014 SIGMOD 6.585234e-05
4,689 LinkClus: Efficient Clustering via Heterogeneous Semantic Links 2006 VLDB 6.5614634e-05
4,812 Clustering by Pattern Similarity in Large Data Sets 2002 SIGMOD 6.4980201e-05
5,335 Clustering Stream Data by Exploring the Evolution of Density Mountain 2018 VLDB 6.2618226e-05
5,895 The 3W Model and Algebra for Unified Data Mining 2000 VLDB 6.0486927e-05
6,205 Outlier-robust Clustering using Independent Components 2008 SIGMOD 5.9443504e-05
6,815 Supporting Ranking and Clustering as Generalized Order-By and Group-By 2007 SIGMOD 5.7652375e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 1 of 1 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
88 Efficient and Effective Clustering Methods for Spatial Data Mining 1994 VLDB 0.00035240327
Previous Page 1 / 1 Next

Semantically Similar Papers