DBScholar

Back to papers

BIRCH: An Efficient Data Clustering Method for Very Large Databases

Summary: BIRCH offers incremental, memory-efficient clustering for very large databases; first DB clustering method to effectively handle noise. Shows strong time/space efficiency and single-scan quality; outperforms CLARANS on large datasets. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hbe1b2fa07b9c5f55
Venue
SIGMOD
Year
1996
Pagerank
0.00049737458
Overall Rank
32 | 99.79%
DOI
10.1145/233269.233324

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{zhang_sigmod96,
        title = {{BIRCH: An Efficient Data Clustering Method for Very Large Databases}},
        author = {Zhang, Tian and Ramakrishnan, Raghu and Livny, Miron},
        series = {{SIGMOD} '96},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/233269.233324},
        url = {https://dl.acm.org/doi/10.1145/233269.233324},
        year = {1996}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 85 citing papers.

Rank Citing Paper Year Venue Pagerank
142 LOF: Identifying Density-Based Local Outliers 2000 SIGMOD 0.00029202746
300 OPTICS: Ordering Points To Identify the Clustering Structure 1999 SIGMOD 0.00021810545
308 Automatic Subspace Clustering of High Dimensional Data for Data Mining Applications 1998 SIGMOD 0.00021473921
326 Storing Semistructured Data with STORED 1999 SIGMOD 0.00020963333
363 CURE: An Efficient Clustering Algorithm for Large Databases 1998 SIGMOD 0.00019987463
583 Efficient Algorithms for Mining Outliers from Large Data Sets 2000 SIGMOD 0.00015960125
621 Efficiently Supporting Ad Hoc Queries in Large Datasets of Time Sequences 1997 SIGMOD 0.00015492309
695 Algorithms for Mining Distance-Based Outliers in Large Datasets 1998 VLDB 0.00014702685
917 Trajectory Clustering: A Partition-and-Group Framework 2007 SIGMOD 0.00013091246
928 A Framework for Clustering Evolving Data Streams 2003 VLDB 0.00013025824
1,036 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.00012377471
1,077 Multi-dimensional Selectivity Estimation Using Compressed Histogram Information 1999 SIGMOD 0.00012154948
1,278 Streaming Pattern Discovery in Multiple Time-Series 2005 VLDB 0.00011231952
1,305 STING: A Statistical Information Grid Approach to Spatial Data Mining 1997 VLDB 0.00011099619
1,634 Local Dimensionality Reduction: A New Approach to Indexing High Dimensional Spaces 2000 VLDB 0.0001001862
1,679 Fast Algorithms for Projected Clustering 1999 SIGMOD 9.9079414e-05
1,749 Clustering Categorical Data: An Approach Based on Dynamical Systems 1998 VLDB 9.7329666e-05
1,756 Semantic Compression and Pattern Extraction with Fascicles 1999 VLDB 9.7140303e-05
1,789 Plan Selection based on Query Clustering 2002 VLDB 9.6293635e-05
1,891 CS2: A New Database Synopsis for Query Estimation 2013 SIGMOD 9.4184294e-05
1,945 Incremental Clustering for Mining in a Data Warehousing Environment 1998 VLDB 9.3257256e-05
1,969 Information-Theoretic Tools for Mining Database Structure from Large Data Sets 2004 SIGMOD 9.2935435e-05
1,986 WaveCluster: A Multi-Resolution Clustering Approach for Very Large Spatial Databases 1998 VLDB 9.2514331e-05
2,003 SQLEM: Fast Clustering in SQL using the EM Algorithm 2000 SIGMOD 9.208656e-05
2,128 Scalable K-Means++ 2012 VLDB 8.9964096e-05
2,242 Automatic Categorization of Query Results 2004 SIGMOD 8.7685327e-05
2,260 Maintaining Variance and k–Medians over Data Stream Windows 2003 PODS 8.7353589e-05
2,304 Approximation Algorithms for Clustering Uncertain Data 2008 PODS 8.6702799e-05
2,650 DEVise: Integrated Querying and Visual Exploration of Large Datasets 1997 SIGMOD 8.1744404e-05
2,747 Approximate XML Joins 2002 SIGMOD 8.0614566e-05
3,016 Indexing the Distance: An Efficient Method to KNN Processing 2001 VLDB 7.7483267e-05
3,246 Approximate XML Query Answers 2004 SIGMOD 7.4946879e-05
3,332 WALRUS: A Similarity Retrieval Algorithm for Image Databases 1999 SIGMOD 7.4163853e-05
3,618 A Monte Carlo Algorithm for Fast Projective Clustering 2002 SIGMOD 7.1559402e-05
3,740 Optimal Grid-Clustering: Towards Breaking the Curse of Dimensionality in High-Dimensional Clustering 1999 VLDB 7.0595359e-05
3,753 VSS: A Storage System for Video Analytics 2021 SIGMOD 7.0502734e-05
3,851 Graph-Based Synopses for Relational Selectivity Estimation 2006 SIGMOD 6.9785886e-05
3,963 Using Trees to Depict a Forest 2009 VLDB 6.8922952e-05
3,987 Compressing Large Boolean Matrices Using Reordering Techniques 2004 VLDB 6.8721548e-05
4,299 Outlier Detection for High Dimensional Data 2001 SIGMOD 6.6789252e-05
4,313 Density Biased Sampling: An Improved Method for Data Mining and Clustering 2000 SIGMOD 6.6699869e-05
4,473 AutoPlait: Automatic Mining of Co-evolving Time Sequences 2014 SIGMOD 6.5841437e-05
4,637 Association Rules over Interval Data 1997 SIGMOD 6.4923571e-05
4,743 Density-based Place Clustering in Geo-Social Networks 2014 SIGMOD 6.4396758e-05
4,785 LinkClus: Efficient Clustering via Heterogeneous Semantic Links 2006 VLDB 6.4150472e-05
4,922 Clustering by Pattern Similarity in Large Data Sets 2002 SIGMOD 6.3523165e-05
5,440 Clustering Stream Data by Exploring the Evolution of Density Mountain 2018 VLDB 6.1276555e-05
6,020 The 3W Model and Algebra for Unified Data Mining 2000 VLDB 5.9129763e-05
6,333 Outlier-robust Clustering using Independent Components 2008 SIGMOD 5.8112446e-05
6,949 Supporting Ranking and Clustering as Generalized Order-By and Group-By 2007 SIGMOD 5.6365631e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 1 of 1 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
94 Efficient and Effective Clustering Methods for Spatial Data Mining 1994 VLDB 0.00034579889
Previous Page 1 / 1 Next

Semantically Similar Papers