DBScholar

Back to papers

BIRCH: An Efficient Data Clustering Method for Very Large Databases

Summary: BIRCH offers incremental, memory-efficient clustering for very large databases; first DB clustering method to effectively handle noise. Shows strong time/space efficiency and single-scan quality; outperforms CLARANS on large datasets. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hbe1b2fa07b9c5f55
Venue
SIGMOD
Year
1996
Pagerank
0.00049714561
Overall Rank
32 | 99.79%
DOI
10.1145/233269.233324

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{zhang_sigmod96,
        title = {{BIRCH: An Efficient Data Clustering Method for Very Large Databases}},
        author = {Zhang, Tian and Ramakrishnan, Raghu and Livny, Miron},
        series = {{SIGMOD} '96},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/233269.233324},
        url = {https://dl.acm.org/doi/10.1145/233269.233324},
        year = {1996}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 85 citing papers.

Rank Citing Paper Year Venue Pagerank
142 LOF: Identifying Density-Based Local Outliers 2000 SIGMOD 0.00029189529
300 OPTICS: Ordering Points To Identify the Clustering Structure 1999 SIGMOD 0.00021800242
308 Automatic Subspace Clustering of High Dimensional Data for Data Mining Applications 1998 SIGMOD 0.00021463972
326 Storing Semistructured Data with STORED 1999 SIGMOD 0.00020953829
364 CURE: An Efficient Clustering Algorithm for Large Databases 1998 SIGMOD 0.00019978187
583 Efficient Algorithms for Mining Outliers from Large Data Sets 2000 SIGMOD 0.00015952617
622 Efficiently Supporting Ad Hoc Queries in Large Datasets of Time Sequences 1997 SIGMOD 0.00015485529
695 Algorithms for Mining Distance-Based Outliers in Large Datasets 1998 VLDB 0.00014696154
917 Trajectory Clustering: A Partition-and-Group Framework 2007 SIGMOD 0.00013085176
928 A Framework for Clustering Evolving Data Streams 2003 VLDB 0.0001301967
1,036 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.00012372946
1,078 Multi-dimensional Selectivity Estimation Using Compressed Histogram Information 1999 SIGMOD 0.00012149796
1,279 Streaming Pattern Discovery in Multiple Time-Series 2005 VLDB 0.00011226666
1,305 STING: A Statistical Information Grid Approach to Spatial Data Mining 1997 VLDB 0.00011094431
1,634 Local Dimensionality Reduction: A New Approach to Indexing High Dimensional Spaces 2000 VLDB 0.00010013999
1,679 Fast Algorithms for Projected Clustering 1999 SIGMOD 9.90334e-05
1,750 Clustering Categorical Data: An Approach Based on Dynamical Systems 1998 VLDB 9.7283911e-05
1,758 Semantic Compression and Pattern Extraction with Fascicles 1999 VLDB 9.7095435e-05
1,788 Plan Selection based on Query Clustering 2002 VLDB 9.6275306e-05
1,892 CS2: A New Database Synopsis for Query Estimation 2013 SIGMOD 9.4150583e-05
1,946 Incremental Clustering for Mining in a Data Warehousing Environment 1998 VLDB 9.3216771e-05
1,970 Information-Theoretic Tools for Mining Database Structure from Large Data Sets 2004 SIGMOD 9.2892451e-05
1,988 WaveCluster: A Multi-Resolution Clustering Approach for Very Large Spatial Databases 1998 VLDB 9.2470974e-05
2,004 SQLEM: Fast Clustering in SQL using the EM Algorithm 2000 SIGMOD 9.205468e-05
2,130 Scalable K-Means++ 2012 VLDB 8.992151e-05
2,243 Automatic Categorization of Query Results 2004 SIGMOD 8.7643883e-05
2,261 Maintaining Variance and k–Medians over Data Stream Windows 2003 PODS 8.7313225e-05
2,307 Approximation Algorithms for Clustering Uncertain Data 2008 PODS 8.6661766e-05
2,650 DEVise: Integrated Querying and Visual Exploration of Large Datasets 1997 SIGMOD 8.1706208e-05
2,747 Approximate XML Joins 2002 SIGMOD 8.0576613e-05
3,017 Indexing the Distance: An Efficient Method to KNN Processing 2001 VLDB 7.7455245e-05
3,248 Approximate XML Query Answers 2004 SIGMOD 7.4912478e-05
3,333 WALRUS: A Similarity Retrieval Algorithm for Image Databases 1999 SIGMOD 7.4129235e-05
3,618 A Monte Carlo Algorithm for Fast Projective Clustering 2002 SIGMOD 7.1525532e-05
3,743 Optimal Grid-Clustering: Towards Breaking the Curse of Dimensionality in High-Dimensional Clustering 1999 VLDB 7.0562784e-05
3,754 VSS: A Storage System for Video Analytics 2021 SIGMOD 7.0473286e-05
3,852 Graph-Based Synopses for Relational Selectivity Estimation 2006 SIGMOD 6.9756022e-05
3,965 Using Trees to Depict a Forest 2009 VLDB 6.8890329e-05
3,989 Compressing Large Boolean Matrices Using Reordering Techniques 2004 VLDB 6.8689183e-05
4,301 Outlier Detection for High Dimensional Data 2001 SIGMOD 6.6757672e-05
4,314 Density Biased Sampling: An Improved Method for Data Mining and Clustering 2000 SIGMOD 6.6668377e-05
4,475 AutoPlait: Automatic Mining of Co-evolving Time Sequences 2014 SIGMOD 6.5810269e-05
4,640 Association Rules over Interval Data 1997 SIGMOD 6.4892961e-05
4,746 Density-based Place Clustering in Geo-Social Networks 2014 SIGMOD 6.4366273e-05
4,788 LinkClus: Efficient Clustering via Heterogeneous Semantic Links 2006 VLDB 6.4120116e-05
4,923 Clustering by Pattern Similarity in Large Data Sets 2002 SIGMOD 6.3493099e-05
5,445 Clustering Stream Data by Exploring the Evolution of Density Mountain 2018 VLDB 6.1247547e-05
6,020 The 3W Model and Algebra for Unified Data Mining 2000 VLDB 5.9101772e-05
6,336 Outlier-robust Clustering using Independent Components 2008 SIGMOD 5.8085039e-05
6,952 Supporting Ranking and Clustering as Generalized Order-By and Group-By 2007 SIGMOD 5.6338956e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 1 of 1 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
94 Efficient and Effective Clustering Methods for Spatial Data Mining 1994 VLDB 0.0003456395
Previous Page 1 / 1 Next

Semantically Similar Papers