Database Paper Browser

Back to papers

BIRCH: An Efficient Data Clustering Method for Very Large Databases

Summary: BIRCH offers incremental, memory-efficient clustering for very large databases; first DB clustering method to effectively handle noise. Shows strong time/space efficiency and single-scan quality; outperforms CLARANS on large datasets. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
2876
Venue
SIGMOD
Year
1996
Pagerank
0.00050802843
Overall Rank
32 | 99.78%
DOI
-

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 50 of 84 citing papers.

Rank Citing Paper Year Venue Pagerank
152 LOF: Identifying Density-Based Local Outliers 2000 SIGMOD 0.00029288871
285 OPTICS: Ordering Points To Identify the Clustering Structure 1999 SIGMOD 0.00022487395
293 Automatic Subspace Clustering of High Dimensional Data for Data Mining Applications 1998 SIGMOD 0.00022194691
311 Storing Semistructured Data with STORED 1999 SIGMOD 0.0002168197
345 CURE: An Efficient Clustering Algorithm for Large Databases 1998 SIGMOD 0.00020727962
607 Efficiently Supporting Ad Hoc Queries in Large Datasets of Time Sequences 1997 SIGMOD 0.00015929565
623 Efficient Algorithms for Mining Outliers from Large Data Sets 2000 SIGMOD 0.00015752968
698 Algorithms for Mining Distance-Based Outliers in Large Datasets 1998 VLDB 0.00015007291
884 A Framework for Clustering Evolving Data Streams 2003 VLDB 0.0001349419
1,041 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.00012565893
1,046 Multi-dimensional Selectivity Estimation Using Compressed Histogram Information 1999 SIGMOD 0.00012552939
1,091 Trajectory Clustering: A Partition-and-Group Framework 2007 SIGMOD 0.00012313678
1,231 Streaming Pattern Discovery in Multiple Time-Series 2005 VLDB 0.00011658901
1,302 STING : A Statistical Information Grid Approach to Spatial Data Mining 1997 VLDB 0.00011321438
1,580 Local Dimensionality Reduction: A New Approach to Indexing High Dimensional Spaces 2000 VLDB 0.00010361788
1,617 Fast Algorithms for Projected Clustering 1999 SIGMOD 0.00010274537
1,691 Clustering Categorical Data: An Approach Based on Dynamical Systems 1998 VLDB 0.00010090912
1,704 Semantic Compression and Pattern Extraction with Fascicles 1999 VLDB 0.00010060319
1,805 Plan Selection based on Query Clustering 2002 VLDB 9.8143517e-05
1,877 Incremental Clustering for Mining in a Data Warehousing Environment 1998 VLDB 9.6577707e-05
1,887 Information-Theoretic Tools for Mining Database Structure from Large Data Sets 2004 SIGMOD 9.6342251e-05
1,909 CS2: A New Database Synopsis for Query Estimation 2013 SIGMOD 9.5931227e-05
1,934 WaveCluster: A Multi-Resolution Clustering Approach for Very Large Spatial Databases 1998 VLDB 9.5346934e-05
2,004 Scalable K-Means++ 2012 VLDB 9.4015684e-05
2,024 SQLEM: Fast Clustering in SQL using the EM Algorithm 2000 SIGMOD 9.3781848e-05
2,164 Automatic Categorization of Query Results 2004 SIGMOD 9.0977344e-05
2,184 Maintaining Variance and k–Medians over Data Stream Windows 2003 PODS 9.0601412e-05
2,596 DEVise: Integrated Querying and Visual Exploration of Large Datasets 1997 SIGMOD 8.4353579e-05
2,651 Approximate XML Joins 2002 SIGMOD 8.3668215e-05
2,758 Approximation Algorithms for Clustering Uncertain Data 2008 PODS 8.2228285e-05
2,939 Indexing the Distance: An Efficient Method to KNN Processing 2001 VLDB 8.0039334e-05
3,143 Approximate XML Query Answers 2004 SIGMOD 7.7696746e-05
3,220 WALRUS: A Similarity Retrieval Algorithm for Image Databases 1999 SIGMOD 7.6967387e-05
3,504 A Monte Carlo Algorithm for Fast Projective Clustering 2002 SIGMOD 7.4309226e-05
3,620 Optimal Grid-Clustering: Towards Breaking the Curse of Dimensionality in High-Dimensional Clustering 1999 VLDB 7.3197965e-05
3,666 VSS: A Storage System for Video Analytics 2021 SIGMOD 7.2781557e-05
3,737 Graph-Based Synopses for Relational Selectivity Estimation 2006 SIGMOD 7.2182521e-05
3,844 Using Trees to Depict a Forest 2009 VLDB 7.1313612e-05
3,846 Compressing Large Boolean Matrices Using Reordering Techniques 2004 VLDB 7.1297616e-05
4,142 Outlier Detection for High Dimensional Data 2001 SIGMOD 6.938149e-05
4,181 Density Biased Sampling: An Improved Method for Data Mining and Clustering 2000 SIGMOD 6.912593e-05
4,309 AutoPlait: Automatic Mining of Co-evolving Time Sequences 2014 SIGMOD 6.8395788e-05
4,477 Association Rules over Interval Data 1997 SIGMOD 6.7374915e-05
4,619 Density-based Place Clustering in Geo-Social Networks 2014 SIGMOD 6.6663058e-05
4,672 LinkClus: Efficient Clustering via Heterogeneous Semantic Links 2006 VLDB 6.6343257e-05
4,753 Clustering by Pattern Similarity in Large Data Sets 2002 SIGMOD 6.5961306e-05
5,304 Clustering Stream Data by Exploring the Evolution of Density Mountain 2018 VLDB 6.343355e-05
5,807 The 3W Model and Algebra for Unified Data Mining 2000 VLDB 6.1423731e-05
6,111 Outlier-robust Clustering using Independent Components 2008 SIGMOD 6.0365798e-05
6,714 Supporting Ranking and Clustering as Generalized Order-By and Group-By 2007 SIGMOD 5.8531292e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 1 of 1 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
88 Efficient and Effective Clustering Methods for Spatial Data Mining 1994 VLDB 0.00035647575
Previous Page 1 / 1 Next

Semantically Similar Papers