Database Paper Browser

Back to papers

A New Sparse Data Clustering Method Based On Frequent Items

Summary: Proposes k-FreqItems, a scalable clustering method for high-dimensional, sparse categorical data using a sparse FreqItem center and Jaccard distance for interpretable clusters. SILK, an LSH-based seeding technique, oversamples frequent co-occurrences to seed k-FreqItems, delivering faster, more effective initialization and billion-object scalability on commodity GPUs (code: https://github.com/HuangQiang/k-FreqItems). (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6509
Venue
SIGMOD
Year
2023
Pagerank
5.2365238e-05
Overall Rank
6,000 | 58.31%
DOI
10.1145/3588685

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
10,531 SBSC: A fast Self-tuned Bipartite proximity graph-based Spectral Clustering 2025 SIGMOD 4.1905499e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 19 of 19 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
33 BIRCH: An Efficient Data Clustering Method for Very Large Databases 1996 SIGMOD 0.00077399244
36 Fast Algorithms for Mining Association Rules 1994 VLDB 0.00076114894
263 OPTICS: Ordering Points To Identify the Clustering Structure 1999 SIGMOD 0.00029955858
340 CURE: An Efficient Clustering Algorithm for Large Databases 1998 SIGMOD 0.00026854084
399 Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search 2007 VLDB 0.00024359304
579 Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting 2012 SIGMOD 0.0001982328
596 Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search 2016 VLDB 0.00019455943
675 Quality and Efficiency in High Dimensional Nearest Neighbor Search 2009 SIGMOD 0.00018304179
858 SRS: Solving c-Approximate Nearest Neighbor Queries in High Dimensional Euclidean Space with a Tiny Index 2015 VLDB 0.00015833075
918 DBSCAN Revisited: Mis-Claim, Un-Fixability, and Approximation 2015 SIGMOD 0.00015286593
1,934 VHP: Approximate Nearest Neighbor Search via Virtual Hypersphere Partitioning 2020 VLDB 0.00010047294
1,966 LazyLSH: Approximate Nearest Neighbor Search for Multiple Distance Functions with a Single Index 2016 SIGMOD 9.9130791e-05
2,146 Scalable K-Means++ 2012 VLDB 9.4341455e-05
2,160 PM-LSH: A Fast and Accurate LSH Framework for High-Dimensional Approximate NN Search 2020 VLDB 9.4037759e-05
2,669 NG-DBSCAN: Scalable Density-Based Clustering for Arbitrary Data 2017 VLDB 8.3429208e-05
4,230 Locality-Sensitive Hashing Scheme based on Longest Circular Co-Substring 2020 SIGMOD 6.3337893e-05
4,869 Point-to-Hyperplane Nearest Neighbor Search Beyond the Unit Hypersphere 2021 SIGMOD 5.8588434e-05
5,690 FARGO: Fast Maximum Inner Product Search via Global Multi-Probing 2023 VLDB 5.3699421e-05
9,267 MQH: Locality Sensitive Hashing on Multi-level Quantization Errors for Point-to-Hyperplane Distances 2023 VLDB 4.3647353e-05
Previous Page 1 / 1 Next

Semantically Similar Papers