DBScholar

Back to papers

A New Sparse Data Clustering Method Based On Frequent Items

Summary: Proposes k-FreqItems, a scalable clustering method for high-dimensional, sparse categorical data using a sparse FreqItem center and Jaccard distance for interpretable clusters. SILK, an LSH-based seeding technique, oversamples frequent co-occurrences to seed k-FreqItems, delivering faster, more effective initialization and billion-object scalability on commodity GPUs (code: https://github.com/HuangQiang/k-FreqItems). (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6570
Venue
SIGMOD
Year
2023
Pagerank
5.7303405e-05
Overall Rank
6,956 | 52.28%
DOI
10.1145/3588685

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{huang_sigmod23,
        title = {{A New Sparse Data Clustering Method Based On Frequent Items}},
        author = {Huang, Qiang and Luo, Pingyi and Tung, Anthony K. H.},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3588685},
        url = {https://dl.acm.org/doi/10.1145/3588685},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
10,794 SBSC: A fast Self-tuned Bipartite proximity graph-based Spectral Clustering 2025 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 19 of 19 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
27 Fast Algorithms for Mining Association Rules 1994 VLDB 0.00052255472
31 BIRCH: An Efficient Data Clustering Method for Very Large Databases 1996 SIGMOD 0.00050347119
287 Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search 2007 VLDB 0.00022323585
291 OPTICS: Ordering Points To Identify the Clustering Structure 1999 SIGMOD 0.00022264197
332 Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search 2016 VLDB 0.00020920444
351 CURE: An Efficient Clustering Algorithm for Large Databases 1998 SIGMOD 0.00020424271
369 Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting 2012 SIGMOD 0.00019945234
580 SRS: Solving c-Approximate Nearest Neighbor Queries in High Dimensional Euclidean Space with a Tiny Index 2015 VLDB 0.00016157635
581 Quality and Efficiency in High Dimensional Nearest Neighbor Search 2009 SIGMOD 0.00016153395
962 DBSCAN Revisited: Mis-Claim, Un-Fixability, and Approximation 2015 SIGMOD 0.00012936472
1,430 VHP: Approximate Nearest Neighbor Search via Virtual Hypersphere Partitioning 2020 VLDB 0.0001080902
1,546 PM-LSH: A Fast and Accurate LSH Framework for High-Dimensional Approximate NN Search 2020 VLDB 0.00010407159
1,572 LazyLSH: Approximate Nearest Neighbor Search for Multiple Distance Functions with a Single Index 2016 SIGMOD 0.00010329197
2,085 Scalable K-Means++ 2012 VLDB 9.1943614e-05
2,556 NG-DBSCAN: Scalable Density-Based Clustering for Arbitrary Data 2017 VLDB 8.4162172e-05
3,279 Locality-Sensitive Hashing Scheme based on Longest Circular Co-Substring 2020 SIGMOD 7.5711218e-05
5,121 Point-to-Hyperplane Nearest Neighbor Search Beyond the Unit Hypersphere 2021 SIGMOD 6.3577317e-05
5,223 FARGO: Fast Maximum Inner Product Search via Global Multi-Probing 2023 VLDB 6.3102417e-05
9,493 MQH: Locality Sensitive Hashing on Multi-level Quantization Errors for Point-to-Hyperplane Distances 2023 VLDB 5.2621754e-05
Previous Page 1 / 1 Next

Semantically Similar Papers