DBScholar

Back to papers

A New Sparse Data Clustering Method Based On Frequent Items

Summary: Proposes k-FreqItems, a scalable clustering method for high-dimensional, sparse categorical data using a sparse FreqItem center and Jaccard distance for interpretable clusters. SILK, an LSH-based seeding technique, oversamples frequent co-occurrences to seed k-FreqItems, delivering faster, more effective initialization and billion-object scalability on commodity GPUs (code: https://github.com/HuangQiang/k-FreqItems). (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h8e9e5688a29cc64d
Venue
SIGMOD
Year
2023
Pagerank
5.601767e-05
Overall Rank
7,096 | 52.30%
DOI
10.1145/3588685

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{huang_sigmod23,
        title = {{A New Sparse Data Clustering Method Based On Frequent Items}},
        author = {Huang, Qiang and Luo, Pingyi and Tung, Anthony K. H.},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3588685},
        url = {https://dl.acm.org/doi/10.1145/3588685},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
11,209 SBSC: A fast Self-tuned Bipartite proximity graph-based Spectral Clustering 2025 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 19 of 19 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
29 Fast Algorithms for Mining Association Rules 1994 VLDB 0.0005121339
32 BIRCH: An Efficient Data Clustering Method for Very Large Databases 1996 SIGMOD 0.00049737458
280 Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search 2007 VLDB 0.0002230467
298 Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search 2016 VLDB 0.00021833987
300 OPTICS: Ordering Points To Identify the Clustering Structure 1999 SIGMOD 0.00021810545
338 Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting 2012 SIGMOD 0.00020585187
363 CURE: An Efficient Clustering Algorithm for Large Databases 1998 SIGMOD 0.00019987463
562 SRS: Solving c-Approximate Nearest Neighbor Queries in High Dimensional Euclidean Space with a Tiny Index 2015 VLDB 0.00016335405
576 Quality and Efficiency in High Dimensional Nearest Neighbor Search 2009 SIGMOD 0.00016121388
986 DBSCAN Revisited: Mis-Claim, Un-Fixability, and Approximation 2015 SIGMOD 0.00012668888
1,328 VHP: Approximate Nearest Neighbor Search via Virtual Hypersphere Partitioning 2020 VLDB 0.00011003106
1,485 PM-LSH: A Fast and Accurate LSH Framework for High-Dimensional Approximate NN Search 2020 VLDB 0.00010523759
1,530 LazyLSH: Approximate Nearest Neighbor Search for Multiple Distance Functions with a Single Index 2016 SIGMOD 0.00010344205
2,128 Scalable K-Means++ 2012 VLDB 8.9964096e-05
2,610 NG-DBSCAN: Scalable Density-Based Clustering for Arbitrary Data 2017 VLDB 8.2276558e-05
2,969 Locality-Sensitive Hashing Scheme based on Longest Circular Co-Substring 2020 SIGMOD 7.8000797e-05
5,238 Point-to-Hyperplane Nearest Neighbor Search Beyond the Unit Hypersphere 2021 SIGMOD 6.2169207e-05
5,268 FARGO: Fast Maximum Inner Product Search via Global Multi-Probing 2023 VLDB 6.2015316e-05
9,674 MQH: Locality Sensitive Hashing on Multi-level Quantization Errors for Point-to-Hyperplane Distances 2023 VLDB 5.1441063e-05
Previous Page 1 / 1 Next

Semantically Similar Papers