DBScholar

Back to papers

A New Sparse Data Clustering Method Based On Frequent Items

Summary: Proposes k-FreqItems, a scalable clustering method for high-dimensional, sparse categorical data using a sparse FreqItem center and Jaccard distance for interpretable clusters. SILK, an LSH-based seeding technique, oversamples frequent co-occurrences to seed k-FreqItems, delivering faster, more effective initialization and billion-object scalability on commodity GPUs (code: https://github.com/HuangQiang/k-FreqItems). (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h8e9e5688a29cc64d
Venue
SIGMOD
Year
2023
Pagerank
5.5991152e-05
Overall Rank
7,098 | 52.30%
DOI
10.1145/3588685
PDF
Download (CC BY-NC-SA 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{huang_sigmod23,
        title = {{A New Sparse Data Clustering Method Based On Frequent Items}},
        author = {Huang, Qiang and Luo, Pingyi and Tung, Anthony K. H.},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3588685},
        url = {https://dl.acm.org/doi/10.1145/3588685},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
11,218 SBSC: A fast Self-tuned Bipartite proximity graph-based Spectral Clustering 2025 SIGMOD 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 19 of 19 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
29 Fast Algorithms for Mining Association Rules 1994 VLDB 0.00051189413
32 BIRCH: An Efficient Data Clustering Method for Very Large Databases 1996 SIGMOD 0.00049714561
278 Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search 2007 VLDB 0.00022310642
297 Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search 2016 VLDB 0.00021849337
300 OPTICS: Ordering Points To Identify the Clustering Structure 1999 SIGMOD 0.00021800242
338 Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting 2012 SIGMOD 0.00020600264
364 CURE: An Efficient Clustering Algorithm for Large Databases 1998 SIGMOD 0.00019978187
562 SRS: Solving c-Approximate Nearest Neighbor Queries in High Dimensional Euclidean Space with a Tiny Index 2015 VLDB 0.00016350316
576 Quality and Efficiency in High Dimensional Nearest Neighbor Search 2009 SIGMOD 0.00016118297
986 DBSCAN Revisited: Mis-Claim, Un-Fixability, and Approximation 2015 SIGMOD 0.00012663024
1,325 VHP: Approximate Nearest Neighbor Search via Virtual Hypersphere Partitioning 2020 VLDB 0.00011016872
1,481 PM-LSH: A Fast and Accurate LSH Framework for High-Dimensional Approximate NN Search 2020 VLDB 0.00010535847
1,527 LazyLSH: Approximate Nearest Neighbor Search for Multiple Distance Functions with a Single Index 2016 SIGMOD 0.00010350688
2,130 Scalable K-Means++ 2012 VLDB 8.992151e-05
2,612 NG-DBSCAN: Scalable Density-Based Clustering for Arbitrary Data 2017 VLDB 8.2237609e-05
2,955 Locality-Sensitive Hashing Scheme based on Longest Circular Co-Substring 2020 SIGMOD 7.8111585e-05
5,222 Point-to-Hyperplane Nearest Neighbor Search Beyond the Unit Hypersphere 2021 SIGMOD 6.2205593e-05
5,274 FARGO: Fast Maximum Inner Product Search via Global Multi-Probing 2023 VLDB 6.1985959e-05
9,681 MQH: Locality Sensitive Hashing on Multi-level Quantization Errors for Point-to-Hyperplane Distances 2023 VLDB 5.1416711e-05
Previous Page 1 / 1 Next

Semantically Similar Papers