DBScholar

Back to papers

Set Similarity Search for Skewed Data

Summary: Analyzes set-similarity search for skewed random 0–1 data, targeting high Pearson correlation. Introduces a recursive data-dependent index whose theoretical guarantees exploit heterogeneous item frequencies, explaining heuristic advantages beyond worst-case analyses. (summarized by gpt-5.6-luna on Jul 26 2026)

Paper ID
1764
Venue
PODS
Year
2018
Pagerank
5.289545e-05
Overall Rank
9,317 | 36.08%
DOI
10.1145/3196959.3196985

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{mccauley_pods18,
        address = {New York, NY, USA},
        series = {{PODS} '18},
        title = {{Set Similarity Search for Skewed Data}},
        url = {https://dl.acm.org/doi/10.1145/3196959.3196985},
        doi = {10.1145/3196959.3196985},
        booktitle = {Proceedings of the {ACM} {SIGMOD} Symposium on {Principles} of {Database} {Systems}},
        publisher = {Association for Computing Machinery},
        author = {McCauley, Samuel and Mikkelsen, Jesper W. and Pagh, Rasmus},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
11,447 A Two-Level Signature Scheme for Stable Set Similarity Joins 2023 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 8 of 8 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers