Set Similarity Search for Skewed Data
Summary: Analyzes set-similarity search for skewed random 0–1 data, targeting high Pearson correlation. Introduces a recursive data-dependent index whose theoretical guarantees exploit heterogeneous item frequencies, explaining heuristic advantages beyond worst-case analyses. (summarized by gpt-5.6-luna on Jul 26 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Samuel McCauley (Basic Algorithms Research Copenhagen; IT University of Copenhagen)
- 2. Jesper W. Mikkelsen (IT University of Copenhagen)
- 3. Rasmus Pagh (Basic Algorithms Research Copenhagen; IT University of Copenhagen)
BibTeX Citation
@inproceedings{mccauley_pods18,
address = {New York, NY, USA},
series = {{PODS} '18},
title = {{Set Similarity Search for Skewed Data}},
url = {https://dl.acm.org/doi/10.1145/3196959.3196985},
doi = {10.1145/3196959.3196985},
booktitle = {Proceedings of the {ACM} {SIGMOD} Symposium on {Principles} of {Database} {Systems}},
publisher = {Association for Computing Machinery},
author = {McCauley, Samuel and Mikkelsen, Jesper W. and Pagh, Rasmus},
year = {2018}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,447 | A Two-Level Signature Scheme for Stable Set Similarity Joins | 2023 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 169 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.0002743469 |
| 975 | Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search | 2012 | SIGMOD | 0.00012870645 |
| 1,886 | Pass-Join: A Partition-based Method for Similarity Joins | 2012 | VLDB | 9.5358137e-05 |
| 2,462 | Output-optimal Parallel Algorithms for Similarity Joins | 2017 | PODS | 8.5487602e-05 |
| 2,501 | An Empirical Evaluation of Set Similarity Join Techniques | 2016 | VLDB | 8.4975661e-05 |
| 3,676 | LEMP: Fast Retrieval of Large Entries in a Matrix Product | 2015 | SIGMOD | 7.2094091e-05 |
| 3,806 | On the Complexity of Inner Product Similarity Join | 2016 | PODS | 7.108802e-05 |
| 5,666 | Smooth Tradeoffs between Insert and Query Complexity in Nearest Neighbor Search | 2015 | PODS | 6.1285141e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,308 | Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries | 2008 | VLDB |
| 2 | 690 | Efficient Similarity Search and Classification via Rank Aggregation | 2003 | SIGMOD |
| 3 | 6,093 | High-Dimensional Vector Similarity Search: From Time Series to Deep Network Embeddings | 2020 | SIGMOD |
| 4 | 3,040 | Leveraging Set Relations in Exact Set Similarity Join | 2017 | VLDB |
| 5 | 3,474 | An Efficient Partition Based Method for Exact Set Similarity Joins | 2016 | VLDB |
| 6 | 7,839 | Set Similarity Join on Probabilistic Data | 2010 | VLDB |
| 7 | 10,638 | A Theoretical Framework for Distribution-Aware Dataset Search | 2025 | PODS |
| 8 | 3,724 | Overlap Set Similarity Joins with Theoretical Guarantees | 2018 | SIGMOD |
| 9 | 169 | Efficient Exact Set-Similarity Joins | 2006 | VLDB |
| 10 | 7,942 | Efficient and Tunable Similar Set Retrieval | 2001 | SIGMOD |