Set Similarity Search for Skewed Data
Summary: Analyzes set-similarity search for skewed random 0–1 data, targeting high Pearson correlation. Introduces a recursive data-dependent index whose theoretical guarantees exploit heterogeneous item frequencies, explaining heuristic advantages beyond worst-case analyses. (summarized by gpt-5.6-luna on Jul 26 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Samuel McCauley (Basic Algorithms Research Copenhagen; IT University of Copenhagen)
- 2. Jesper W. Mikkelsen (IT University of Copenhagen)
- 3. Rasmus Pagh (Basic Algorithms Research Copenhagen; IT University of Copenhagen)
BibTeX Citation
@inproceedings{mccauley_pods18,
address = {New York, NY, USA},
series = {{PODS} '18},
title = {{Set Similarity Search for Skewed Data}},
url = {https://dl.acm.org/doi/10.1145/3196959.3196985},
doi = {10.1145/3196959.3196985},
booktitle = {Proceedings of the {ACM} {SIGMOD} Symposium on {Principles} of {Database} {Systems}},
publisher = {Association for Computing Machinery},
author = {McCauley, Samuel and Mikkelsen, Jesper W. and Pagh, Rasmus},
year = {2018}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,759 | A Two-Level Signature Scheme for Stable Set Similarity Joins | 2023 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 168 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.00027163517 |
| 963 | Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search | 2012 | SIGMOD | 0.00012816649 |
| 1,933 | Pass-Join: A Partition-based Method for Similarity Joins | 2012 | VLDB | 9.3459285e-05 |
| 2,464 | Output-optimal Parallel Algorithms for Similarity Joins | 2017 | PODS | 8.4260608e-05 |
| 2,513 | An Empirical Evaluation of Set Similarity Join Techniques | 2016 | VLDB | 8.3679178e-05 |
| 3,611 | LEMP: Fast Retrieval of Large Entries in a Matrix Product | 2015 | SIGMOD | 7.1628753e-05 |
| 3,809 | On the Complexity of Inner Product Similarity Join | 2016 | PODS | 7.0089846e-05 |
| 5,796 | Smooth Tradeoffs between Insert and Query Complexity in Nearest Neighbor Search | 2015 | PODS | 5.9910143e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,354 | Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries | 2008 | VLDB |
| 2 | 674 | Efficient Similarity Search and Classification via Rank Aggregation | 2003 | SIGMOD |
| 3 | 6,181 | High-Dimensional Vector Similarity Search: From Time Series to Deep Network Embeddings | 2020 | SIGMOD |
| 4 | 2,921 | Leveraging Set Relations in Exact Set Similarity Join | 2017 | VLDB |
| 5 | 3,350 | An Efficient Partition Based Method for Exact Set Similarity Joins | 2016 | VLDB |
| 6 | 7,998 | Set Similarity Join on Probabilistic Data | 2010 | VLDB |
| 7 | 11,081 | A Theoretical Framework for Distribution-Aware Dataset Search | 2025 | PODS |
| 8 | 3,596 | Overlap Set Similarity Joins with Theoretical Guarantees | 2018 | SIGMOD |
| 9 | 168 | Efficient Exact Set-Similarity Joins | 2006 | VLDB |
| 10 | 8,110 | Efficient and Tunable Similar Set Retrieval | 2001 | SIGMOD |