Similarity Join Size Estimation using Locality Sensitive Hashing
Summary: Introduces LSH-SS, a sampling-based VSJ estimator leveraging Locality-Sensitive Hashing to enable accurate sampling at high similarity thresholds, generalizing SSJ to vector representations (e.g., TF-IDF). Empirical results show LSH-SS delivers higher accuracy and lower variance than random sampling and an adapted SSJ baseline across thresholds on real datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Hongrae Lee (University of British Columbia)
- 2. Raymond T. Ng (University of British Columbia)
- 3. Kyuseok Shim (Seoul National University)
BibTeX Citation
@article{lee_vldb11,
title = {{Similarity Join Size Estimation using Locality Sensitive Hashing}},
author = {Lee, Hongrae and Ng, Raymond T. and Shim, Kyuseok},
journal = {PVLDB},
series = {{VLDB} '11},
volume = {4},
number = {6},
pages = {338--349},
doi = {10.14778/2212351.2212357},
url = {https://doi.org/10.14778/2212351.2212357},
year = {2011}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 975 | Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search | 2012 | SIGMOD | 0.00012870645 |
| 2,186 | String Similarity Joins: An Experimental Evaluation | 2014 | VLDB | 9.0001436e-05 |
| 2,203 | Estimating Join Selectivities using Bandwidth-Optimized Kernel Density Models | 2017 | VLDB | 8.9610447e-05 |
| 4,617 | Learned Cardinality Estimation for Similarity Queries | 2021 | SIGMOD | 6.604437e-05 |
| 5,088 | String Similarity Measures and Joins with Synonyms | 2013 | SIGMOD | 6.3673276e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 54 | On Random Sampling over Joins | 1999 | SIGMOD | 0.00040810225 |
| 76 | Practical Selectivity Estimation through Adaptive Sampling | 1990 | SIGMOD | 0.00037054261 |
| 89 | On the Propagation of Errors in the Size of Join Results | 1991 | SIGMOD | 0.00035031529 |
| 169 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.0002743469 |
| 200 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025597287 |
| 418 | Tracking Join and Self-Join Sizes in Limited Storage | 1999 | PODS | 0.00018812821 |
| 730 | Bifocal Sampling for Skew-Resistant Join Size Estimation | 1996 | SIGMOD | 0.00014539362 |
| 1,186 | Fixed-Precision Estimation of Join Selectivity | 1993 | PODS | 0.00011764128 |
| 2,308 | Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries | 2008 | VLDB | 8.7738996e-05 |
| 4,450 | Power-Law Based Estimation of Set Similarity Join Size | 2009 | VLDB | 6.6972929e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,040 | Leveraging Set Relations in Exact Set Similarity Join | 2017 | VLDB |
| 2 | 6,014 | Scaling Similarity Joins over Tree-Structured Data | 2015 | VLDB |
| 3 | 5,666 | Smooth Tradeoffs between Insert and Query Complexity in Nearest Neighbor Search | 2015 | PODS |
| 4 | 369 | Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting | 2012 | SIGMOD |
| 5 | 3,806 | On the Complexity of Inner Product Similarity Join | 2016 | PODS |
| 6 | 200 | Efficient set joins on similarity predicates | 2004 | SIGMOD |
| 7 | 6,583 | Distance-Sensitive Hashing | 2018 | PODS |
| 8 | 2,308 | Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries | 2008 | VLDB |
| 9 | 4,450 | Power-Law Based Estimation of Set Similarity Join Size | 2009 | VLDB |
| 10 | 9,052 | Fast Approximate Similarity Join in Vector Databases | 2025 | SIGMOD |