Similarity Join Size Estimation using Locality Sensitive Hashing
Summary: Introduces LSH-SS, a sampling-based VSJ estimator leveraging Locality-Sensitive Hashing to enable accurate sampling at high similarity thresholds, generalizing SSJ to vector representations (e.g., TF-IDF). Empirical results show LSH-SS delivers higher accuracy and lower variance than random sampling and an adapted SSJ baseline across thresholds on real datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Hongrae Lee (University of British Columbia)
- 2. Raymond T. Ng (University of British Columbia)
- 3. Kyuseok Shim (Seoul National University)
BibTeX Citation
@article{lee_vldb11,
title = {{Similarity Join Size Estimation using Locality Sensitive Hashing}},
author = {Lee, Hongrae and Ng, Raymond T. and Shim, Kyuseok},
journal = {PVLDB},
series = {{VLDB} '11},
volume = {4},
number = {6},
pages = {338--349},
doi = {10.14778/2212351.2212357},
url = {https://doi.org/10.14778/2212351.2212357},
year = {2011}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 965 | Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search | 2012 | SIGMOD | 0.00012810695 |
| 2,217 | Estimating Join Selectivities using Bandwidth-Optimized Kernel Density Models | 2017 | VLDB | 8.8151982e-05 |
| 2,223 | String Similarity Joins: An Experimental Evaluation | 2014 | VLDB | 8.8105347e-05 |
| 4,690 | Learned Cardinality Estimation for Similarity Queries | 2021 | SIGMOD | 6.4667478e-05 |
| 5,210 | String Similarity Measures and Joins with Synonyms | 2013 | SIGMOD | 6.2243094e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 57 | On Random Sampling over Joins | 1999 | SIGMOD | 0.00040095727 |
| 79 | Practical Selectivity Estimation through Adaptive Sampling | 1990 | SIGMOD | 0.00036476265 |
| 91 | On the Propagation of Errors in the Size of Join Results | 1991 | SIGMOD | 0.00034748721 |
| 168 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.00027151132 |
| 201 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025319937 |
| 429 | Tracking Join and Self-Join Sizes in Limited Storage | 1999 | PODS | 0.00018445263 |
| 746 | Bifocal Sampling for Skew-Resistant Join Size Estimation | 1996 | SIGMOD | 0.00014282427 |
| 1,203 | Fixed-Precision Estimation of Join Selectivity | 1993 | PODS | 0.00011543634 |
| 2,354 | Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries | 2008 | VLDB | 8.5872598e-05 |
| 4,538 | Power-Law Based Estimation of Set Similarity Join Size | 2009 | VLDB | 6.5526273e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,923 | Leveraging Set Relations in Exact Set Similarity Join | 2017 | VLDB |
| 2 | 6,140 | Scaling Similarity Joins over Tree-Structured Data | 2015 | VLDB |
| 3 | 5,798 | Smooth Tradeoffs between Insert and Query Complexity in Nearest Neighbor Search | 2015 | PODS |
| 4 | 338 | Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting | 2012 | SIGMOD |
| 5 | 3,811 | On the Complexity of Inner Product Similarity Join | 2016 | PODS |
| 6 | 201 | Efficient set joins on similarity predicates | 2004 | SIGMOD |
| 7 | 6,708 | Distance-Sensitive Hashing | 2018 | PODS |
| 8 | 2,354 | Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries | 2008 | VLDB |
| 9 | 4,538 | Power-Law Based Estimation of Set Similarity Join Size | 2009 | VLDB |
| 10 | 8,416 | Fast Approximate Similarity Join in Vector Databases | 2025 | SIGMOD |