DBScholar

Back to papers

Pass-Join: A Partition-based Method for Similarity Joins

Summary: Pass-Join adaptively handles edit-distance similarity joins over both short and long strings via partitioning, inverted indices, and provably minimal substring selection. Novel pruning accelerates candidate verification, outperforming prior methods on real datasets. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hf44ffca463306256
Venue
VLDB
Year
2012
Pagerank
9.341845e-05
Overall Rank
1,934 | 87.01%
DOI
10.14778/2078331.2078340

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{li_vldb12,
        title = {{Pass-Join: A Partition-based Method for Similarity Joins}},
        author = {Li, Guoliang and Deng, Dong and Wang, Jiannan and Feng, Jianhua},
        journal = {PVLDB},
        series = {{VLDB} '12},
        volume = {5},
        number = {3},
        pages = {253--264},
        doi = {10.14778/2078331.2078340},
        url = {https://doi.org/10.14778/2078331.2078340},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 39 of 39 citing papers.

Rank Citing Paper Year Venue Pagerank
694 JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes 2019 SIGMOD 0.00014721161
965 Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search 2012 SIGMOD 0.00012810695
1,894 iDEC: Indexable Distance Estimating Codes for Approximate Nearest Neighbor Search 2020 VLDB 9.4129766e-05
2,223 String Similarity Joins: An Experimental Evaluation 2014 VLDB 8.8105347e-05
3,350 An Efficient Partition Based Method for Exact Set Similarity Joins 2016 VLDB 7.3875743e-05
3,598 Overlap Set Similarity Joins with Theoretical Guarantees 2018 SIGMOD 7.1756405e-05
3,811 On the Complexity of Inner Product Similarity Join 2016 PODS 7.0061203e-05
4,147 Local Similarity Search for Unstructured Text 2016 SIGMOD 6.7771533e-05
4,275 Set Similarity Joins on MapReduce: An Experimental Survey 2018 VLDB 6.6887374e-05
4,496 Approximate String Joins with Abbreviations 2018 VLDB 6.5741786e-05
4,573 SILKMOTH: An Efficient Method for Finding Related Sets with Maximum Matching Constraints 2017 VLDB 6.5257221e-05
4,690 Learned Cardinality Estimation for Similarity Queries 2021 SIGMOD 6.4667478e-05
4,972 Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples 2021 SIGMOD 6.3300201e-05
5,115 SEAL: Spatio-Textual Similarity Search 2012 VLDB 6.2654746e-05
5,210 String Similarity Measures and Joins with Synonyms 2013 SIGMOD 6.2243094e-05
5,449 Pigeonring: A Principle for Faster Thresholded Similarity Search 2019 VLDB 6.123954e-05
5,914 Dima: A Distributed In-Memory Similarity-Based Query Processing System 2017 VLDB 5.9472869e-05
6,122 The Communication Complexity of Distributed Set-Joins with Applications to Matrix Multiplication 2015 PODS 5.8773159e-05
6,140 Scaling Similarity Joins over Tree-Structured Data 2015 VLDB 5.8708494e-05
6,400 Human-in-the-loop Data Integration 2017 VLDB 5.7962311e-05
6,608 A Pivotal Prefix Based Filtering Algorithm for String Similarity Search 2014 SIGMOD 5.73402e-05
7,032 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.6147056e-05
7,175 Efficient Error-tolerant Query Autocompletion 2013 VLDB 5.5915423e-05
7,208 Boosting Graph Similarity Search through Pre-Computation 2021 SIGMOD 5.5836037e-05
7,625 SyncSignature: A Simple, Efficient, Parallelizable Framework for Tree Similarity Joins 2023 VLDB 5.4806154e-05
7,716 Scalable Column Concept Determination for Web Tables Using Large Knowledge Bases 2013 VLDB 5.4692136e-05
7,837 Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts 2021 SIGMOD 5.4410891e-05
7,910 Near-Duplicate Text Alignment with One Permutation Hashing 2024 SIGMOD 5.426303e-05
8,697 TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection 2022 SIGMOD 5.2880532e-05
9,136 Balance-Aware Distributed String Similarity-Based Query Processing System 2019 VLDB 5.2213971e-05
9,506 Set Similarity Search for Skewed Data 2018 PODS 5.168414e-05
9,886 META: An Efficient Matching-Based Method for Error-Tolerant Autocompletion 2016 VLDB 5.115241e-05
10,220 Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation 2023 SIGMOD 5.0572653e-05
10,303 Local Filtering: Improving the Performance of Approximate Queries on String Collections 2015 SIGMOD 5.0394806e-05
10,304 Efficient and Effective KNN Sequence Search with Approximate n-grams 2014 VLDB 5.0394806e-05
11,347 Extensible and Robust Evaluation of Similarity Queries 2025 VLDB 4.9769913e-05
11,503 Similarity Joins of Sparse Features 2024 SIGMOD 4.9769913e-05
11,621 Dealing with Acronyms, Abbreviations, and Typos in Real-World Entity Matching 2024 VLDB 4.9769913e-05
12,474 Similarity Joins for Uncertain Strings 2014 SIGMOD 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
108 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.00033040246
161 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00027705594
168 Efficient Exact Set-Similarity Joins 2006 VLDB 0.00027151132
201 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025319937
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020001237
929 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013014029
2,074 Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance 2010 SIGMOD 9.0821759e-05
2,354 Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries 2008 VLDB 8.5872598e-05
2,969 Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance 2007 VLDB 7.7993378e-05
3,501 An Efficient Filter for Approximate Membership Checking 2008 SIGMOD 7.2497317e-05
3,512 Efficient Approximate Entity Extraction with Edit Distance Constraints 2009 SIGMOD 7.2405277e-05
3,817 Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints 2010 VLDB 7.0039415e-05
4,441 Incremental Maintenance of Length Normalized Indexes for Approximate String Matching 2009 SIGMOD 6.5966903e-05
4,538 Power-Law Based Estimation of Set Similarity Join Size 2009 VLDB 6.5526273e-05
4,737 Scalable Ad-hoc Entity Extraction from Text Collections 2008 VLDB 6.442156e-05
5,551 Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction 2011 SIGMOD 6.0852227e-05
Previous Page 1 / 1 Next

Semantically Similar Papers