DBScholar

Back to papers

Pass-Join: A Partition-based Method for Similarity Joins

Summary: Pass-Join adaptively handles edit-distance similarity joins over both short and long strings via partitioning, inverted indices, and provably minimal substring selection. Novel pruning accelerates candidate verification, outperforming prior methods on real datasets. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hf44ffca463306256
Venue
VLDB
Year
2012
Pagerank
9.3459285e-05
Overall Rank
1,933 | 87.01%
DOI
10.14778/2078331.2078340

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{li_vldb12,
        title = {{Pass-Join: A Partition-based Method for Similarity Joins}},
        author = {Li, Guoliang and Deng, Dong and Wang, Jiannan and Feng, Jianhua},
        journal = {PVLDB},
        series = {{VLDB} '12},
        volume = {5},
        number = {3},
        pages = {253--264},
        doi = {10.14778/2078331.2078340},
        url = {https://doi.org/10.14778/2078331.2078340},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 39 of 39 citing papers.

Rank Citing Paper Year Venue Pagerank
694 JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes 2019 SIGMOD 0.00014727089
963 Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search 2012 SIGMOD 0.00012816649
1,897 iDEC: Indexable Distance Estimating Codes for Approximate Nearest Neighbor Search 2020 VLDB 9.4092345e-05
2,220 String Similarity Joins: An Experimental Evaluation 2014 VLDB 8.8146984e-05
3,350 An Efficient Partition Based Method for Exact Set Similarity Joins 2016 VLDB 7.3910669e-05
3,596 Overlap Set Similarity Joins with Theoretical Guarantees 2018 SIGMOD 7.1790375e-05
3,809 On the Complexity of Inner Product Similarity Join 2016 PODS 7.0089846e-05
4,147 Local Similarity Search for Unstructured Text 2016 SIGMOD 6.7803631e-05
4,274 Set Similarity Joins on MapReduce: An Experimental Survey 2018 VLDB 6.6918599e-05
4,493 Approximate String Joins with Abbreviations 2018 VLDB 6.5772891e-05
4,571 SILKMOTH: An Efficient Method for Finding Related Sets with Maximum Matching Constraints 2017 VLDB 6.5286445e-05
4,688 Learned Cardinality Estimation for Similarity Queries 2021 SIGMOD 6.4697463e-05
4,970 Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples 2021 SIGMOD 6.3329595e-05
5,112 SEAL: Spatio-Textual Similarity Search 2012 VLDB 6.268442e-05
5,207 String Similarity Measures and Joins with Synonyms 2013 SIGMOD 6.2272365e-05
5,444 Pigeonring: A Principle for Faster Thresholded Similarity Search 2019 VLDB 6.1268526e-05
5,912 Dima: A Distributed In-Memory Similarity-Based Query Processing System 2017 VLDB 5.9501002e-05
6,121 The Communication Complexity of Distributed Set-Joins with Applications to Matrix Multiplication 2015 PODS 5.8800995e-05
6,137 Scaling Similarity Joins over Tree-Structured Data 2015 VLDB 5.8736299e-05
6,397 Human-in-the-loop Data Integration 2017 VLDB 5.7989499e-05
6,605 A Pivotal Prefix Based Filtering Algorithm for String Similarity Search 2014 SIGMOD 5.7367357e-05
7,030 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.6173605e-05
7,172 Efficient Error-tolerant Query Autocompletion 2013 VLDB 5.5941906e-05
7,206 Boosting Graph Similarity Search through Pre-Computation 2021 SIGMOD 5.5862482e-05
7,619 SyncSignature: A Simple, Efficient, Parallelizable Framework for Tree Similarity Joins 2023 VLDB 5.4832111e-05
7,709 Scalable Column Concept Determination for Web Tables Using Large Knowledge Bases 2013 VLDB 5.4717883e-05
7,833 Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts 2021 SIGMOD 5.443666e-05
7,906 Near-Duplicate Text Alignment with One Permutation Hashing 2024 SIGMOD 5.428873e-05
8,689 TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection 2022 SIGMOD 5.2905577e-05
9,126 Balance-Aware Distributed String Similarity-Based Query Processing System 2019 VLDB 5.22387e-05
9,495 Set Similarity Search for Skewed Data 2018 PODS 5.1708619e-05
9,879 META: An Efficient Matching-Based Method for Error-Tolerant Autocompletion 2016 VLDB 5.1176637e-05
10,214 Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation 2023 SIGMOD 5.0596605e-05
10,298 Local Filtering: Improving the Performance of Approximate Queries on String Collections 2015 SIGMOD 5.0418674e-05
10,299 Efficient and Effective KNN Sequence Search with Approximate n-grams 2014 VLDB 5.0418674e-05
11,339 Extensible and Robust Evaluation of Similarity Queries 2025 VLDB 4.9793485e-05
11,497 Similarity Joins of Sparse Features 2024 SIGMOD 4.9793485e-05
11,615 Dealing with Acronyms, Abbreviations, and Typos in Real-World Entity Matching 2024 VLDB 4.9793485e-05
12,468 Similarity Joins for Uncertain Strings 2014 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
108 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.0003305531
161 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00027718195
168 Efficient Exact Set-Similarity Joins 2006 VLDB 0.00027163517
201 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025331535
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020009936
929 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013020115
2,072 Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance 2010 SIGMOD 9.0864612e-05
2,354 Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries 2008 VLDB 8.5896515e-05
2,967 Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance 2007 VLDB 7.8029216e-05
3,501 An Efficient Filter for Approximate Membership Checking 2008 SIGMOD 7.2531528e-05
3,512 Efficient Approximate Entity Extraction with Edit Distance Constraints 2009 SIGMOD 7.2439455e-05
3,816 Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints 2010 VLDB 7.0072547e-05
4,439 Incremental Maintenance of Length Normalized Indexes for Approximate String Matching 2009 SIGMOD 6.5998035e-05
4,536 Power-Law Based Estimation of Set Similarity Join Size 2009 VLDB 6.5557137e-05
4,736 Scalable Ad-hoc Entity Extraction from Text Collections 2008 VLDB 6.4451961e-05
5,549 Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction 2011 SIGMOD 6.088104e-05
Previous Page 1 / 1 Next

Semantically Similar Papers