DBScholar

Back to papers

Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search

Summary: Adaptive framework for similarity join and search that selects per-object prefixes via a cost model, instead of fixed prefix-filtering. Efficient indexes enable dynamic prefix selection, yielding gains vs traditional prefix-filtering baselines. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
4576
Venue
SIGMOD
Year
2012
Pagerank
0.00012870645
Overall Rank
975 | 93.32%
DOI
10.1145/2213836.2213847

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{wang_sigmod12,
        title = {{Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search}},
        author = {Wang, Jiannan and Li, Guoliang and Feng, Jianhua},
        series = {{SIGMOD} '12},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2213836.2213847},
        url = {https://dl.acm.org/doi/10.1145/2213836.2213847},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 42 of 42 citing papers.

Rank Citing Paper Year Venue Pagerank
779 JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes 2019 SIGMOD 0.00014092047
2,186 String Similarity Joins: An Experimental Evaluation 2014 VLDB 9.0001436e-05
2,216 Open Data Integration 2018 VLDB 8.9374127e-05
2,501 An Empirical Evaluation of Set Similarity Join Techniques 2016 VLDB 8.4975661e-05
2,541 Locality-Sensitive Hashing for Earthquake Detection: A Case Study of Scaling Data-Driven Science 2018 VLDB 8.4500033e-05
3,040 Leveraging Set Relations in Exact Set Similarity Join 2017 VLDB 7.8262287e-05
3,474 An Efficient Partition Based Method for Exact Set Similarity Joins 2016 VLDB 7.3859271e-05
3,724 Overlap Set Similarity Joins with Theoretical Guarantees 2018 SIGMOD 7.1715735e-05
3,806 On the Complexity of Inner Product Similarity Join 2016 PODS 7.108802e-05
3,904 QASCA: A Quality-Aware Task Assignment System for Crowdsourcing Applications 2015 SIGMOD 7.0304212e-05
4,052 Local Similarity Search for Unstructured Text 2016 SIGMOD 6.935988e-05
4,396 Approximate String Joins with Abbreviations 2018 VLDB 6.7268636e-05
4,617 Learned Cardinality Estimation for Similarity Queries 2021 SIGMOD 6.604437e-05
5,088 String Similarity Measures and Joins with Synonyms 2013 SIGMOD 6.3673276e-05
5,134 Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples 2021 SIGMOD 6.3532024e-05
5,671 Question Answering Over Knowledge Graphs: Question Understanding Via Template Decomposition 2018 VLDB 6.1262839e-05
5,726 Pigeonring: A Principle for Faster Thresholded Similarity Search 2019 VLDB 6.1079184e-05
6,290 Dima: A Distributed In-Memory Similarity-Based Query Processing System 2017 VLDB 5.9253163e-05
6,484 A Pivotal Prefix Based Filtering Algorithm for String Similarity Search 2014 SIGMOD 5.8665833e-05
6,906 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.7418509e-05
7,476 SyncSignature: A Simple, Efficient, Parallelizable Framework for Tree Similarity Joins 2023 VLDB 5.6090634e-05
7,493 Human-in-the-loop Data Integration 2017 VLDB 5.6046905e-05
7,575 Scalable Column Concept Determination for Web Tables Using Large Knowledge Bases 2013 VLDB 5.5937684e-05
7,676 Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts 2021 SIGMOD 5.5686107e-05
8,522 TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection 2022 SIGMOD 5.4119882e-05
8,658 Nexus: Correlation Discovery over Collections of Spatio-Temporal Tabular Data 2024 SIGMOD 5.3904679e-05
9,317 Set Similarity Search for Skewed Data 2018 PODS 5.289545e-05
9,613 On-the-Fly Token Similarity Joins in Relational Databases 2014 SIGMOD 5.2447096e-05
9,702 Towards a Unified Framework for String Similarity Joins 2019 VLDB 5.2351259e-05
9,979 Balance-Aware Distributed String Similarity-Based Query Processing System 2019 VLDB 5.1845938e-05
10,024 Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation 2023 SIGMOD 5.1757914e-05
10,085 Local Filtering: Improving the Performance of Approximate Queries on String Collections 2015 SIGMOD 5.1559617e-05
10,086 Efficient and Effective KNN Sequence Search with Approximate n-grams 2014 VLDB 5.1559617e-05
10,951 Extensible and Robust Evaluation of Similarity Queries 2025 VLDB 5.093636e-05
11,293 Dealing with Acronyms, Abbreviations, and Typos in Real-World Entity Matching 2024 VLDB 5.093636e-05
11,381 Grouping Time Series for Efficient Columnar Storage 2023 SIGMOD 5.093636e-05
11,447 A Two-Level Signature Scheme for Stable Set Similarity Joins 2023 VLDB 5.093636e-05
11,504 TokenJoin: Efficient Filtering for Set Similarity Join with Maximum Weighted Bipartite Matching 2023 VLDB 5.093636e-05
11,545 OpenTFV: An Open Domain Table-Based Fact Verification System 2022 SIGMOD 5.093636e-05
11,702 LES3: Learning-based Exact Set Similarity Search 2021 VLDB 5.093636e-05
11,929 ZigZag: Supporting Similarity Queries on Vector Space Models 2018 SIGMOD 5.093636e-05
12,283 RCSI: Scalable similarity search in thousand(s) of genomes 2013 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
107 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.00033511706
158 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00028199923
169 Efficient Exact Set-Similarity Joins 2006 VLDB 0.0002743469
200 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025597287
356 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020303289
911 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013283031
1,040 VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams 2007 VLDB 0.00012466499
1,886 Pass-Join: A Partition-based Method for Similarity Joins 2012 VLDB 9.5358137e-05
1,935 ATLAS: A Probabilistic Algorithm for High Dimensional Similarity Search 2011 SIGMOD 9.4560124e-05
2,036 Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance 2010 SIGMOD 9.2782094e-05
2,262 n-Gram/2L: A Space and Time Efficient Two-Level n-Gram Inverted Index Structure 2005 VLDB 8.8440146e-05
2,308 Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries 2008 VLDB 8.7738996e-05
3,193 Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme 2011 SIGMOD 7.649474e-05
3,731 Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints 2010 VLDB 7.1663952e-05
4,035 Selectivity Estimation for Fuzzy String Predicates in Large Data Sets 2005 VLDB 6.9439151e-05
4,450 Power-Law Based Estimation of Set Similarity Join Size 2009 VLDB 6.6972929e-05
4,857 Similarity Join Size Estimation using Locality Sensitive Hashing 2011 VLDB 6.4752373e-05
5,412 Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction 2011 SIGMOD 6.2272563e-05
Previous Page 1 / 1 Next

Semantically Similar Papers