DBScholar

Back to papers

Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search

Summary: Adaptive framework for similarity join and search that selects per-object prefixes via a cost model, instead of fixed prefix-filtering. Efficient indexes enable dynamic prefix selection, yielding gains vs traditional prefix-filtering baselines. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h2b0ed84fa10c1dea
Venue
SIGMOD
Year
2012
Pagerank
0.00012816649
Overall Rank
963 | 93.53%
DOI
10.1145/2213836.2213847

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{wang_sigmod12,
        title = {{Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search}},
        author = {Wang, Jiannan and Li, Guoliang and Feng, Jianhua},
        series = {{SIGMOD} '12},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2213836.2213847},
        url = {https://dl.acm.org/doi/10.1145/2213836.2213847},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 43 of 43 citing papers.

Rank Citing Paper Year Venue Pagerank
694 JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes 2019 SIGMOD 0.00014727089
2,158 Open Data Integration 2018 VLDB 8.941016e-05
2,220 String Similarity Joins: An Experimental Evaluation 2014 VLDB 8.8146984e-05
2,513 An Empirical Evaluation of Set Similarity Join Techniques 2016 VLDB 8.3679178e-05
2,568 Locality-Sensitive Hashing for Earthquake Detection: A Case Study of Scaling Data-Driven Science 2018 VLDB 8.2870435e-05
2,921 Leveraging Set Relations in Exact Set Similarity Join 2017 VLDB 7.8519256e-05
3,350 An Efficient Partition Based Method for Exact Set Similarity Joins 2016 VLDB 7.3910669e-05
3,596 Overlap Set Similarity Joins with Theoretical Guarantees 2018 SIGMOD 7.1790375e-05
3,809 On the Complexity of Inner Product Similarity Join 2016 PODS 7.0089846e-05
3,839 QASCA: A Quality-Aware Task Assignment System for Crowdsourcing Applications 2015 SIGMOD 6.9921604e-05
4,147 Local Similarity Search for Unstructured Text 2016 SIGMOD 6.7803631e-05
4,493 Approximate String Joins with Abbreviations 2018 VLDB 6.5772891e-05
4,688 Learned Cardinality Estimation for Similarity Queries 2021 SIGMOD 6.4697463e-05
4,970 Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples 2021 SIGMOD 6.3329595e-05
5,207 String Similarity Measures and Joins with Synonyms 2013 SIGMOD 6.2272365e-05
5,444 Pigeonring: A Principle for Faster Thresholded Similarity Search 2019 VLDB 6.1268526e-05
5,764 Question Answering Over Knowledge Graphs: Question Understanding Via Template Decomposition 2018 VLDB 6.0030951e-05
5,912 Dima: A Distributed In-Memory Similarity-Based Query Processing System 2017 VLDB 5.9501002e-05
6,397 Human-in-the-loop Data Integration 2017 VLDB 5.7989499e-05
6,491 Nexus: Correlation Discovery over Collections of Spatio-Temporal Tabular Data 2024 SIGMOD 5.7674551e-05
6,605 A Pivotal Prefix Based Filtering Algorithm for String Similarity Search 2014 SIGMOD 5.7367357e-05
7,030 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.6173605e-05
7,619 SyncSignature: A Simple, Efficient, Parallelizable Framework for Tree Similarity Joins 2023 VLDB 5.4832111e-05
7,709 Scalable Column Concept Determination for Web Tables Using Large Knowledge Bases 2013 VLDB 5.4717883e-05
7,833 Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts 2021 SIGMOD 5.443666e-05
8,689 TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection 2022 SIGMOD 5.2905577e-05
9,126 Balance-Aware Distributed String Similarity-Based Query Processing System 2019 VLDB 5.22387e-05
9,495 Set Similarity Search for Skewed Data 2018 PODS 5.1708619e-05
9,741 LES3: Learning-based Exact Set Similarity Search 2021 VLDB 5.1349531e-05
9,788 On-the-Fly Token Similarity Joins in Relational Databases 2014 SIGMOD 5.1272311e-05
9,878 Towards a Unified Framework for String Similarity Joins 2019 VLDB 5.1176637e-05
10,214 Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation 2023 SIGMOD 5.0596605e-05
10,298 Local Filtering: Improving the Performance of Approximate Queries on String Collections 2015 SIGMOD 5.0418674e-05
10,299 Efficient and Effective KNN Sequence Search with Approximate n-grams 2014 VLDB 5.0418674e-05
10,859 Pail: Efficient kNN Search on Set-Valued Attributes 2026 VLDB 4.9793485e-05
11,339 Extensible and Robust Evaluation of Similarity Queries 2025 VLDB 4.9793485e-05
11,615 Dealing with Acronyms, Abbreviations, and Typos in Real-World Entity Matching 2024 VLDB 4.9793485e-05
11,696 Grouping Time Series for Efficient Columnar Storage 2023 SIGMOD 4.9793485e-05
11,759 A Two-Level Signature Scheme for Stable Set Similarity Joins 2023 VLDB 4.9793485e-05
11,813 TokenJoin: Efficient Filtering for Set Similarity Join with Maximum Weighted Bipartite Matching 2023 VLDB 4.9793485e-05
11,854 OpenTFV: An Open Domain Table-Based Fact Verification System 2022 SIGMOD 4.9793485e-05
12,228 ZigZag: Supporting Similarity Queries on Vector Space Models 2018 SIGMOD 4.9793485e-05
12,574 RCSI: Scalable similarity search in thousand(s) of genomes 2013 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
108 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.0003305531
161 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00027718195
168 Efficient Exact Set-Similarity Joins 2006 VLDB 0.00027163517
201 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025331535
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020009936
929 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013020115
1,061 VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams 2007 VLDB 0.00012213729
1,923 ATLAS: A Probabilistic Algorithm for High Dimensional Similarity Search 2011 SIGMOD 9.3750526e-05
1,933 Pass-Join: A Partition-based Method for Similarity Joins 2012 VLDB 9.3459285e-05
2,072 Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance 2010 SIGMOD 9.0864612e-05
2,313 n-Gram/2L: A Space and Time Efficient Two-Level n-Gram Inverted Index Structure 2005 VLDB 8.6550779e-05
2,354 Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries 2008 VLDB 8.5896515e-05
3,255 Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme 2011 SIGMOD 7.4890984e-05
3,816 Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints 2010 VLDB 7.0072547e-05
4,117 Selectivity Estimation for Fuzzy String Predicates in Large Data Sets 2005 VLDB 6.7958711e-05
4,536 Power-Law Based Estimation of Set Similarity Join Size 2009 VLDB 6.5557137e-05
4,958 Similarity Join Size Estimation using Locality Sensitive Hashing 2011 VLDB 6.3394776e-05
5,549 Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction 2011 SIGMOD 6.088104e-05
Previous Page 1 / 1 Next

Semantically Similar Papers