DBScholar

Back to papers

Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search

Summary: Adaptive framework for similarity join and search that selects per-object prefixes via a cost model, instead of fixed prefix-filtering. Efficient indexes enable dynamic prefix selection, yielding gains vs traditional prefix-filtering baselines. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h2b0ed84fa10c1dea
Venue
SIGMOD
Year
2012
Pagerank
0.00012810695
Overall Rank
965 | 93.52%
DOI
10.1145/2213836.2213847

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{wang_sigmod12,
        title = {{Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search}},
        author = {Wang, Jiannan and Li, Guoliang and Feng, Jianhua},
        series = {{SIGMOD} '12},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2213836.2213847},
        url = {https://dl.acm.org/doi/10.1145/2213836.2213847},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 43 of 43 citing papers.

Rank Citing Paper Year Venue Pagerank
694 JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes 2019 SIGMOD 0.00014721161
2,160 Open Data Integration 2018 VLDB 8.9373643e-05
2,223 String Similarity Joins: An Experimental Evaluation 2014 VLDB 8.8105347e-05
2,513 An Empirical Evaluation of Set Similarity Join Techniques 2016 VLDB 8.3640659e-05
2,568 Locality-Sensitive Hashing for Earthquake Detection: A Case Study of Scaling Data-Driven Science 2018 VLDB 8.2839378e-05
2,923 Leveraging Set Relations in Exact Set Similarity Join 2017 VLDB 7.8482459e-05
3,350 An Efficient Partition Based Method for Exact Set Similarity Joins 2016 VLDB 7.3875743e-05
3,598 Overlap Set Similarity Joins with Theoretical Guarantees 2018 SIGMOD 7.1756405e-05
3,811 On the Complexity of Inner Product Similarity Join 2016 PODS 7.0061203e-05
3,840 QASCA: A Quality-Aware Task Assignment System for Crowdsourcing Applications 2015 SIGMOD 6.988858e-05
4,147 Local Similarity Search for Unstructured Text 2016 SIGMOD 6.7771533e-05
4,496 Approximate String Joins with Abbreviations 2018 VLDB 6.5741786e-05
4,690 Learned Cardinality Estimation for Similarity Queries 2021 SIGMOD 6.4667478e-05
4,972 Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples 2021 SIGMOD 6.3300201e-05
5,210 String Similarity Measures and Joins with Synonyms 2013 SIGMOD 6.2243094e-05
5,449 Pigeonring: A Principle for Faster Thresholded Similarity Search 2019 VLDB 6.123954e-05
5,765 Question Answering Over Knowledge Graphs: Question Understanding Via Template Decomposition 2018 VLDB 6.0002533e-05
5,914 Dima: A Distributed In-Memory Similarity-Based Query Processing System 2017 VLDB 5.9472869e-05
6,400 Human-in-the-loop Data Integration 2017 VLDB 5.7962311e-05
6,493 Nexus: Correlation Discovery over Collections of Spatio-Temporal Tabular Data 2024 SIGMOD 5.7647249e-05
6,608 A Pivotal Prefix Based Filtering Algorithm for String Similarity Search 2014 SIGMOD 5.73402e-05
7,032 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.6147056e-05
7,625 SyncSignature: A Simple, Efficient, Parallelizable Framework for Tree Similarity Joins 2023 VLDB 5.4806154e-05
7,716 Scalable Column Concept Determination for Web Tables Using Large Knowledge Bases 2013 VLDB 5.4692136e-05
7,837 Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts 2021 SIGMOD 5.4410891e-05
8,697 TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection 2022 SIGMOD 5.2880532e-05
9,136 Balance-Aware Distributed String Similarity-Based Query Processing System 2019 VLDB 5.2213971e-05
9,506 Set Similarity Search for Skewed Data 2018 PODS 5.168414e-05
9,746 LES3: Learning-based Exact Set Similarity Search 2021 VLDB 5.1325223e-05
9,794 On-the-Fly Token Similarity Joins in Relational Databases 2014 SIGMOD 5.1248055e-05
9,885 Towards a Unified Framework for String Similarity Joins 2019 VLDB 5.115241e-05
10,220 Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation 2023 SIGMOD 5.0572653e-05
10,303 Local Filtering: Improving the Performance of Approximate Queries on String Collections 2015 SIGMOD 5.0394806e-05
10,304 Efficient and Effective KNN Sequence Search with Approximate n-grams 2014 VLDB 5.0394806e-05
10,868 Pail: Efficient kNN Search on Set-Valued Attributes 2026 VLDB 4.9769913e-05
11,347 Extensible and Robust Evaluation of Similarity Queries 2025 VLDB 4.9769913e-05
11,621 Dealing with Acronyms, Abbreviations, and Typos in Real-World Entity Matching 2024 VLDB 4.9769913e-05
11,702 Grouping Time Series for Efficient Columnar Storage 2023 SIGMOD 4.9769913e-05
11,765 A Two-Level Signature Scheme for Stable Set Similarity Joins 2023 VLDB 4.9769913e-05
11,819 TokenJoin: Efficient Filtering for Set Similarity Join with Maximum Weighted Bipartite Matching 2023 VLDB 4.9769913e-05
11,860 OpenTFV: An Open Domain Table-Based Fact Verification System 2022 SIGMOD 4.9769913e-05
12,234 ZigZag: Supporting Similarity Queries on Vector Space Models 2018 SIGMOD 4.9769913e-05
12,580 RCSI: Scalable similarity search in thousand(s) of genomes 2013 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
108 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.00033040246
161 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00027705594
168 Efficient Exact Set-Similarity Joins 2006 VLDB 0.00027151132
201 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025319937
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020001237
929 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013014029
1,062 VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams 2007 VLDB 0.00012208031
1,922 ATLAS: A Probabilistic Algorithm for High Dimensional Similarity Search 2011 SIGMOD 9.3726231e-05
1,934 Pass-Join: A Partition-based Method for Similarity Joins 2012 VLDB 9.341845e-05
2,074 Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance 2010 SIGMOD 9.0821759e-05
2,316 n-Gram/2L: A Space and Time Efficient Two-Level n-Gram Inverted Index Structure 2005 VLDB 8.6509997e-05
2,354 Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries 2008 VLDB 8.5872598e-05
3,256 Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme 2011 SIGMOD 7.4855571e-05
3,817 Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints 2010 VLDB 7.0039415e-05
4,118 Selectivity Estimation for Fuzzy String Predicates in Large Data Sets 2005 VLDB 6.7927394e-05
4,538 Power-Law Based Estimation of Set Similarity Join Size 2009 VLDB 6.5526273e-05
4,960 Similarity Join Size Estimation using Locality Sensitive Hashing 2011 VLDB 6.3365273e-05
5,551 Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction 2011 SIGMOD 6.0852227e-05
Previous Page 1 / 1 Next

Semantically Similar Papers