Back to papers
A Pivotal Prefix Based Filtering Algorithm for String Similarity Search
Summary: Introduces a pivotal prefix filter for string similarity search under edit distance, drastically reducing signatures and boosting pruning power. A dynamic-programming method selects high-quality pivotal prefixes to prune non-consecutive errors, while an alignment filter prunes consecutive errors; experiments on real datasets show order-of-magnitude speedups over baselines.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 4806
- Venue
- SIGMOD
- Year
- 2014
- Pagerank
- 4.9436522e-05
- Overall Rank
- 6,730 | 53.23%
- DOI
-
10.1145/2588555.2593675
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 14 of 14 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 2,729 |
String Similarity Joins: An Experimental Evaluation |
2014 |
VLDB |
8.2175463e-05 |
| 4,247 |
Local Similarity Search for Unstructured Text |
2016 |
SIGMOD |
6.3180334e-05 |
| 4,350 |
Overlap Set Similarity Joins with Theoretical Guarantees |
2018 |
SIGMOD |
6.2576191e-05 |
| 5,180 |
SilkMoth: An Efficient Method for Finding Related Sets with Maximum Matching Constraints |
2017 |
VLDB |
5.6377866e-05 |
| 5,297 |
Fast Subtrajectory Similarity Search in Road Networks under Weighted Edit Distance Constraints |
2020 |
VLDB |
5.5772847e-05 |
| 6,080 |
Pigeonring: A Principle for Faster Thresholded Similarity Search |
2019 |
VLDB |
5.219249e-05 |
| 7,106 |
Efficient Similarity Join and Search on Multi-Attribute Data |
2015 |
SIGMOD |
4.8250163e-05 |
| 7,636 |
Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts |
2021 |
SIGMOD |
4.6863871e-05 |
| 8,284 |
TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection |
2022 |
SIGMOD |
4.5392079e-05 |
| 9,566 |
META: An Efficient Matching-Based Method for Error-Tolerant Autocompletion |
2016 |
VLDB |
4.3212967e-05 |
| 9,831 |
Balance-Aware Distributed String Similarity-Based Query Processing System |
2019 |
VLDB |
4.2710095e-05 |
| 9,875 |
Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation |
2023 |
SIGMOD |
4.2626861e-05 |
| 9,933 |
Local Filtering: Improving the Performance of Approximate Queries on String Collections |
2015 |
SIGMOD |
4.245954e-05 |
| 11,730 |
ZigZag: Supporting Similarity Queries on Vector Space Models |
2018 |
SIGMOD |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 15 of 15 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 125 |
Approximate String Joins in a Database (Almost) for Free |
2001 |
VLDB |
0.00044946098 |
| 155 |
Robust and Efficient Fuzzy Match for Online Data Cleaning |
2003 |
SIGMOD |
0.00040666376 |
| 264 |
Efficient Exact Set-Similarity Joins |
2006 |
VLDB |
0.00029950264 |
| 1,203 |
VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams |
2007 |
VLDB |
0.00013317317 |
| 1,232 |
Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints |
2008 |
VLDB |
0.00013133604 |
| 1,396 |
Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search |
2012 |
SIGMOD |
0.00012215253 |
| 2,077 |
Extending Autocompletion To Tolerate Errors |
2009 |
SIGMOD |
9.6038296e-05 |
| 2,387 |
Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance |
2010 |
SIGMOD |
8.9075095e-05 |
| 2,588 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
8.4872437e-05 |
| 2,729 |
String Similarity Joins: An Experimental Evaluation |
2014 |
VLDB |
8.2175463e-05 |
| 3,779 |
Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme |
2011 |
SIGMOD |
6.7709545e-05 |
| 4,215 |
Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints |
2010 |
VLDB |
6.3464157e-05 |
| 4,878 |
Incremental Maintenance of Length Normalized Indexes for Approximate String Matching |
2009 |
SIGMOD |
5.8537044e-05 |
| 7,142 |
Efficient Error-tolerant Query Autocompletion |
2013 |
VLDB |
4.8151659e-05 |
| 7,707 |
Efficient Top-k Algorithms for Approximate Substring Matching |
2013 |
SIGMOD |
4.6676985e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 2,588 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
8.4872437e-05 |
| 9,934 |
Efficient and Effective KNN Sequence Search with Approximate n-grams |
2014 |
VLDB |
4.245954e-05 |
| 11,987 |
Similarity Joins for Uncertain Strings |
2014 |
SIGMOD |
4.1905499e-05 |
| 1,232 |
Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints |
2008 |
VLDB |
0.00013133604 |
| 7,106 |
Efficient Similarity Join and Search on Multi-Attribute Data |
2015 |
SIGMOD |
4.8250163e-05 |
| 1,396 |
Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search |
2012 |
SIGMOD |
0.00012215253 |
| 4,215 |
Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints |
2010 |
VLDB |
6.3464157e-05 |
| 9,933 |
Local Filtering: Improving the Performance of Approximate Queries on String Collections |
2015 |
SIGMOD |
4.245954e-05 |
| 3,779 |
Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme |
2011 |
SIGMOD |
6.7709545e-05 |
| 7,707 |
Efficient Top-k Algorithms for Approximate Substring Matching |
2013 |
SIGMOD |
4.6676985e-05 |