Efficient Top-k Algorithms for Approximate Substring Matching
Summary: Proposes top-k approximate substring matching over long strings under edit-distance with containment search. It uses novel filtering based on q-grams and inverted q-gram indexes to prune distance computations, achieving scalable, efficient performance on real datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Younghoon Kim (Seoul National University)
- 2. Kyuseok Shim (Seoul National University)
BibTeX Citation
@inproceedings{kim_sigmod13,
title = {{Efficient Top-k Algorithms for Approximate Substring Matching}},
author = {Kim, Younghoon and Shim, Kyuseok},
series = {{SIGMOD} '13},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2463676.2465324},
url = {https://dl.acm.org/doi/10.1145/2463676.2465324},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,997 | Efficient and Effective Similar Subtrajectory Search with Deep Reinforcement Learning | 2020 | VLDB | 6.0146869e-05 |
| 6,484 | A Pivotal Prefix Based Filtering Algorithm for String Similarity Search | 2014 | SIGMOD | 5.8665833e-05 |
| 7,268 | Cardinality Estimation of Approximate Substring Queries using Deep Learning | 2022 | VLDB | 5.6597123e-05 |
| 9,844 | Cardinality Estimation of LIKE Predicate Queries using Deep Learning | 2025 | SIGMOD | 5.2094602e-05 |
| 10,505 | The Case For Language Model Approximated LIKE Predicate | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,899 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance | 2007 | VLDB |
| 2 | 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB |
| 3 | 3,193 | Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme | 2011 | SIGMOD |
| 4 | 6,858 | A Generic Framework for Efficient and Effective Subsequence Retrieval | 2012 | VLDB |
| 5 | 6,484 | A Pivotal Prefix Based Filtering Algorithm for String Similarity Search | 2014 | SIGMOD |
| 6 | 4,776 | An Efficient Index Structure for String Databases | 2001 | VLDB |
| 7 | 911 | Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints | 2008 | VLDB |
| 8 | 10,085 | Local Filtering: Improving the Performance of Approximate Queries on String Collections | 2015 | SIGMOD |
| 9 | 8,114 | Approximate Substring Matching over Uncertain Strings | 2011 | VLDB |
| 10 | 10,086 | Efficient and Effective KNN Sequence Search with Approximate n-grams | 2014 | VLDB |