Efficient Top-k Algorithms for Approximate Substring Matching
Summary: Proposes top-k approximate substring matching over long strings under edit-distance with containment search. It uses novel filtering based on q-grams and inverted q-gram indexes to prune distance computations, achieving scalable, efficient performance on real datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Younghoon Kim (Seoul National University)
- 2. Kyuseok Shim (Seoul National University)
BibTeX Citation
@inproceedings{kim_sigmod13,
title = {{Efficient Top-k Algorithms for Approximate Substring Matching}},
author = {Kim, Younghoon and Shim, Kyuseok},
series = {{SIGMOD} '13},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2463676.2465324},
url = {https://dl.acm.org/doi/10.1145/2463676.2465324},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,031 | Efficient and Effective Similar Subtrajectory Search with Deep Reinforcement Learning | 2020 | VLDB | 5.9104518e-05 |
| 6,605 | A Pivotal Prefix Based Filtering Algorithm for String Similarity Search | 2014 | SIGMOD | 5.7367357e-05 |
| 7,417 | Cardinality Estimation of Approximate Substring Queries using Deep Learning | 2022 | VLDB | 5.5327978e-05 |
| 10,027 | Cardinality Estimation of LIKE Predicate Queries using Deep Learning | 2025 | SIGMOD | 5.0925739e-05 |
| 10,691 | The Case For Language Model Approximated LIKE Predicate | 2026 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,967 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance | 2007 | VLDB |
| 2 | 108 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB |
| 3 | 3,255 | Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme | 2011 | SIGMOD |
| 4 | 7,006 | A Generic Framework for Efficient and Effective Subsequence Retrieval | 2012 | VLDB |
| 5 | 6,605 | A Pivotal Prefix Based Filtering Algorithm for String Similarity Search | 2014 | SIGMOD |
| 6 | 4,889 | An Efficient Index Structure for String Databases | 2001 | VLDB |
| 7 | 929 | Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints | 2008 | VLDB |
| 8 | 10,298 | Local Filtering: Improving the Performance of Approximate Queries on String Collections | 2015 | SIGMOD |
| 9 | 8,294 | Approximate Substring Matching over Uncertain Strings | 2011 | VLDB |
| 10 | 10,299 | Efficient and Effective KNN Sequence Search with Approximate n-grams | 2014 | VLDB |