Database Paper Browser

Back to papers

VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams

Summary: VGRAM selects variable-length grams to speed up approximate string queries. It derives query grams from selected grams and links gram-set similarity to edit distance, enabling adoption by algorithms with minimal changes; experiments show speedups. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
9586
Venue
VLDB
Year
2007
Pagerank
0.00013317317
Overall Rank
1,203 | 91.65%
DOI
-

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 28 of 28 citing papers.

Rank Citing Paper Year Venue Pagerank
507 On Active Learning of Record Matching Packages 2010 SIGMOD 0.00021474096
942 Framework for Evaluating Clustering Algorithms in Duplicate Detection 2009 VLDB 0.00015143877
1,232 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013133604
1,396 Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search 2012 SIGMOD 0.00012215253
1,947 WHAM: A High-throughput Sequence Alignment Method 2011 SIGMOD 9.9963826e-05
2,196 Cost-Based Variable-Length-Gram Selection for String Collections to Support Approximate Queries Efficiently 2008 SIGMOD 9.318552e-05
3,583 Efficient Approximate Entity Extraction with Edit Distance Constraints 2009 SIGMOD 6.944299e-05
3,779 Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme 2011 SIGMOD 6.7709545e-05
4,215 Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints 2010 VLDB 6.3464157e-05
4,352 Astrid: Accurate Selectivity Estimation for String Predicates using Deep Learning 2021 VLDB 6.2542257e-05
4,432 Sampling Dirty Data for Matching Attributes 2010 SIGMOD 6.1858589e-05
4,907 Probabilistic String Similarity Joins 2010 SIGMOD 5.8365853e-05
5,071 Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction 2011 SIGMOD 5.7122974e-05
5,297 Fast Subtrajectory Similarity Search in Road Networks under Weighted Edit Distance Constraints 2020 VLDB 5.5772847e-05
5,823 Reference-Based Alignment in Large Sequence Databases 2009 VLDB 5.3122255e-05
5,892 Efficient Approximate Search on String Collections (Tutorial) 2009 VLDB 5.2829767e-05
6,080 Pigeonring: A Principle for Faster Thresholded Similarity Search 2019 VLDB 5.219249e-05
6,351 SigMatch: Fast and Scalable Multi-Pattern Matching 2010 VLDB 5.0956764e-05
6,730 A Pivotal Prefix Based Filtering Algorithm for String Similarity Search 2014 SIGMOD 4.9436522e-05
6,982 A Generic Framework for Efficient and Effective Subsequence Retrieval 2012 VLDB 4.8685999e-05
7,106 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 4.8250163e-05
7,707 Efficient Top-k Algorithms for Approximate Substring Matching 2013 SIGMOD 4.6676985e-05
9,444 On-the-Fly Token Similarity Joins in Relational Databases 2014 SIGMOD 4.3382418e-05
9,831 Balance-Aware Distributed String Similarity-Based Query Processing System 2019 VLDB 4.2710095e-05
9,933 Local Filtering: Improving the Performance of Approximate Queries on String Collections 2015 SIGMOD 4.245954e-05
9,934 Efficient and Effective KNN Sequence Search with Approximate n-grams 2014 VLDB 4.245954e-05
10,216 The Case For Language Model Approximated LIKE Predicate 2026 SIGMOD 4.1905499e-05
11,730 ZigZag: Supporting Similarity Queries on Vector Space Models 2018 SIGMOD 4.1905499e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 8 of 8 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers