Cost-Based Variable-Length-Gram Selection for String Collections to Support Approximate Queries Efficiently
Summary: Cost-based selection of variable-length grams for approximate string queries with VGRAM-style indexing. Dynamic programming yields tight lower bounds on shared grams and enables automatic gram discovery for workloads, linking gram choice to index structure and performance. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Xiaochun Yang (Northeastern University)
- 2. Bin Wang (Northeastern University)
- 3. Chen Li (University of California Irvine)
BibTeX Citation
@inproceedings{yang_sigmod08,
title = {{Cost-Based Variable-Length-Gram Selection for String Collections to Support Approximate Queries Efficiently}},
author = {Yang, Xiaochun and Wang, Bin and Li, Chen},
series = {{SIGMOD} '08},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1376616.1376655},
url = {https://dl.acm.org/doi/10.1145/1376616.1376655},
year = {2008}
}
Incoming Citations (Sorted by Pagerank)
Showing 17 of 17 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 108 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.0003305531 |
| 161 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD | 0.00027718195 |
| 201 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025331535 |
| 262 | Record Linkage: Similarity Measures and Algorithms | 2006 | SIGMOD | 0.00022821392 |
| 1,061 | VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams | 2007 | VLDB | 0.00012213729 |
| 1,163 | Estimating Alphanumeric Selectivity in the Presence of Wildcards | 1996 | SIGMOD | 0.00011750813 |
| 1,461 | Substring Selectivity Estimation | 1999 | PODS | 0.00010583833 |
| 2,313 | n-Gram/2L: A Space and Time Efficient Two-Level n-Gram Inverted Index Structure | 2005 | VLDB | 8.6550779e-05 |
| 2,967 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance | 2007 | VLDB | 7.8029216e-05 |
| 4,117 | Selectivity Estimation for Fuzzy String Predicates in Large Data Sets | 2005 | VLDB | 6.7958711e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,959 | Dynamic Data Structures for Document Collections and Graphs | 2015 | PODS |
| 2 | 10,298 | Local Filtering: Improving the Performance of Approximate Queries on String Collections | 2015 | SIGMOD |
| 3 | 11,437 | An Evaluation of N-Gram Selection Strategies for Regular Expression Indexing in Contemporary Text Analysis Tasks | 2025 | VLDB |
| 4 | 4,439 | Incremental Maintenance of Length Normalized Indexes for Approximate String Matching | 2009 | SIGMOD |
| 5 | 6,952 | Space-efficient Substring Occurrence Estimation | 2011 | PODS |
| 6 | 108 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB |
| 7 | 4,889 | An Efficient Index Structure for String Databases | 2001 | VLDB |
| 8 | 2,967 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance | 2007 | VLDB |
| 9 | 7,842 | Efficient Top-k Algorithms for Approximate Substring Matching | 2013 | SIGMOD |
| 10 | 1,061 | VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams | 2007 | VLDB |