Cost-Based Variable-Length-Gram Selection for String Collections to Support Approximate Queries Efficiently
Summary: Cost-based selection of variable-length grams for approximate string queries with VGRAM-style indexing. Dynamic programming yields tight lower bounds on shared grams and enables automatic gram discovery for workloads, linking gram choice to index structure and performance. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Xiaochun Yang (Northeastern University)
- 2. Bin Wang (Northeastern University)
- 3. Chen Li (University of California Irvine)
BibTeX Citation
@inproceedings{yang_sigmod08,
title = {{Cost-Based Variable-Length-Gram Selection for String Collections to Support Approximate Queries Efficiently}},
author = {Yang, Xiaochun and Wang, Bin and Li, Chen},
series = {{SIGMOD} '08},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1376616.1376655},
url = {https://dl.acm.org/doi/10.1145/1376616.1376655},
year = {2008}
}
Incoming Citations (Sorted by Pagerank)
Showing 17 of 17 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.00033511706 |
| 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD | 0.00028199923 |
| 200 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025597287 |
| 254 | Record Linkage: Similarity Measures and Algorithms | 2006 | SIGMOD | 0.00023199211 |
| 1,040 | VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams | 2007 | VLDB | 0.00012466499 |
| 1,140 | Estimating Alphanumeric Selectivity in the Presence of Wildcards | 1996 | SIGMOD | 0.0001200574 |
| 1,427 | Substring Selectivity Estimation | 1999 | PODS | 0.00010812749 |
| 2,262 | n-Gram/2L: A Space and Time Efficient Two-Level n-Gram Inverted Index Structure | 2005 | VLDB | 8.8440146e-05 |
| 2,899 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance | 2007 | VLDB | 7.97814e-05 |
| 4,035 | Selectivity Estimation for Fuzzy String Predicates in Large Data Sets | 2005 | VLDB | 6.9439151e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,810 | Dynamic Data Structures for Document Collections and Graphs | 2015 | PODS |
| 2 | 10,085 | Local Filtering: Improving the Performance of Approximate Queries on String Collections | 2015 | SIGMOD |
| 3 | 11,081 | An Evaluation of N-Gram Selection Strategies for Regular Expression Indexing in Contemporary Text Analysis Tasks | 2025 | VLDB |
| 4 | 4,359 | Incremental Maintenance of Length Normalized Indexes for Approximate String Matching | 2009 | SIGMOD |
| 5 | 6,818 | Space-efficient Substring Occurrence Estimation | 2011 | PODS |
| 6 | 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB |
| 7 | 4,776 | An Efficient Index Structure for String Databases | 2001 | VLDB |
| 8 | 2,899 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance | 2007 | VLDB |
| 9 | 7,692 | Efficient Top-k Algorithms for Approximate Substring Matching | 2013 | SIGMOD |
| 10 | 1,040 | VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams | 2007 | VLDB |