DBScholar

Back to papers

The Case For Language Model Approximated LIKE Predicate

Summary: SMILE reframes wildcard LIKE as neural pattern decoding: a compact column-local language model translates complex LIKE predicates into small candidate sets, then verifies via hash lookups. Yields asymptotic, dataset-size-invariant evaluation, robust to drift, outperforming trigram/B+-tree indexes by large margins. (summarized by gpt-5.4-mini on Apr 11 2026)

Paper ID
h3100bb9b6fb2de85
Venue
SIGMOD
Year
2026
Pagerank
4.9769913e-05
Overall Rank
10,701 | 28.08%
DOI
10.1145/3786703
PDF
Download (CC BY 4.0)

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{li_sigmod26,
        title = {{The Case For Language Model Approximated LIKE Predicate}},
        author = {Li, Yingze and Wang, Dong and Wang, Zixuan and Zhou, Yingli and Yan, Yu and Geng, Jian and Wang, Xinyue and Zeng, Ziqing and Wang, Hongzhi},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3786703},
        url = {https://dl.acm.org/doi/10.1145/3786703},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 43 of 43 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
40 The Case for Learned Index Structures 2018 SIGMOD 0.00046363107
85 Learned Cardinalities: Estimating Correlated Joins with Deep Learning 2019 CIDR 0.00035876108
95 Potter's Wheel: An Interactive Data Cleaning System 2001 VLDB 0.00034367518
108 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.00033040246
161 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00027705594
309 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021376597
406 Deep Unsupervised Cardinality Estimation 2020 VLDB 0.00019050182
510 NeuroCard: One Cardinality Estimator for All Tables 2021 VLDB 0.00017059914
643 Evaluating Top-k Selection Queries 1999 VLDB 0.0001521012
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014687805
768 FITing-Tree: A Data-aware Index Structure 2019 SIGMOD 0.00014107655
962 RankSQL: Query Algebra and Optimization for Relational Top-k Queries 2005 SIGMOD 0.00012818013
981 Cardinality Estimation in DBMS: A Comprehensive Benchmark Evaluation 2022 VLDB 0.00012713454
1,062 VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams 2007 VLDB 0.00012208031
1,269 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.00011254742
1,525 Updatable Learned Index with Precise Positions 2021 VLDB 0.00010355133
1,739 Extending Autocompletion To Tolerate Errors 2009 SIGMOD 9.7448851e-05
2,008 IO-Top-k: Index-access Optimized Top-k Query Processing 2006 VLDB 9.1956849e-05
2,147 Cost-Based Variable-Length-Gram Selection for String Collections to Support Approximate Queries Efficiently 2008 SIGMOD 8.9566452e-05
2,784 Interaction between Record Matching and Data Repairing 2011 SIGMOD 8.0147123e-05
2,839 SNARF: A Learning-Enhanced Range Filter 2022 VLDB 7.9490277e-05
3,424 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.3084429e-05
3,489 Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL 2024 VLDB 7.2594757e-05
3,563 Astrid: Accurate Selectivity Estimation for String Predicates using Deep Learning 2021 VLDB 7.2023194e-05
3,565 Auto-WLM: Machine Learning Enhanced Workload Management in Amazon Redshift 2023 SIGMOD 7.200937e-05
3,708 The Case for a Learned Sorting Algorithm 2020 SIGMOD 7.0781032e-05
4,118 Selectivity Estimation for Fuzzy String Predicates in Large Data Sets 2005 VLDB 6.7927394e-05
4,213 Magneto: Combining Small and Large Language Models for Schema Matching 2025 VLDB 6.7281505e-05
4,299 ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic Workloads 2024 VLDB 6.6766173e-05
4,845 Pattern Functional Dependencies for Data Cleaning 2020 VLDB 6.3808938e-05
4,890 Can Learned Models Replace Hash Functions? 2023 VLDB 6.3663299e-05
5,216 Stage: Query Execution Time Prediction in Amazon Redshift 2024 SIGMOD 6.2218868e-05
5,551 Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction 2011 SIGMOD 6.0852227e-05
6,400 Human-in-the-loop Data Integration 2017 VLDB 5.7962311e-05
6,889 LPLM: A Neural Language Model for Cardinality Estimation of LIKE-Queries 2024 SIGMOD 5.6536655e-05
7,043 Human-in-the-loop Outlier Detection 2020 SIGMOD 5.611642e-05
7,082 LITS: An Optimized Learned Index for Strings 2024 VLDB 5.6010119e-05
7,604 SageDB: An Instance-Optimized Data Analytics System 2022 VLDB 5.4846038e-05
7,846 Efficient Top-k Algorithms for Approximate Substring Matching 2013 SIGMOD 5.4400167e-05
9,154 One Seed, Two Birds: A Unified Learned Structure for Exact and Approximate Counting 2024 SIGMOD 5.2176438e-05
9,710 Repairing Data through Regular Expressions 2016 VLDB 5.1352996e-05
10,032 Cardinality Estimation of LIKE Predicate Queries using Deep Learning 2025 SIGMOD 5.0901632e-05
10,324 Stop Word and Related Problems in Web Interface Integration 2009 VLDB 5.0348633e-05
Previous Page 1 / 1 Next

Semantically Similar Papers