DBScholar

Back to papers

The Case For Language Model Approximated LIKE Predicate

Summary: SMILE reframes wildcard LIKE as neural pattern decoding: a compact column-local language model translates complex LIKE predicates into small candidate sets, then verifies via hash lookups. Yields asymptotic, dataset-size-invariant evaluation, robust to drift, outperforming trigram/B+-tree indexes by large margins. (summarized by gpt-5.4-mini on Apr 11 2026)

Paper ID
h3100bb9b6fb2de85
Venue
SIGMOD
Year
2026
Pagerank
4.9793485e-05
Overall Rank
10,691 | 28.12%
DOI
10.1145/3786703

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{li_sigmod26,
        title = {{The Case For Language Model Approximated LIKE Predicate}},
        author = {Li, Yingze and Wang, Dong and Wang, Zixuan and Zhou, Yingli and Yan, Yu and Geng, Jian and Wang, Xinyue and Zeng, Ziqing and Wang, Hongzhi},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3786703},
        url = {https://dl.acm.org/doi/10.1145/3786703},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 43 of 43 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
40 The Case for Learned Index Structures 2018 SIGMOD 0.00046284649
85 Learned Cardinalities: Estimating Correlated Joins with Deep Learning 2019 CIDR 0.00035864347
95 Potter's Wheel: An Interactive Data Cleaning System 2001 VLDB 0.00034382643
108 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.0003305531
161 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00027718195
309 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021384073
406 Deep Unsupervised Cardinality Estimation 2020 VLDB 0.00019045544
512 NeuroCard: One Cardinality Estimator for All Tables 2021 VLDB 0.00017050173
643 Evaluating Top-k Selection Queries 1999 VLDB 0.00015217076
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014694048
779 FITing-Tree: A Data-aware Index Structure 2019 SIGMOD 0.00014030069
961 RankSQL: Query Algebra and Optimization for Relational Top-k Queries 2005 SIGMOD 0.0001282305
982 Cardinality Estimation in DBMS: A Comprehensive Benchmark Evaluation 2022 VLDB 0.00012714044
1,061 VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams 2007 VLDB 0.00012213729
1,268 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.0001126007
1,550 Updatable Learned Index with Precise Positions 2021 VLDB 0.00010282449
1,737 Extending Autocompletion To Tolerate Errors 2009 SIGMOD 9.7494976e-05
2,006 IO-Top-k: Index-access Optimized Top-k Query Processing 2006 VLDB 9.199795e-05
2,144 Cost-Based Variable-Length-Gram Selection for String Collections to Support Approximate Queries Efficiently 2008 SIGMOD 8.9608583e-05
2,784 Interaction between Record Matching and Data Repairing 2011 SIGMOD 8.018482e-05
2,837 SNARF: A Learning-Enhanced Range Filter 2022 VLDB 7.9527715e-05
3,424 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.3117029e-05
3,489 Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL 2024 VLDB 7.2627807e-05
3,562 Astrid: Accurate Selectivity Estimation for String Predicates using Deep Learning 2021 VLDB 7.2046519e-05
3,563 Auto-WLM: Machine Learning Enhanced Workload Management in Amazon Redshift 2023 SIGMOD 7.2042148e-05
3,714 The Case for a Learned Sorting Algorithm 2020 SIGMOD 7.0769061e-05
4,117 Selectivity Estimation for Fuzzy String Predicates in Large Data Sets 2005 VLDB 6.7958711e-05
4,218 Magneto: Combining Small and Large Language Models for Schema Matching 2025 VLDB 6.7270169e-05
4,311 ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic Workloads 2024 VLDB 6.6727978e-05
4,843 Pattern Functional Dependencies for Data Cleaning 2020 VLDB 6.3839159e-05
4,890 Can Learned Models Replace Hash Functions? 2023 VLDB 6.3682031e-05
5,214 Stage: Query Execution Time Prediction in Amazon Redshift 2024 SIGMOD 6.2248104e-05
5,549 Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction 2011 SIGMOD 6.088104e-05
6,397 Human-in-the-loop Data Integration 2017 VLDB 5.7989499e-05
6,884 LPLM: A Neural Language Model for Cardinality Estimation of LIKE-Queries 2024 SIGMOD 5.6563432e-05
7,042 Human-in-the-loop Outlier Detection 2020 SIGMOD 5.6142998e-05
7,372 LITS: An Optimized Learned Index for Strings 2024 VLDB 5.5403496e-05
7,598 SageDB: An Instance-Optimized Data Analytics System 2022 VLDB 5.4871733e-05
7,842 Efficient Top-k Algorithms for Approximate Substring Matching 2013 SIGMOD 5.4425932e-05
9,145 One Seed, Two Birds: A Unified Learned Structure for Exact and Approximate Counting 2024 SIGMOD 5.220115e-05
9,705 Repairing Data through Regular Expressions 2016 VLDB 5.1377317e-05
10,027 Cardinality Estimation of LIKE Predicate Queries using Deep Learning 2025 SIGMOD 5.0925739e-05
10,317 Stop Word and Related Problems in Web Interface Integration 2009 VLDB 5.0372479e-05
Previous Page 1 / 1 Next

Semantically Similar Papers