DBScholar

Back to papers

The Case For Language Model Approximated LIKE Predicate

Summary: SMILE reframes wildcard LIKE as neural pattern decoding: a compact column-local language model translates complex LIKE predicates into small candidate sets, then verifies via hash lookups. Yields asymptotic, dataset-size-invariant evaluation, robust to drift, outperforming trigram/B+-tree indexes by large margins. (summarized by gpt-5.4-mini on Apr 11 2026)

Paper ID
7719
Venue
SIGMOD
Year
2026
Pagerank
5.093636e-05
Overall Rank
10,505 | 27.93%
DOI
10.1145/3786703

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{li_sigmod26,
        title = {{The Case For Language Model Approximated LIKE Predicate}},
        author = {Li, Yingze and Wang, Dong and Wang, Zixuan and Zhou, Yingli and Yan, Yu and Geng, Jian and Wang, Xinyue and Zeng, Ziqing and Wang, Hongzhi},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3786703},
        url = {https://dl.acm.org/doi/10.1145/3786703},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 43 of 43 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
43 The Case for Learned Index Structures 2018 SIGMOD 0.00046060254
84 Learned Cardinalities: Estimating Correlated Joins with Deep Learning 2019 CIDR 0.00035838391
94 Potter's Wheel: An Interactive Data Cleaning System 2001 VLDB 0.00034616103
107 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.00033511706
158 Robust and Efficient Fuzzy Match for Online Data Cleaning 2003 SIGMOD 0.00028199923
307 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021792475
401 Deep Unsupervised Cardinality Estimation 2020 VLDB 0.00019092557
513 NeuroCard: One Cardinality Estimator for All Tables 2021 VLDB 0.00017190574
635 Evaluating Top-k Selection Queries 1999 VLDB 0.00015527042
725 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014617251
790 FITing-Tree: A Data-aware Index Structure 2019 SIGMOD 0.0001401445
973 RankSQL: Query Algebra and Optimization for Relational Top-k Queries 2005 SIGMOD 0.00012874284
1,040 VGRAM: Improving Performance of Approximate Queries on String Collections Using Variable-Length Grams 2007 VLDB 0.00012466499
1,122 Cardinality Estimation in DBMS: A Comprehensive Benchmark Evaluation 2022 VLDB 0.0001209124
1,466 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.0001068941
1,551 Updatable Learned Index with Precise Positions 2021 VLDB 0.00010381398
1,701 Extending Autocompletion To Tolerate Errors 2009 SIGMOD 9.9707165e-05
1,967 IO-Top-k: Index-access Optimized Top-k Query Processing 2006 VLDB 9.3804693e-05
2,103 Cost-Based Variable-Length-Gram Selection for String Collections to Support Approximate Queries Efficiently 2008 SIGMOD 9.1621686e-05
2,782 Interaction between Record Matching and Data Repairing 2011 SIGMOD 8.1280264e-05
2,919 SNARF: A Learning-Enhanced Range Filter 2022 VLDB 7.9628657e-05
3,366 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.4748604e-05
3,545 Astrid: Accurate Selectivity Estimation for String Predicates using Deep Learning 2021 VLDB 7.3249967e-05
3,787 Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL 2024 VLDB 7.1249098e-05
3,809 Auto-WLM: Machine Learning Enhanced Workload Management in Amazon Redshift 2023 SIGMOD 7.1074195e-05
3,865 The Case for a Learned Sorting Algorithm 2020 SIGMOD 7.0621718e-05
4,035 Selectivity Estimation for Fuzzy String Predicates in Large Data Sets 2005 VLDB 6.9439151e-05
4,349 ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic Workloads 2024 VLDB 6.7504619e-05
4,757 Pattern Functional Dependencies for Data Cleaning 2020 VLDB 6.5220005e-05
4,780 Can Learned Models Replace Hash Functions? 2023 VLDB 6.5118885e-05
5,107 Stage: Query Execution Time Prediction in Amazon Redshift 2024 SIGMOD 6.3623786e-05
5,293 Magneto: Combining Small and Large Language Models for Schema Matching 2025 VLDB 6.279939e-05
5,412 Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction 2011 SIGMOD 6.2272563e-05
6,760 LPLM: A Neural Language Model for Cardinality Estimation of LIKE-Queries 2024 SIGMOD 5.7826781e-05
6,899 Human-in-the-loop Outlier Detection 2020 SIGMOD 5.7431609e-05
7,224 LITS: An Optimized Learned Index for Strings 2024 VLDB 5.6675134e-05
7,493 Human-in-the-loop Data Integration 2017 VLDB 5.6046905e-05
7,692 Efficient Top-k Algorithms for Approximate Substring Matching 2013 SIGMOD 5.5664611e-05
8,323 SageDB: An Instance-Optimized Data Analytics System 2022 VLDB 5.4539294e-05
8,986 One Seed, Two Birds: A Unified Learned Structure for Exact and Approximate Counting 2024 SIGMOD 5.3387783e-05
9,523 Repairing Data through Regular Expressions 2016 VLDB 5.2556545e-05
9,844 Cardinality Estimation of LIKE Predicate Queries using Deep Learning 2025 SIGMOD 5.2094602e-05
10,095 Stop Word and Related Problems in Web Interface Integration 2009 VLDB 5.1528643e-05
Previous Page 1 / 1 Next

Semantically Similar Papers