Accelerating String-Heavy Queries with LLM Token Tables
Summary: Repurposes GPT-4’s tokenizer as a globally shared string token table, letting joins and aggregations operate directly on encoded values. Umbra achieves up to 2× faster string-heavy queries, 1.65× lower storage/memory, and >6 GB/s decoding. (summarized by gpt-5.6-luna on Aug 28 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Tobias Schmidt (Technical University of Munich)
- 2. Nicolas Schmitt (Technical University of Munich)
- 3. Thomas Neumann (Technical University of Munich)
- 4. Andreas Kipf (University of Technology Nuremberg)
BibTeX Citation
@article{schmidt_vldb26,
title = {{Accelerating String-Heavy Queries with LLM Token Tables}},
author = {Schmidt, Tobias and Schmitt, Nicolas and Neumann, Thomas and Kipf, Andreas},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {11},
pages = {3635--3648},
doi = {10.14778/3836663.3836714},
url = {https://doi.org/10.14778/3836663.3836714},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 28 of 28 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,016 | Generating Succinct Descriptions of Database Schemata for Cost-Efficient Prompting of Large Language Models | 2024 | VLDB |
| 2 | 8,563 | Robust and Budget-Constrained Encoding Configurations for In-Memory Database Systems | 2022 | VLDB |
| 3 | 9,483 | Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point Compression | 2025 | VLDB |
| 4 | 5,489 | Compressed Representations of Conjunctive Query Results | 2018 | PODS |
| 5 | 9,863 | GPU Acceleration of SQL Analytics on Compressed Data | 2026 | VLDB |
| 6 | 910 | Dictionary-based Order-preserving String Compression for Main Memory Column Stores | 2009 | SIGMOD |
| 7 | 10,651 | Improving LZ4 for Effective Compression and Efficient Query | 2026 | SIGMOD |
| 8 | 6,737 | CompressDB: Enabling Efficient Compressed Data Direct Processing for Various Databases | 2022 | SIGMOD |
| 9 | 900 | Query Optimization In Compressed Database Systems | 2001 | SIGMOD |
| 10 | 13,576 | Waiting to Decompress: The Economics of LLM-Based Compression | 2026 | CIDR |