DBScholar

Back to papers

SQLStorm: Taking Database Benchmarking into the LLM Era

Summary: LLM-driven methodology generates SQLStorm: 18K+ realistic queries across 1–220 GB for only $15. Broad SQL coverage beyond TPC-H/DS/JOB enables compatibility testing, bug discovery, optimizer/cardinality-estimation research, and robust performance evaluation. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h12a91e8aa8604d4f
Venue
VLDB
Year
2025
Pagerank
6.9320915e-05
Overall Rank
3,900 | 73.79%
DOI
10.14778/3749646.3749683
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{schmidt_vldb25,
        title = {{SQLStorm: Taking Database Benchmarking into the LLM Era}},
        author = {Schmidt, Tobias and Leis, Viktor and Boncz, Peter and Neumann, Thomas},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {11},
        pages = {4144--4157},
        doi = {10.14778/3749646.3749683},
        url = {https://doi.org/10.14778/3749646.3749683},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 16 of 16 citing papers.

Rank Citing Paper Year Venue Pagerank
3,321 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 7.427185e-05
9,229 Finding Missed Optimizations in DBMSs through Unbalanced Short-Circuit Query Construction 2026 SIGMOD 5.2032182e-05
9,463 SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL Workloads 2026 SIGMOD 5.1712291e-05
10,107 Still Asking: How Good Are Query Optimizers, Really? 2025 VLDB 5.0765311e-05
10,248 Redbench: Workload Synthesis From Cloud Traces 2026 VLDB 5.0501823e-05
10,353 Rethinking Query Optimization for Multi-Agent Systems 2027 VLDB 4.9769913e-05
10,369 On the Vexing Difficulty of Evaluating IN Predicates 2026 CIDR 4.9769913e-05
10,427 Automated Discovery of Test Oracles for Database Management Systems Using LLMs 2026 SIGMOD 4.9769913e-05
10,448 Dialect-Agnostic SQL Parsing via LLM-Based Segmentation 2026 SIGMOD 4.9769913e-05
10,638 CorrBound: Cardinality Estimation Accounting for Inter- and Intra-relation Correlations 2026 SIGMOD 4.9769913e-05
10,723 Robust Predicate Transfer with Dynamic Execution 2026 VLDB 4.9769913e-05
10,761 Toward Drift-Aware Database Benchmarking 2026 VLDB 4.9769913e-05
10,842 ReSequel: Robust LLM-assisted Query Rewriting and Optimization using Templatization and Sampling 2026 VLDB 4.9769913e-05
10,895 Accelerating String-Heavy Queries with LLM Token Tables 2026 VLDB 4.9769913e-05
10,964 How Out-of-Bounds Are Your Cardinality Estimates? 2026 VLDB 4.9769913e-05
11,005 Persona-Conditioned Query Generation for Selectivity-Controlled Workloads 2026 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 31 of 31 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
15 How Good Are Query Optimizers, Really? 2016 VLDB 0.00061067652
85 Learned Cardinalities: Estimating Correlated Joins with Deep Learning 2019 CIDR 0.00035876108
329 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00020867521
361 Bao: Making Learned Query Optimization Practical 2021 SIGMOD 0.00020000855
386 Preventing Bad Plans by Bounding the Impact of Cardinality Estimation Errors 2009 VLDB 0.00019446558
981 Cardinality Estimation in DBMS: A Comprehensive Benchmark Evaluation 2022 VLDB 0.00012713454
1,245 DB-BERT: A Database Tuning Tool that "Reads the Manual" 2022 SIGMOD 0.0001136308
1,515 DSB: A Decision Support Benchmark for Workload-Driven and Traditional Database Systems 2021 VLDB 0.00010418766
1,659 D-Bot: Database Diagnosis System using Large Language Models 2024 VLDB 9.9625133e-05
1,834 Why TPC Is Not Enough: An Analysis of the Amazon Redshift Fleet 2024 VLDB 9.5349903e-05
1,949 Quantifying TPC-H Choke Points and Their Optimizations 2020 VLDB 9.3172855e-05
1,975 CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 Codex 2022 VLDB 9.2801545e-05
2,037 LLM-R^2: A Large Language Model Enhanced Rule-based Rewrite System for Boosting Query Efficiency 2025 VLDB 9.1494269e-05
2,227 GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian Optimization 2024 VLDB 8.8007923e-05
3,311 Cloud Analytics Benchmark 2023 VLDB 7.4364696e-05
3,321 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 7.427185e-05
3,362 Exact Cardinality Query Optimization for Optimizer Testing 2009 VLDB 7.3719912e-05
3,462 DIAMetrics: Benchmarking Query Engines at Scale 2020 VLDB 7.280037e-05
4,113 Automatic Database Configuration Debugging using Retrieval-Augmented Language Models 2025 SIGMOD 6.7964307e-05
4,346 Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs 2024 VLDB 6.646096e-05
4,384 R-Bot: An LLM-based Query Rewrite System 2025 VLDB 6.6232918e-05
4,727 LLM for Data Management 2024 VLDB 6.4461562e-05
4,945 Debunking the Myth of Join Ordering: Toward Robust SQL Analytics 2025 SIGMOD 6.3418058e-05
5,020 Analyzing the Impact of Cardinality Estimation on Execution Plans in Microsoft SQL Server 2023 VLDB 6.3096708e-05
5,895 Yannakakis+: Practical Acyclic Query Evaluation with Theoretical Guarantees 2025 SIGMOD 5.9534254e-05
5,910 Join Size Bounds using l_p-Norms on Degree Sequences 2024 PODS 5.9478947e-05
6,575 Can Large Language Models Be Query Optimizer for Relational Databases? 2026 SIGMOD 5.7428777e-05
6,646 Workload Insights From The Snowflake Data Cloud: What Do Production Analytic Queries Really Look Like? 2025 VLDB 5.7214206e-05
7,810 Parachute: Single-Pass Bi-Directional Information Passing 2025 VLDB 5.4477354e-05
8,984 DataLoom: Simplifying Data Loading with LLMs 2024 VLDB 5.2428922e-05
9,635 Low Rank Learning for Offline Query Optimization 2025 SIGMOD 5.1453041e-05
Previous Page 1 / 1 Next

Semantically Similar Papers