DBScholar

Back to papers

SQLStorm: Taking Database Benchmarking into the LLM Era

Summary: LLM-driven methodology generates SQLStorm: 18K+ realistic queries across 1–220 GB for only $15. Broad SQL coverage beyond TPC-H/DS/JOB enables compatibility testing, bug discovery, optimizer/cardinality-estimation research, and robust performance evaluation. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h12a91e8aa8604d4f
Venue
VLDB
Year
2025
Pagerank
6.7885553e-05
Overall Rank
4,132 | 72.22%
DOI
10.14778/3749646.3749683

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{schmidt_vldb25,
        title = {{SQLStorm: Taking Database Benchmarking into the LLM Era}},
        author = {Schmidt, Tobias and Leis, Viktor and Boncz, Peter and Neumann, Thomas},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {11},
        pages = {4144--4157},
        doi = {10.14778/3749646.3749683},
        url = {https://doi.org/10.14778/3749646.3749683},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 15 of 15 citing papers.

Rank Citing Paper Year Venue Pagerank
3,336 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 7.4137763e-05
9,219 Finding Missed Optimizations in DBMSs through Unbalanced Short-Circuit Query Construction 2026 SIGMOD 5.2056825e-05
9,454 SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL Workloads 2026 SIGMOD 5.1736782e-05
10,103 Still Asking: How Good Are Query Optimizers, Really? 2025 VLDB 5.0789354e-05
10,242 Redbench: Workload Synthesis From Cloud Traces 2026 VLDB 5.0525742e-05
10,357 On the Vexing Difficulty of Evaluating IN Predicates 2026 CIDR 4.9793485e-05
10,415 Automated Discovery of Test Oracles for Database Management Systems Using LLMs 2026 SIGMOD 4.9793485e-05
10,436 Dialect-Agnostic SQL Parsing via LLM-Based Segmentation 2026 SIGMOD 4.9793485e-05
10,627 CorrBound: Cardinality Estimation Accounting for Inter- and Intra-relation Correlations 2026 SIGMOD 4.9793485e-05
10,713 Robust Predicate Transfer with Dynamic Execution 2026 VLDB 4.9793485e-05
10,751 Toward Drift-Aware Database Benchmarking 2026 VLDB 4.9793485e-05
10,832 ReSequel: Robust LLM-assisted Query Rewriting and Optimization using Templatization and Sampling 2026 VLDB 4.9793485e-05
10,886 Accelerating String-Heavy Queries with LLM Token Tables 2026 VLDB 4.9793485e-05
10,955 How Out-of-Bounds Are Your Cardinality Estimates? 2026 VLDB 4.9793485e-05
10,996 Persona-Conditioned Query Generation for Selectivity-Controlled Workloads 2026 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 31 of 31 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
15 How Good Are Query Optimizers, Really? 2016 VLDB 0.00061066921
85 Learned Cardinalities: Estimating Correlated Joins with Deep Learning 2019 CIDR 0.00035864347
329 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00020858443
362 Bao: Making Learned Query Optimization Practical 2021 SIGMOD 0.00019989474
386 Preventing Bad Plans by Bounding the Impact of Cardinality Estimation Errors 2009 VLDB 0.00019444411
982 Cardinality Estimation in DBMS: A Comprehensive Benchmark Evaluation 2022 VLDB 0.00012714044
1,250 DB-BERT: A Database Tuning Tool that "Reads the Manual" 2022 SIGMOD 0.00011339256
1,515 DSB: A Decision Support Benchmark for Workload-Driven and Traditional Database Systems 2021 VLDB 0.00010417728
1,658 D-Bot: Database Diagnosis System using Large Language Models 2024 VLDB 9.9642078e-05
1,839 Why TPC Is Not Enough: An Analysis of the Amazon Redshift Fleet 2024 VLDB 9.5304799e-05
1,952 Quantifying TPC-H Choke Points and Their Optimizations 2020 VLDB 9.3189525e-05
1,975 CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 Codex 2022 VLDB 9.2807031e-05
2,039 LLM-R^2: A Large Language Model Enhanced Rule-based Rewrite System for Boosting Query Efficiency 2025 VLDB 9.1493268e-05
2,231 GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian Optimization 2024 VLDB 8.7982985e-05
3,316 Cloud Analytics Benchmark 2023 VLDB 7.4373161e-05
3,336 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 7.4137763e-05
3,367 Exact Cardinality Query Optimization for Optimizer Testing 2009 VLDB 7.3719456e-05
3,465 DIAMetrics: Benchmarking Query Engines at Scale 2020 VLDB 7.2808932e-05
4,114 Automatic Database Configuration Debugging using Retrieval-Augmented Language Models 2025 SIGMOD 6.7971182e-05
4,385 R-Bot: An LLM-based Query Rewrite System 2025 VLDB 6.6235293e-05
4,718 Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs 2024 VLDB 6.4551989e-05
4,950 Debunking the Myth of Join Ordering: Toward Robust SQL Analytics 2025 SIGMOD 6.3421691e-05
5,022 Analyzing the Impact of Cardinality Estimation on Execution Plans in Microsoft SQL Server 2023 VLDB 6.3100988e-05
5,149 LLM for Data Management 2024 VLDB 6.2551644e-05
5,902 Yannakakis+: Practical Acyclic Query Evaluation with Theoretical Guarantees 2025 SIGMOD 5.9536872e-05
5,918 Join Size Bounds using l_p-Norms on Degree Sequences 2024 PODS 5.9481539e-05
6,586 Can Large Language Models Be Query Optimizer for Relational Databases? 2026 SIGMOD 5.7430662e-05
6,648 Workload Insights From The Snowflake Data Cloud: What Do Production Analytic Queries Really Look Like? 2025 VLDB 5.7215989e-05
7,817 Parachute: Single-Pass Bi-Directional Information Passing 2025 VLDB 5.4477841e-05
8,986 DataLoom: Simplifying Data Loading with LLMs 2024 VLDB 5.2428439e-05
9,670 Low Rank Learning for Offline Query Optimization 2025 SIGMOD 5.1452097e-05
Previous Page 1 / 1 Next

Semantically Similar Papers