ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL Systems
Summary: ScienceBenchmark introduces the first domain-expert-validated NL-to-SQL benchmark over three complex scientific databases. It combines scarce human NL/SQL pairs with GPT-3 synthetic data, exposing severe weaknesses of Spider-trained systems under realistic schema and domain complexity. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yi Zhang (Zurich University of Applied Sciences)
- 2. Jan Deriu (Zurich University of Applied Sciences)
- 3. George Katsogiannis-Meimarakis (Athena Research Center)
- 4. Catherine Kosten (Zurich University of Applied Sciences)
- 5. Georgia Koutrika (Athena Research Center)
- 6. Kurt Stockinger (Zurich University of Applied Sciences)
BibTeX Citation
@article{zhang_vldb24,
title = {{ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL Systems}},
author = {Zhang, Yi and Deriu, Jan and Katsogiannis-Meimarakis, George and Kosten, Catherine and Koutrika, Georgia and Stockinger, Kurt},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {4},
pages = {685--698},
doi = {10.14778/3636218.3636225},
url = {https://doi.org/10.14778/3636218.3636225},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 179 | Constructing an Interactive Natural Language Interface for Relational Databases | 2015 | VLDB | 0.00026838277 |
| 459 | ATHENA: An Ontology-Driven System for Natural Language Querying over Relational Data Stores | 2016 | VLDB | 0.00018101318 |
| 555 | NaLIR: An Interactive Natural Language Interface for Querying Relational Databases | 2014 | SIGMOD | 0.00016568053 |
| 1,214 | SODA: Generating SQL for Business Users | 2012 | VLDB | 0.0001163751 |
| 1,719 | The SDSS SkyServer - Public Access to the Sloan Digital Sky Survey Data | 2002 | SIGMOD | 9.9268286e-05 |
| 2,079 | DBPal: A Fully Pluggable NL2SQL Training Pipeline | 2020 | SIGMOD | 9.2060425e-05 |
| 3,354 | MT-Teql: Evaluating and Augmenting Neural NLIDB on Real-world Linguistic and Schema Variations | 2022 | VLDB | 7.4918609e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,870 | Natural Language to SQL: State of the Art and Open Problems | 2025 | VLDB |
| 2 | 10,122 | BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation | 2026 | CIDR |
| 3 | 5,323 | SNAILS: Schema Naming Assessments for Improved LLM-Based SQL Inference | 2025 | SIGMOD |
| 4 | 4,809 | Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks | 2021 | SIGMOD |
| 5 | 1,444 | CatSQL: Towards Real World Natural Language to SQL Applications | 2023 | VLDB |
| 6 | 2,602 | NL2SQL is a solved problem... Not! | 2024 | CIDR |
| 7 | 279 | Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation | 2024 | VLDB |
| 8 | 2,852 | The Dawn of Natural Language to SQL: Are We Fully Ready? | 2024 | VLDB |
| 9 | 865 | Natural language to SQL: Where are we today? | 2020 | VLDB |
| 10 | 10,510 | NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions | 2026 | VLDB |