Text-to-SQL Evaluation Toolkit
Summary: Text2SQL-Eval is an open-source, modular toolkit unifying 12+ metrics—from execution and syntactic equivalence to LLM judging—with inference, profiling, error analysis, and live dashboard comparison. It enables rigorous, diagnosis-oriented evaluation across public and enterprise benchmarks. (summarized by gpt-5.6-luna on Aug 28 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Oktie Hassanzadeh (IBM)
- 2. Yotam Perlitz (IBM)
- 3. Nhan Pham (IBM)
- 4. Tanvi Kaple (IBM)
- 5. Karolina Źróbek (IBM)
- 6. Long Vu (IBM)
- 7. Michael Glass (IBM)
- 8. Dharmashankar Subramanian (IBM)
- 9. Mohammadreza Pourreza (University of Alberta)
- 10. Davood Rafiei (University of Alberta)
BibTeX Citation
@article{hassanzadeh_vldb26,
title = {{Text-to-SQL Evaluation Toolkit}},
author = {Hassanzadeh, Oktie and Perlitz, Yotam and Pham, Nhan and Kaple, Tanvi and Źróbek, Karolina and Vu, Long and Glass, Michael and Subramanian, Dharmashankar and Pourreza, Mohammadreza and Rafiei, Davood},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {12},
pages = {4582--4585},
doi = {10.14778/3827998.3828071},
url = {https://doi.org/10.14778/3827998.3828071},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 1 of 1 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 174 | Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation | 2024 | VLDB | 0.00026790979 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,893 | Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL Generation | 2026 | VLDB |
| 2 | 13,607 | ParSEval: Interactive Counterexample-driven Evaluation for Text-to-SQL | 2026 | VLDB |
| 3 | 9,210 | Natural Language to SQL: State of the Art and Open Problems | 2025 | VLDB |
| 4 | 10,738 | OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision | 2026 | VLDB |
| 5 | 2,348 | The Dawn of Natural Language to SQL: Are We Fully Ready? | 2024 | VLDB |
| 6 | 11,160 | RTS+: Reliable Text to SQL | 2025 | SIGMOD |
| 7 | 776 | Natural language to SQL: Where are we today? | 2020 | VLDB |
| 8 | 174 | Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation | 2024 | VLDB |
| 9 | 5,155 | An In-Depth Benchmarking of Text-to-SQL Systems | 2021 | SIGMOD |
| 10 | 10,695 | NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions | 2026 | VLDB |