DBScholar

Back to papers

Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation

Summary: Systematic benchmark of prompt-engineering components (question representation, example selection/organization) and token-efficiency for LLM-based Text-to-SQL. Proposes DAIL-SQL (86.6% execution on Spider) and evaluates open-source LLMs with supervised fine-tuning, revealing accuracy/efficiency/cost trade-offs. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
h37b8ad61902e832d
Venue
VLDB
Year
2024
Pagerank
0.00026790979
Overall Rank
174 | 98.84%
DOI
10.14778/3641204.3641221

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{gao_vldb24,
        title = {{Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation}},
        author = {Gao, Dawei and Wang, Haibin and Qian, Yichen and Li, Yaliang and Ding, Bolin and Sun, Xiuyu and Zhou, Jingren},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {5},
        pages = {1132--1145},
        doi = {10.14778/3641204.3641221},
        url = {https://doi.org/10.14778/3641204.3641221},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 14 of 64 citing papers.

Rank Citing Paper Year Venue Pagerank
10,893 Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL Generation 2026 VLDB 4.9793485e-05
10,940 Bridging NL2SQL for CAN Signal Analytics at NIO: From Externalized Schemas to CTE Pipelines 2026 VLDB 4.9793485e-05
10,958 SpatialSQL: A Multi-Agent System for Interactive and Observable Spatial Text-to-SQL 2026 VLDB 4.9793485e-05
10,969 Text-to-SQL Evaluation Toolkit 2026 VLDB 4.9793485e-05
10,996 Persona-Conditioned Query Generation for Selectivity-Controlled Workloads 2026 VLDB 4.9793485e-05
11,020 ATLAS: Adaptive Text-to-SQL with Lifecycle-Aware Self-Maintaining Context 2026 VLDB 4.9793485e-05
11,071 Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards 2026 VLDB 4.9793485e-05
11,146 DataDazzle: Intelligent Data Exploration through Natural Language 2025 SIGMOD 4.9793485e-05
11,192 PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models 2025 SIGMOD 4.9793485e-05
11,284 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.9793485e-05
11,303 LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round Annotation 2025 VLDB 4.9793485e-05
11,328 Evoschema: Towards Text-To-Sql Robustness Against Schema Evolution 2025 VLDB 4.9793485e-05
11,364 CEDAR: A System for Cost-Efficient Data-Driven Claim Verification 2025 VLDB 4.9793485e-05
11,379 Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models 2025 VLDB 4.9793485e-05
Previous Page 2 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers