DBScholar

Back to papers

Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation

Summary: Systematic benchmark of prompt-engineering components (question representation, example selection/organization) and token-efficiency for LLM-based Text-to-SQL. Proposes DAIL-SQL (86.6% execution on Spider) and evaluates open-source LLMs with supervised fine-tuning, revealing accuracy/efficiency/cost trade-offs. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
13550
Venue
VLDB
Year
2024
Pagerank
0.00022468369
Overall Rank
279 | 98.09%
DOI
10.14778/3641204.3641221

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{gao_vldb24,
        title = {{Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation}},
        author = {Gao, Dawei and Wang, Haibin and Qian, Yichen and Li, Yaliang and Ding, Bolin and Sun, Xiuyu and Zhou, Jingren},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {5},
        pages = {1132--1145},
        doi = {10.14778/3641204.3641221},
        url = {https://doi.org/10.14778/3641204.3641221},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 55 citing papers.

Rank Citing Paper Year Venue Pagerank
756 CodeS: Towards Building Open-source Language Models for Text-to-SQL 2024 SIGMOD 0.0001431656
2,395 OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale 2025 VLDB 8.6355093e-05
2,710 OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency Alignment 2025 SIGMOD 8.2206448e-05
2,852 The Dawn of Natural Language to SQL: Are We Fully Ready? 2024 VLDB 8.0455088e-05
3,787 Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL 2024 VLDB 7.1249098e-05
4,251 FinSQL: Model-Agnostic LLMs-based Text-to-SQL Framework for Financial Analysis 2024 SIGMOD 6.8031106e-05
4,363 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 6.7423909e-05
4,944 Hybrid Querying Over Relational Databases and Large Language Models 2025 CIDR 6.432467e-05
5,323 SNAILS: Schema Naming Assessments for Improved LLM-Based SQL Inference 2025 SIGMOD 6.2655413e-05
6,211 GenEdit: Compounding Operators and Continuous Improvement to Tackle Text-to-SQL in the Enterprise 2025 CIDR 5.9425753e-05
6,267 Automated Validating and Fixing of Text-to-SQL Translation with Execution Consistency 2025 SIGMOD 5.9348282e-05
6,371 QUEST: Query Optimization in Unstructured Document Analysis 2025 VLDB 5.8962187e-05
6,936 Reliable Text-to-SQL with Adaptive Abstention 2025 SIGMOD 5.7349543e-05
7,332 Is Long Context All You Need? Leveraging LLM's Extended Context for NL2SQL 2025 VLDB 5.6433374e-05
7,865 Generating Succinct Descriptions of Database Schemata for Cost-Efficient Prompting of Large Language Models 2024 VLDB 5.5282752e-05
8,627 ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries 2025 VLDB 5.3969543e-05
8,725 mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs 2025 VLDB 5.3766157e-05
8,906 Unveiling Challenges for LLMs in Enterprise Data Engineering 2026 VLDB 5.3483178e-05
9,047 SQL-Factory: A Multi-Agent Framework for High-Quality and Large-Scale SQL Generation 2026 VLDB 5.3251649e-05
9,136 Sphinteract: Resolving Ambiguities in NL2SQL Through User Interaction 2025 VLDB 5.3166292e-05
9,305 The Power of Constraints in Natural Language to SQL Translation 2025 VLDB 5.289545e-05
9,464 Self-Enhancing Video Data Management System for Compositional Events with Large Language Models 2025 SIGMOD 5.2634238e-05
9,466 LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data 2025 VLDB 5.2634238e-05
9,870 Natural Language to SQL: State of the Art and Open Problems 2025 VLDB 5.2043672e-05
9,945 ParSEval: Plan-aware Test Database Generation for SQL Equivalence Evaluation 2025 VLDB 5.1915905e-05
10,109 QueryArtisan: Generating Data Manipulation Codes for Ad-hoc Analysis in Data Lakes 2025 VLDB 5.1347137e-05
10,141 Text-to-SQL Benchmarks are Broken: An In-Depth Analysis of Annotation Errors 2026 CIDR 5.093636e-05
10,216 DBugScribe: Automatic Database Bug Reproduction from Community Reports 2026 SIGMOD 5.093636e-05
10,217 DeepEye-SQL: A Software-Engineering-Inspired Text-to-SQL Framework 2026 SIGMOD 5.093636e-05
10,220 Dialect-Agnostic SQL Parsing via LLM-Based Segmentation 2026 SIGMOD 5.093636e-05
10,248 Generalized Entity Matching with Adaptivity via Large Language Models 2026 SIGMOD 5.093636e-05
10,276 OctoSelector: Efficient and Effective Batch-Aware Model Selection for Large Language Models 2026 SIGMOD 5.093636e-05
10,285 Reward-SQL: Boosting Text-to-SQL via Stepwise Execution-Aware Reasoning and Process-Supervised Rewards 2026 SIGMOD 5.093636e-05
10,340 AgentTune: An Agent-Based Large Language Model Framework for Database Knob Tuning 2026 SIGMOD 5.093636e-05
10,344 Are Your LLM-based Text-to-SQL Models Secure? Exploring SQL Injection via Backdoor Attacks 2026 SIGMOD 5.093636e-05
10,360 Drama: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries 2026 SIGMOD 5.093636e-05
10,389 PLForge: Enhancing Language Models for Natural Language to Procedural Extensions of SQL 2026 SIGMOD 5.093636e-05
10,398 Reliable Answers for Recurring Questions: Boosting Text-to-SQL Accuracy with Template Constrained Decoding 2026 SIGMOD 5.093636e-05
10,401 SEFRQO: A Self-Evolving Fine-Tuned RAG-Based Query Optimizer 2026 SIGMOD 5.093636e-05
10,444 DIVER: A Robust Text-to-SQL System with Dynamic Interactive Value Linking and Evidence Reasoning 2026 SIGMOD 5.093636e-05
10,483 PRISM: Navigating Cost–Accuracy Trade-offs for NL2SQL 2026 SIGMOD 5.093636e-05
10,510 NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions 2026 VLDB 5.093636e-05
10,530 SQL-Exchange: Transforming SQL Queries Across Domains 2026 VLDB 5.093636e-05
10,556 OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision 2026 VLDB 5.093636e-05
10,565 LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and Configuration 2026 VLDB 5.093636e-05
10,618 ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines 2026 VLDB 5.093636e-05
10,625 Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards 2026 VLDB 5.093636e-05
10,714 DataDazzle: Intelligent Data Exploration through Natural Language 2025 SIGMOD 5.093636e-05
10,774 PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models 2025 SIGMOD 5.093636e-05
10,856 Optimized Batch Prompting for Cost-effective LLMs 2025 VLDB 5.093636e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers