DBScholar

Back to papers

OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale

Summary: Scalable synthesis framework producing SynSQL‑2.5M: 2.5M text-to-SQL samples across ~16k synthetic databases, each with DB, SQL, NL question, and chain-of-thought, addressing data scarcity and reliance on closed-source prompting. Trains OmniSQL (7B/14B/32B), open-source, matching or surpassing larger closed/open LLMs (e.g., GPT‑4o, DeepSeek‑V3). (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
14266
Venue
VLDB
Year
2025
Pagerank
8.6355093e-05
Overall Rank
2,395 | 83.57%
DOI
10.14778/3749646.3749723

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{li_vldb25,
        title = {{OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale}},
        author = {Li, Haoyang and Wu, Shang and Zhang, Xiaokang and Huang, Xinmei and Zhang, Jing and Jiang, Fuxin and Wang, Shuai and Zhang, Tieying and Chen, Jianjun and Shi, Rui and Chen, Hong and Li, Cuiping},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {11},
        pages = {4695--4709},
        doi = {10.14778/3749646.3749723},
        url = {https://doi.org/10.14778/3749646.3749723},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 12 of 12 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 8 of 8 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers