Hybrid Querying Over Relational Databases and Large Language Models
Summary: Presents SWAN: the first cross-domain benchmark of 120 beyond-database questions over four real-world relational schemas for hybrid DB+LLM querying. Proposes schema-expansion and UDF-based integration, evaluates GPT‑4 Turbo (≤40% exec accuracy, 48.2% factuality) and exposes optimization needs and accuracy/factuality gaps. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Fuheng Zhao (University of California Santa Barbara)
- 2. Divyakant Agrawal (University of California Santa Barbara)
- 3. Amr El Abbadi (University of California Santa Barbara)
BibTeX Citation
@inproceedings{zhao_cidr25,
address = {Amsterdam, Netherlands},
series = {{CIDR} '25},
title = {{Hybrid Querying Over Relational Databases and Large Language Models}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Zhao, Fuheng and Agrawal, Divyakant and Abbadi, Amr El},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,045 | Logical and Physical Optimizations for SQL Query Execution over Large Language Models | 2025 | SIGMOD | 6.9394654e-05 |
| 9,136 | Sphinteract: Resolving Ambiguities in NL2SQL Through User Interaction | 2025 | VLDB | 5.3166292e-05 |
| 10,733 | ScaleLLM: A Technique for Scalable LLM-augmented Data Systems | 2025 | SIGMOD | 5.093636e-05 |
| 10,737 | SwellDB: Dynamic Query-Driven Table Generation with Large Language Models | 2025 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 90 | CrowdDB: Answering Queries with Crowdsourcing | 2011 | SIGMOD | 0.00034951786 |
| 103 | DuckDB: an Embeddable Analytical Database | 2019 | SIGMOD | 0.00034161428 |
| 250 | Answering Queries using Humans, Algorithms and Databases | 2011 | CIDR | 0.00023261164 |
| 251 | Crowdsourced Databases: Query Processing with People | 2011 | CIDR | 0.00023261113 |
| 279 | Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation | 2024 | VLDB | 0.00022468369 |
| 420 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00018789852 |
| 2,117 | Obtaining Complete Answers from Incomplete Databases | 1996 | VLDB | 9.1434045e-05 |
| 2,409 | Deco: A System for Declarative Crowdsourcing | 2012 | VLDB | 8.6145078e-05 |
| 2,602 | NL2SQL is a solved problem... Not! | 2024 | CIDR | 8.3535452e-05 |
| 4,050 | Revisiting Prompt Engineering via Declarative Crowdsourcing | 2024 | CIDR | 6.9368666e-05 |
| 8,826 | What Should A Database Know? | 1988 | PODS | 5.3616554e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,787 | Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL | 2024 | VLDB |
| 2 | 9,542 | Database as Runtime: Compiling LLMs to SQL for In-database Model Serving | 2025 | SIGMOD |
| 3 | 10,556 | OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision | 2026 | VLDB |
| 4 | 10,510 | NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions | 2026 | VLDB |
| 5 | 6,259 | R-Bot: An LLM-based Query Rewrite System | 2025 | VLDB |
| 6 | 7,809 | Can Large Language Models Be Query Optimizer for Relational Databases? | 2026 | SIGMOD |
| 7 | 11,120 | Welding Natural Language Queries to Analytics IRs with LLMs | 2024 | CIDR |
| 8 | 10,139 | Leveraging Query Optimizers to Verify the Soundness of LLM-based Query Rewrites for Real-World Workloads, and More! | 2026 | CIDR |
| 9 | 4,045 | Logical and Physical Optimizations for SQL Query Execution over Large Language Models | 2025 | SIGMOD |
| 10 | 6,936 | Reliable Text-to-SQL with Adaptive Abstention | 2025 | SIGMOD |