ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
Summary: ELT-Bench is an end-to-end benchmark for AI agents orchestrating realistic ELT workflows across databases, tools, code, and SQL, spanning 100 pipelines, 835 sources, and 203 models. State-of-the-art agents generate only 11.3% of models, exposing substantial automation gaps. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Tengjun Jin (University of Illinois Urbana-Champaign)
- 2. Yuxuan Zhu (University of Illinois Urbana-Champaign)
- 3. Daniel Kang (University of Illinois Urbana-Champaign)
BibTeX Citation
@article{jin_vldb26,
title = {{ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines}},
author = {Jin, Tengjun and Zhu, Yuxuan and Kang, Daniel},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {2},
pages = {84--98},
doi = {10.14778/3773749.3773750},
url = {https://doi.org/10.14778/3773749.3773750},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,648 | BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents | 2026 | VLDB | 5.0723324e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 67 | The Snowflake Elastic Data Warehouse | 2016 | SIGMOD | 0.00038495191 |
| 266 | Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation | 2024 | VLDB | 0.00022795833 |
| 679 | CodeS: Towards Building Open-source Language Models for Text-to-SQL | 2024 | SIGMOD | 0.00015005329 |
| 1,178 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB | 0.00011775585 |
| 2,811 | Text2SQL is Not Enough: Unifying AI and Databases with TAG | 2025 | CIDR | 8.0782961e-05 |
| 3,586 | Cloud Analytics Benchmark | 2023 | VLDB | 7.2756541e-05 |
| 4,877 | TPC-DI: The First Industry Benchmark for Data Integration | 2014 | VLDB | 6.4488377e-05 |
| 5,044 | TPCx-AI - An Industry Standard Benchmark for Artificial Intelligence and Machine Learning Systems | 2023 | VLDB | 6.3714012e-05 |
| 5,496 | Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System | 2025 | SIGMOD | 6.1825966e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,183 | Welding Natural Language Queries to Analytics IRs with LLMs | 2024 | CIDR |
| 2 | 11,090 | A Demonstration of QueryArtisan: Real-Time Data Lake Analysis via Dynamically Generated Data Manipulation Code | 2025 | VLDB |
| 3 | 5,517 | ELEET: Efficient Learned Query Execution over Text and Tables | 2024 | VLDB |
| 4 | 8,937 | Unveiling Challenges for LLMs in Enterprise Data Engineering | 2026 | VLDB |
| 5 | 13,356 | Demonstrating CatDB: LLM-based Generation of Data-centric ML Pipelines | 2025 | SIGMOD |
| 6 | 1,996 | ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL Systems | 2024 | VLDB |
| 7 | 10,164 | BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation | 2026 | CIDR |
| 8 | 266 | Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation | 2024 | VLDB |
| 9 | 1,178 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB |
| 10 | 10,548 | NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions | 2026 | VLDB |