Unstructured Data Analysis Using LLMs: A Comprehensive Benchmark
Summary: Bench-U is a comprehensive benchmark for LLM-based unstructured data analysis, pairing six diverse datasets with manually curated relational ground truth and rich analytical workloads. It enables interface-agnostic evaluation and exposes tradeoffs across query interfaces, optimization, operators, and processing. (summarized by gpt-5.6-luna on Aug 17 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Qiyan Deng (Beijing Institute of Technology)
- 2. Jianhui Li (Beijing Institute of Technology)
- 3. Chengliang Chai (Beijing Institute of Technology)
- 4. Ye Yuan (Beijing Institute of Technology)
- 5. Jinqi Liu (Beijing Institute of Technology)
- 6. Junzhi She (Beijing Institute of Technology)
- 7. Kaisen Jin (Beijing Institute of Technology)
- 8. Zhaoze Sun (Beijing Institute of Technology)
- 9. Yuhao Deng (Beijing Institute of Technology)
- 10. Jia Yuan (University of Arizona)
- 11. Yuping Wang (Beijing Institute of Technology)
- 12. Xu Zhou (Hunan University)
- 13. Guoren Wang (Beijing Institute of Technology)
- 14. Lei Cao (University of Arizona)
BibTeX Citation
@article{deng_vldb26,
title = {{Unstructured Data Analysis Using LLMs: A Comprehensive Benchmark}},
author = {Deng, Qiyan and Li, Jianhui and Chai, Chengliang and Yuan, Ye and Liu, Jinqi and She, Junzhi and Jin, Kaisen and Sun, Zhaoze and Deng, Yuhao and Yuan, Jia and Wang, Yuping and Zhou, Xu and Wang, Guoren and Cao, Lei},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {9},
pages = {2398--2410},
doi = {10.14778/3819518.3819559},
url = {https://doi.org/10.14778/3819518.3819559},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,960 | QuWARTS: Query Workload Aware Relational Table Synthesis from Unstructured Text | 2026 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 501 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB | 0.00017267905 |
| 669 | CAESURA: Language Models as Multi-Modal Query Planners | 2024 | CIDR | 0.0001495987 |
| 683 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB | 0.00014817539 |
| 1,392 | Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes | 2023 | CIDR | 0.00010807936 |
| 1,846 | From Natural Language Processing to Neural Databases | 2021 | VLDB | 9.513089e-05 |
| 2,450 | ThalamusDB: Approximate Query Processing on Multi-Modal Data | 2024 | SIGMOD | 8.4474092e-05 |
| 4,731 | QUEST: Query Optimization in Unstructured Document Analysis | 2025 | VLDB | 6.4462032e-05 |
| 6,869 | Doctopus: Budget-aware Structural Table Extraction from Unstructured Documents | 2025 | VLDB | 5.659905e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,155 | An In-Depth Benchmarking of Text-to-SQL Systems | 2021 | SIGMOD |
| 2 | 10,620 | AutoDDG: Automated Dataset Description Generation using Large Language Models | 2026 | SIGMOD |
| 3 | 8,679 | Relational Deep Dive: Error-Aware Queries Over Unstructured Data | 2026 | VLDB |
| 4 | 5,501 | SemBench: A Benchmark for Semantic Query Processing Engines | 2026 | VLDB |
| 5 | 4,731 | QUEST: Query Optimization in Unstructured Document Analysis | 2025 | VLDB |
| 6 | 174 | Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation | 2024 | VLDB |
| 7 | 10,695 | NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions | 2026 | VLDB |
| 8 | 6,379 | Unify: A System For Unstructured Data Analytics | 2025 | VLDB |
| 9 | 8,632 | DocDB: A Database for Unstructured Document Analysis | 2025 | VLDB |
| 10 | 11,529 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |