DBScholar

Back to papers

DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing

Summary: DocETL declaratively optimizes complex LLM document-processing pipelines via agent-generated logical rewrites (“rewrite directives”) and latency-aware plan evaluation. It trades single-call execution for decomposition and empirical plan search, improving accuracy 21–80% on four real tasks. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
14128
Venue
VLDB
Year
2025
Pagerank
0.00011095866
Overall Rank
1,343 | 90.79%
DOI
10.14778/3746405.3746426

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{shankar_vldb25,
        title = {{DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing}},
        author = {Shankar, Shreya and Chambers, Tristan and Shah, Tarak and Parameswaran, Aditya G. and Wu, Eugene},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {9},
        pages = {3035--3048},
        doi = {10.14778/3746405.3746426},
        url = {https://doi.org/10.14778/3746405.3746426},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 26 of 26 citing papers.

Rank Citing Paper Year Venue Pagerank
2,956 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 7.9203461e-05
4,045 Logical and Physical Optimizations for SQL Query Execution over Large Language Models 2025 SIGMOD 6.9394654e-05
4,081 Abacus: A Cost-Based Optimizer for Semantic Operator Systems 2026 VLDB 6.9165634e-05
6,778 VectraFlow: Integrating Vectors into Stream Processing 2025 CIDR 5.7772535e-05
7,568 Semantic Operators and Their Optimization: Enabling LLM-Based Data Processing with Accuracy Guarantees in LOTUS 2025 VLDB 5.5953379e-05
7,694 PalimpChat: Declarative and Interactive AI analytics 2025 SIGMOD 5.5661389e-05
8,829 Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees 2026 SIGMOD 5.3613784e-05
8,868 SemBench: A Benchmark for Semantic Query Processing Engines 2026 VLDB 5.3546848e-05
9,841 Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems 2025 VLDB 5.2101877e-05
10,116 Please Don't Kill My Vibe: Empowering Agents with Data Flow Control 2026 CIDR 5.093636e-05
10,120 KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration 2026 CIDR 5.093636e-05
10,132 Making Prompts First-Class Citizens for Adaptive LLM Pipelines 2026 CIDR 5.093636e-05
10,137 Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics 2026 CIDR 5.093636e-05
10,184 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models: [Experiments & Analysis] 2026 SIGMOD 5.093636e-05
10,199 Automated Discovery of Test Oracles for Database Management Systems Using LLMs 2026 SIGMOD 5.093636e-05
10,206 Bridging the Gap: Cardinality Estimation for Semantic Queries on Unstructured Data 2026 SIGMOD 5.093636e-05
10,286 ScaleDoc: Scaling LLM-based Predicates over Large Document Collections 2026 SIGMOD 5.093636e-05
10,360 Drama: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries 2026 SIGMOD 5.093636e-05
10,405 AixelAsk: A Stepwise-Guided Retrieval and Reasoning Framework for Large Table QA 2026 SIGMOD 5.093636e-05
10,414 Visual Template Inference for Data Extraction from Documents 2026 SIGMOD 5.093636e-05
10,433 Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization 2026 SIGMOD 5.093636e-05
10,483 PRISM: Navigating Cost–Accuracy Trade-offs for NL2SQL 2026 SIGMOD 5.093636e-05
10,504 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 5.093636e-05
10,572 Relational Deep Dive: Error-Aware Queries Over Unstructured Data 2026 VLDB 5.093636e-05
10,618 ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines 2026 VLDB 5.093636e-05
10,623 KEN: An Execution Engine for Unstructured Database Systems 2026 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
61 Maintaining Views Incrementally 1993 SIGMOD 0.00039026867
90 CrowdDB: Answering Queries with Crowdsourcing 2011 SIGMOD 0.00034951786
251 Crowdsourced Databases: Query Processing with People 2011 CIDR 0.00023261113
284 NoScope: Optimizing Neural Network Queries over Video at Scale 2017 VLDB 0.00022370521
290 An Overview of Query Optimization in Relational Systems 1998 PODS 0.0002227038
713 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00014672521
950 CAESURA: Language Models as Multi-Modal Query Planners 2024 CIDR 0.0001302491
1,245 Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing 2025 CIDR 0.00011507415
1,337 DB-BERT: A Database Tuning Tool that "Reads the Manual" 2022 SIGMOD 0.00011117488
2,242 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 8.8823802e-05
2,956 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 7.9203461e-05
3,073 DeepJoin: Joinable Table Discovery with Pre-trained Language Models 2023 VLDB 7.785842e-05
3,536 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.3297343e-05
4,050 Revisiting Prompt Engineering via Declarative Crowdsourcing 2024 CIDR 6.9368666e-05
4,081 Abacus: A Cost-Based Optimizer for Semantic Operator Systems 2026 VLDB 6.9165634e-05
4,621 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.6024412e-05
5,211 CDB: A Crowd-Powered Database System 2018 VLDB 6.3161892e-05
5,816 Observatory: Characterizing Embeddings of Relational Tables 2024 VLDB 6.0776454e-05
Previous Page 1 / 1 Next

Semantically Similar Papers