DBScholar

Back to papers

DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing

Summary: DocETL declaratively optimizes complex LLM document-processing pipelines via agent-generated logical rewrites (“rewrite directives”) and latency-aware plan evaluation. It trades single-call execution for decomposition and empirical plan search, improving accuracy 21–80% on four real tasks. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h3f14751160b5f392
Venue
VLDB
Year
2025
Pagerank
0.00015058738
Overall Rank
658 | 95.58%
DOI
10.14778/3746405.3746426
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{shankar_vldb25,
        title = {{DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing}},
        author = {Shankar, Shreya and Chambers, Tristan and Shah, Tarak and Parameswaran, Aditya G. and Wu, Eugene},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {9},
        pages = {3035--3048},
        doi = {10.14778/3746405.3746426},
        url = {https://doi.org/10.14778/3746405.3746426},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 42 of 42 citing papers.

Rank Citing Paper Year Venue Pagerank
2,273 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 8.7093949e-05
2,909 Semantic Operators and Their Optimization: Enabling LLM-Based Data Processing with Accuracy Guarantees in LOTUS 2025 VLDB 7.8705905e-05
3,049 Abacus: A Cost-Based Optimizer for Semantic Operator Systems 2026 VLDB 7.7087759e-05
3,717 Logical and Physical Optimizations for SQL Query Execution over Large Language Models 2025 SIGMOD 7.0712441e-05
5,309 SemBench: A Benchmark for Semantic Query Processing Engines 2026 VLDB 6.1841176e-05
6,607 Multi-Objective Agentic Rewrites for Unstructured Data Processing 2026 VLDB 5.73539e-05
6,777 VectraFlow: Integrating Vectors into Stream Processing 2025 CIDR 5.6833587e-05
6,939 PalimpChat: Declarative and Interactive AI analytics 2025 SIGMOD 5.6372718e-05
7,865 Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems 2025 VLDB 5.4361737e-05
8,297 ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines 2026 VLDB 5.3587491e-05
8,526 Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees 2026 SIGMOD 5.3215522e-05
8,687 Relational Deep Dive: Error-Aware Queries Over Unstructured Data 2026 VLDB 5.2880532e-05
8,688 Unstructured Data Analysis Using LLMs: A Comprehensive Benchmark 2026 VLDB 5.2880532e-05
10,148 Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics 2026 CIDR 5.0691578e-05
10,198 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 5.0599411e-05
10,199 Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees 2026 VLDB 5.0599411e-05
10,351 Evergreen: Efficient Claim Verification for Semantic Aggregates 2027 VLDB 4.9769913e-05
10,353 Rethinking Query Optimization for Multi-Agent Systems 2027 VLDB 4.9769913e-05
10,355 Please Don't Kill My Vibe: Empowering Agents with Data Flow Control 2026 CIDR 4.9769913e-05
10,358 KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration 2026 CIDR 4.9769913e-05
10,365 Making Prompts First-Class Citizens for Adaptive LLM Pipelines 2026 CIDR 4.9769913e-05
10,412 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models: [Experiments & Analysis] 2026 SIGMOD 4.9769913e-05
10,427 Automated Discovery of Test Oracles for Database Management Systems Using LLMs 2026 SIGMOD 4.9769913e-05
10,434 Bridging the Gap: Cardinality Estimation for Semantic Queries on Unstructured Data 2026 SIGMOD 4.9769913e-05
10,509 ScaleDoc: Scaling LLM-based Predicates over Large Document Collections 2026 SIGMOD 4.9769913e-05
10,571 Drama: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries 2026 SIGMOD 4.9769913e-05
10,609 AixelAsk: A Stepwise-Guided Retrieval and Reasoning Framework for Large Table QA 2026 SIGMOD 4.9769913e-05
10,616 Visual Template Inference for Data Extraction from Documents 2026 SIGMOD 4.9769913e-05
10,633 Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization 2026 SIGMOD 4.9769913e-05
10,681 PRISM: Navigating Cost–Accuracy Trade-offs for NL2SQL 2026 SIGMOD 4.9769913e-05
10,801 BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents 2026 VLDB 4.9769913e-05
10,814 Document-to-Database: Extraction Meets Relational Semantics 2026 VLDB 4.9769913e-05
10,867 Sema: A High-performance System for LLM-based Semantic Query Processing 2026 VLDB 4.9769913e-05
10,966 Credo: Declarative Control of LLM Pipelines via Beliefs and Policies 2026 VLDB 4.9769913e-05
10,969 QuWARTS: Query Workload Aware Relational Table Synthesis from Unstructured Text 2026 VLDB 4.9769913e-05
10,987 CADENZA in Action: Breaking the Monolith with Intent-Dependent Plan Spaces for Semantic Queries 2026 VLDB 4.9769913e-05
10,989 Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries 2026 VLDB 4.9769913e-05
10,995 Demonstrating Samsara: A Super-Optimizer for Multimodal Stream Processing 2026 VLDB 4.9769913e-05
11,017 iPDB: SQL with ML and LLM Predicates (Towards a Database Engine for AI) 2026 VLDB 4.9769913e-05
11,033 Bridging LLMs and Database Systems: A Deep Dive into Enhanced Relational Operators 2026 VLDB 4.9769913e-05
11,043 Data Agents: Rethinking Data Systems in the AI Agent Era 2026 VLDB 4.9769913e-05
11,078 KEN: An Execution Engine for Unstructured Database Systems 2026 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
62 Maintaining Views Incrementally 1993 SIGMOD 0.00039040346
92 CrowdDB: Answering Queries with Crowdsourcing 2011 SIGMOD 0.00034670735
257 Crowdsourced Databases: Query Processing with People 2011 CIDR 0.00022958404
271 NoScope: Optimizing Neural Network Queries over Video at Scale 2017 VLDB 0.0002256866
272 An Overview of Query Optimization in Relational Systems 1998 PODS 0.0002251422
496 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017318538
655 CAESURA: Language Models as Multi-Modal Query Planners 2024 CIDR 0.00015090153
717 Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing 2025 CIDR 0.00014528323
1,245 DB-BERT: A Database Tuning Tool that "Reads the Manual" 2022 SIGMOD 0.0001136308
1,929 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 9.3551286e-05
2,273 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 8.7093949e-05
2,378 DeepJoin: Joinable Table Discovery with Pre-trained Language Models 2023 VLDB 8.5495442e-05
3,049 Abacus: A Cost-Based Optimizer for Semantic Operator Systems 2026 VLDB 7.7087759e-05
3,236 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.4996147e-05
3,668 Revisiting Prompt Engineering via Declarative Crowdsourcing 2024 CIDR 7.1144218e-05
4,678 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.4720352e-05
5,050 CDB: A Crowd-Powered Database System 2018 VLDB 6.2955765e-05
5,322 Observatory: Characterizing Embeddings of Relational Tables 2024 VLDB 6.1809951e-05
Previous Page 1 / 1 Next

Semantically Similar Papers