DBScholar

Back to papers

DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing

Summary: DocETL declaratively optimizes complex LLM document-processing pipelines via agent-generated logical rewrites (“rewrite directives”) and latency-aware plan evaluation. It trades single-call execution for decomposition and empirical plan search, improving accuracy 21–80% on four real tasks. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h3f14751160b5f392
Venue
VLDB
Year
2025
Pagerank
0.00014817539
Overall Rank
683 | 95.41%
DOI
10.14778/3746405.3746426

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{shankar_vldb25,
        title = {{DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing}},
        author = {Shankar, Shreya and Chambers, Tristan and Shah, Tarak and Parameswaran, Aditya G. and Wu, Eugene},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {9},
        pages = {3035--3048},
        doi = {10.14778/3746405.3746426},
        url = {https://doi.org/10.14778/3746405.3746426},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 40 of 40 citing papers.

Rank Citing Paper Year Venue Pagerank
2,341 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 8.6066183e-05
2,982 Semantic Operators and Their Optimization: Enabling LLM-Based Data Processing with Accuracy Guarantees in LOTUS 2025 VLDB 7.7845174e-05
3,126 Abacus: A Cost-Based Optimizer for Semantic Operator Systems 2026 VLDB 7.6185225e-05
3,743 Logical and Physical Optimizations for SQL Query Execution over Large Language Models 2025 SIGMOD 7.0586112e-05
5,501 SemBench: A Benchmark for Semantic Query Processing Engines 2026 VLDB 6.1027188e-05
6,789 VectraFlow: Integrating Vectors into Stream Processing 2025 CIDR 5.6813135e-05
6,886 Multi-Objective Agentic Rewrites for Unstructured Data Processing 2026 VLDB 5.6551172e-05
6,945 PalimpChat: Declarative and Interactive AI analytics 2025 SIGMOD 5.6378669e-05
7,885 Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems 2025 VLDB 5.433531e-05
8,291 ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines 2026 VLDB 5.3612871e-05
8,679 Relational Deep Dive: Error-Aware Queries Over Unstructured Data 2026 VLDB 5.2905577e-05
8,680 Unstructured Data Analysis Using LLMs: A Comprehensive Benchmark 2026 VLDB 5.2905577e-05
8,997 Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees 2026 SIGMOD 5.2410834e-05
10,144 Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics 2026 CIDR 5.0715586e-05
10,343 Please Don't Kill My Vibe: Empowering Agents with Data Flow Control 2026 CIDR 4.9793485e-05
10,346 KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration 2026 CIDR 4.9793485e-05
10,353 Making Prompts First-Class Citizens for Adaptive LLM Pipelines 2026 CIDR 4.9793485e-05
10,400 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models: [Experiments & Analysis] 2026 SIGMOD 4.9793485e-05
10,415 Automated Discovery of Test Oracles for Database Management Systems Using LLMs 2026 SIGMOD 4.9793485e-05
10,422 Bridging the Gap: Cardinality Estimation for Semantic Queries on Unstructured Data 2026 SIGMOD 4.9793485e-05
10,498 ScaleDoc: Scaling LLM-based Predicates over Large Document Collections 2026 SIGMOD 4.9793485e-05
10,560 Drama: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries 2026 SIGMOD 4.9793485e-05
10,598 AixelAsk: A Stepwise-Guided Retrieval and Reasoning Framework for Large Table QA 2026 SIGMOD 4.9793485e-05
10,605 Visual Template Inference for Data Extraction from Documents 2026 SIGMOD 4.9793485e-05
10,622 Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization 2026 SIGMOD 4.9793485e-05
10,670 PRISM: Navigating Cost–Accuracy Trade-offs for NL2SQL 2026 SIGMOD 4.9793485e-05
10,690 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 4.9793485e-05
10,791 BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents 2026 VLDB 4.9793485e-05
10,804 Document-to-Database: Extraction Meets Relational Semantics 2026 VLDB 4.9793485e-05
10,845 Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees 2026 VLDB 4.9793485e-05
10,858 Sema: A High-performance System for LLM-based Semantic Query Processing 2026 VLDB 4.9793485e-05
10,957 Credo: Declarative Control of LLM Pipelines via Beliefs and Policies 2026 VLDB 4.9793485e-05
10,960 QuWARTS: Query Workload Aware Relational Table Synthesis from Unstructured Text 2026 VLDB 4.9793485e-05
10,978 CADENZA in Action: Breaking the Monolith with Intent-Dependent Plan Spaces for Semantic Queries 2026 VLDB 4.9793485e-05
10,980 Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries 2026 VLDB 4.9793485e-05
10,986 Demonstrating Samsara: A Super-Optimizer for Multimodal Stream Processing 2026 VLDB 4.9793485e-05
11,008 iPDB: SQL with ML and LLM Predicates (Towards a Database Engine for AI) 2026 VLDB 4.9793485e-05
11,024 Bridging LLMs and Database Systems: A Deep Dive into Enhanced Relational Operators 2026 VLDB 4.9793485e-05
11,034 Data Agents: Rethinking Data Systems in the AI Agent Era 2026 VLDB 4.9793485e-05
11,069 KEN: An Execution Engine for Unstructured Database Systems 2026 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
62 Maintaining Views Incrementally 1993 SIGMOD 0.00039045511
92 CrowdDB: Answering Queries with Crowdsourcing 2011 SIGMOD 0.00034672523
257 Crowdsourced Databases: Query Processing with People 2011 CIDR 0.00022962347
271 NoScope: Optimizing Neural Network Queries over Video at Scale 2017 VLDB 0.00022560564
272 An Overview of Query Optimization in Relational Systems 1998 PODS 0.00022509573
501 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017267905
669 CAESURA: Language Models as Multi-Modal Query Planners 2024 CIDR 0.0001495987
748 Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing 2025 CIDR 0.00014281926
1,250 DB-BERT: A Database Tuning Tool that "Reads the Manual" 2022 SIGMOD 0.00011339256
1,932 CHORUS: Foundation Models for Unified Data Discovery and Exploration 2024 VLDB 9.34643e-05
2,341 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 8.6066183e-05
2,382 DeepJoin: Joinable Table Discovery with Pre-trained Language Models 2023 VLDB 8.5458532e-05
3,126 Abacus: A Cost-Based Optimizer for Semantic Operator Systems 2026 VLDB 7.6185225e-05
3,250 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.4938546e-05
3,695 Revisiting Prompt Engineering via Declarative Crowdsourcing 2024 CIDR 7.092445e-05
4,675 spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines 2024 VLDB 6.474947e-05
5,061 CDB: A Crowd-Powered Database System 2018 VLDB 6.2912543e-05
5,332 Observatory: Characterizing Embeddings of Relational Tables 2024 VLDB 6.1766186e-05
Previous Page 1 / 1 Next

Semantically Similar Papers