Back to papers
Task Cascades for Efficient Unstructured Data Processing
Summary: Task cascades generalize model cascades for LLM-based document processing by varying not only the model, but also the queried span and even the operation, exploiting simpler correlated sub-tasks and partial evidence. An iterative optimizer plus statistical accuracy guarantees yields 36% lower cost than standard cascades at 90% target accuracy.
(summarized by gpt-5.4-mini on Apr 11 2026)
Paper ID
7718
Venue
SIGMOD
Year
2026
Pagerank
5.093636e-05
Overall Rank
10,504 | 27.94%
DOI
10.1145/3786702
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
BibTeX Citation
Copy BibTeX
@inproceedings{shankar_sigmod26,
title = {{Task Cascades for Efficient Unstructured Data Processing}},
author = {Shankar, Shreya and Zeighami, Sepanta and Parameswaran, Aditya G.},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3786702},
url = {https://dl.acm.org/doi/10.1145/3786702},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
Outgoing Citations (Sorted by Pagerank)
Showing 29 of 29 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
132
Predicate Migration: Optimizing Queries with Expensive Predicates
1993
SIGMOD
0.00030378624
284
NoScope: Optimizing Neural Network Queries over Video at Scale
2017
VLDB
0.00022370521
295
Accelerating Machine Learning Inference with Probabilistic Predicates
2018
SIGMOD
0.00022238183
363
Approximate Query Processing: Taming the TeraBytes! A Tutorial
2001
VLDB
0.0002005475
508
Random Sampling for Histogram Construction: How much is enough?
1998
SIGMOD
0.00017275873
569
BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics
2020
VLDB
0.00016348191
713
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
2024
VLDB
0.00014672521
950
CAESURA: Language Models as Multi-Modal Query Planners
2024
CIDR
0.0001302491
1,108
Approximate Query Processing: No Silver Bullet
2017
SIGMOD
0.00012145154
1,245
Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing
2025
CIDR
0.00011507415
1,343
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
2025
VLDB
0.00011095866
2,553
LLM-R^2: A Large Language Model Enhanced Rule-based Rewrite System for Boosting Query Efficiency
2025
VLDB
8.4283807e-05
2,898
Approximate Selection with Guarantees using Proxies
2020
VLDB
7.978725e-05
2,956
The Design of an LLM-powered Unstructured Analytics System
2025
CIDR
7.9203461e-05
3,684
ThalamusDB: Approximate Query Processing on Multi-Modal Data
2024
SIGMOD
7.2033959e-05
3,766
TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data
2022
SIGMOD
7.1430942e-05
3,907
Accelerating Approximate Aggregation Queries with Expensive Predicates
2021
VLDB
7.0278233e-05
4,081
Abacus: A Cost-Based Optimizer for Semantic Operator Systems
2026
VLDB
6.9165634e-05
4,375
FiGO: Fine-Grained Query Optimization in Video Analytics
2022
SIGMOD
6.7369552e-05
4,435
Filtering with Approximate Predicates
1998
VLDB
6.7078883e-05
5,485
Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System
2025
SIGMOD
6.2013279e-05
6,118
ELEET: Efficient Learned Query Execution over Text and Tables
2024
VLDB
5.9698795e-05
6,778
VectraFlow: Integrating Vectors into Stream Processing
2025
CIDR
5.7772535e-05
7,316
SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint
2025
SIGMOD
5.646695e-05
7,439
AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries
2025
CIDR
5.6181103e-05
7,568
Semantic Operators and Their Optimization: Enabling LLM-Based Data Processing with Accuracy Guarantees in LOTUS
2025
VLDB
5.5953379e-05
7,648
Accelerating Aggregation Queries on Unstructured Streams of Data
2023
VLDB
5.575838e-05
8,453
A Learned Query Rewrite System
2023
VLDB
5.4229225e-05
8,627
ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries
2025
VLDB
5.3969543e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
10,318
In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration
2026
SIGMOD
2
713
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
2024
VLDB
3
10,856
Optimized Batch Prompting for Cost-effective LLMs
2025
VLDB
4
11,186
Unstructured Data Fusion for Schema and Data Extraction
2024
SIGMOD
5
10,733
ScaleLLM: A Technique for Scalable LLM-augmented Data Systems
2025
SIGMOD
6
13,339
DocDB: A Database for Unstructured Document Analysis
2025
VLDB
7
1,343
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
2025
VLDB
8
6,371
QUEST: Query Optimization in Unstructured Document Analysis
2025
VLDB
9
10,286
ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
2026
SIGMOD
10
8,829
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees
2026
SIGMOD