Back to papers
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees
Summary: Targets LLM cascade cost-quality tradeoffs by provable selection of when to use cheaper LLMs for record processing, addressing weak quality estimation in prior confidence-based cascades. BARGAIN uses adaptive sampling and statistical estimation tuned to data/task to give tight theoretical guarantees (accuracy/precision/recall) and empirically reduces cost up to 86% vs. state-of-the-art.
(summarized by gpt-5-mini on Feb 11 2026)
- Paper ID
- 7373
- Venue
- SIGMOD
- Year
- 2026
- Pagerank
- 4.1905499e-05
- Overall Rank
- 10,064 | 30.06%
- DOI
-
10.1145/3769776
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
Outgoing Citations (Sorted by Pagerank)
Showing 22 of 22 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 317 |
NoScope: Optimizing Neural Network Queries over Video at Scale |
2017 |
VLDB |
0.0002798145 |
| 332 |
Accelerating Machine Learning Inference with Probabilistic Predicates |
2018 |
SIGMOD |
0.00027173479 |
| 516 |
Can Foundation Models Wrangle Your Data? |
2023 |
VLDB |
0.00021194444 |
| 694 |
BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics |
2020 |
VLDB |
0.00018031141 |
| 997 |
CAESURA: Language Models as Multi-Modal Query Planners |
2024 |
CIDR |
0.00014726927 |
| 1,088 |
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes |
2024 |
VLDB |
0.00014158762 |
| 1,839 |
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing |
2025 |
VLDB |
0.00010351287 |
| 3,553 |
Approximate Selection with Guarantees using Proxies |
2020 |
VLDB |
6.9763548e-05 |
| 3,639 |
The Design of an LLM-powered Unstructured Analytics System |
2025 |
CIDR |
6.8886648e-05 |
| 3,982 |
How Large Language Models Will Disrupt Data Management |
2023 |
VLDB |
6.5595332e-05 |
| 4,492 |
TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data |
2022 |
SIGMOD |
6.1374891e-05 |
| 4,703 |
Accelerating Approximate Aggregation Queries with Expensive Predicates |
2021 |
VLDB |
5.9793615e-05 |
| 4,865 |
OTIF: Efficient Tracker Pre-processing over Large Video Datasets |
2022 |
SIGMOD |
5.8627966e-05 |
| 5,168 |
FiGO: Fine-Grained Query Optimization in Video Analytics |
2022 |
SIGMOD |
5.6446115e-05 |
| 5,470 |
RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes |
2024 |
VLDB |
5.4894925e-05 |
| 7,369 |
ELEET: Efficient Learned Query Execution over Text and Tables |
2024 |
VLDB |
4.7452331e-05 |
| 7,703 |
AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries |
2025 |
CIDR |
4.668568e-05 |
| 7,869 |
Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach |
2023 |
SIGMOD |
4.6275089e-05 |
| 7,911 |
Accelerating Aggregation Queries on Unstructured Streams of Data |
2023 |
VLDB |
4.6143141e-05 |
| 9,240 |
ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries |
2025 |
VLDB |
4.3648789e-05 |
| 9,311 |
On Efficient Approximate Queries over Machine Learning Models |
2023 |
VLDB |
4.3535588e-05 |
| 9,728 |
Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems |
2025 |
VLDB |
4.2901665e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 13,152 |
Database Perspective on LLM Inference Systems |
2025 |
VLDB |
- |
| 1,088 |
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes |
2024 |
VLDB |
0.00014158762 |
| 3,803 |
Revisiting Prompt Engineering via Declarative Crowdsourcing |
2024 |
CIDR |
6.7498941e-05 |
| 10,328 |
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning |
2026 |
VLDB |
4.1905499e-05 |
| 7,016 |
LLM for Data Management |
2024 |
VLDB |
4.8561622e-05 |
| 10,462 |
ScaleLLM: A Technique for Scalable LLM-augmented Data Systems |
2025 |
SIGMOD |
4.1905499e-05 |
| 7,335 |
SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint |
2025 |
SIGMOD |
4.7533835e-05 |
| 10,215 |
Task Cascades for Efficient Unstructured Data Processing |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,603 |
Optimized Batch Prompting for Cost-effective LLMs |
2025 |
VLDB |
4.1905499e-05 |
| 9,240 |
ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries |
2025 |
VLDB |
4.3648789e-05 |