DBScholar

Back to papers

Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees

Summary: Targets LLM cascade cost-quality tradeoffs by provable selection of when to use cheaper LLMs for record processing, addressing weak quality estimation in prior confidence-based cascades. BARGAIN uses adaptive sampling and statistical estimation tuned to data/task to give tight theoretical guarantees (accuracy/precision/recall) and empirically reduces cost up to 86% vs. state-of-the-art. (summarized by gpt-5-mini on Feb 11 2026)

Paper ID
7562
Venue
SIGMOD
Year
2026
Pagerank
5.3613784e-05
Overall Rank
8,829 | 39.43%
DOI
10.1145/3769776

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{zeighami_sigmod26,
        title = {{Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees}},
        author = {Zeighami, Sepanta and Shankar, Shreya and Parameswaran, Aditya},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3769776},
        url = {https://dl.acm.org/doi/10.1145/3769776},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Rank Citing Paper Year Venue Pagerank
10,199 Automated Discovery of Test Oracles for Database Management Systems Using LLMs 2026 SIGMOD 5.093636e-05
10,286 ScaleDoc: Scaling LLM-based Predicates over Large Document Collections 2026 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 22 of 22 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
284 NoScope: Optimizing Neural Network Queries over Video at Scale 2017 VLDB 0.00022370521
295 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022238183
420 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00018789852
569 BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics 2020 VLDB 0.00016348191
713 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00014672521
950 CAESURA: Language Models as Multi-Modal Query Planners 2024 CIDR 0.0001302491
1,343 DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing 2025 VLDB 0.00011095866
2,898 Approximate Selection with Guarantees using Proxies 2020 VLDB 7.978725e-05
2,956 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 7.9203461e-05
3,536 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.3297343e-05
3,766 TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data 2022 SIGMOD 7.1430942e-05
3,907 Accelerating Approximate Aggregation Queries with Expensive Predicates 2021 VLDB 7.0278233e-05
4,298 OTIF: Efficient Tracker Pre-processing over Large Video Datasets 2022 SIGMOD 6.7750533e-05
4,375 FiGO: Fine-Grained Query Optimization in Video Analytics 2022 SIGMOD 6.7369552e-05
5,682 RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes 2024 VLDB 6.123486e-05
6,118 ELEET: Efficient Learned Query Execution over Text and Tables 2024 VLDB 5.9698795e-05
7,220 Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach 2023 SIGMOD 5.6679948e-05
7,439 AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries 2025 CIDR 5.6181103e-05
7,648 Accelerating Aggregation Queries on Unstructured Streams of Data 2023 VLDB 5.575838e-05
8,627 ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries 2025 VLDB 5.3969543e-05
9,353 On Efficient Approximate Queries over Machine Learning Models 2023 VLDB 5.2829539e-05
9,841 Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems 2025 VLDB 5.2101877e-05
Previous Page 1 / 1 Next

Semantically Similar Papers