DBScholar

Back to papers

Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees

Summary: Targets LLM cascade cost-quality tradeoffs by provable selection of when to use cheaper LLMs for record processing, addressing weak quality estimation in prior confidence-based cascades. BARGAIN uses adaptive sampling and statistical estimation tuned to data/task to give tight theoretical guarantees (accuracy/precision/recall) and empirically reduces cost up to 86% vs. state-of-the-art. (summarized by gpt-5-mini on Feb 11 2026)

Paper ID
h0b08d826353bd4cb
Venue
SIGMOD
Year
2026
Pagerank
5.2410834e-05
Overall Rank
8,997 | 39.51%
DOI
10.1145/3769776

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{zeighami_sigmod26,
        title = {{Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees}},
        author = {Zeighami, Sepanta and Shankar, Shreya and Parameswaran, Aditya},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3769776},
        url = {https://dl.acm.org/doi/10.1145/3769776},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 22 of 22 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
271 NoScope: Optimizing Neural Network Queries over Video at Scale 2017 VLDB 0.00022560564
281 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022295232
329 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00020858443
501 Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2024 VLDB 0.00017267905
541 BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics 2020 VLDB 0.00016657685
669 CAESURA: Language Models as Multi-Modal Query Planners 2024 CIDR 0.0001495987
683 DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing 2025 VLDB 0.00014817539
2,341 The Design of an LLM-powered Unstructured Analytics System 2025 CIDR 8.6066183e-05
2,776 Approximate Selection with Guarantees using Proxies 2020 VLDB 8.0309448e-05
3,250 How Large Language Models Will Disrupt Data Management 2023 VLDB 7.4938546e-05
3,650 TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data 2022 SIGMOD 7.1341771e-05
3,874 Accelerating Approximate Aggregation Queries with Expensive Predicates 2021 VLDB 6.953738e-05
3,884 FiGO: Fine-Grained Query Optimization in Video Analytics 2022 SIGMOD 6.948464e-05
3,948 OTIF: Efficient Tracker Pre-processing over Large Video Datasets 2022 SIGMOD 6.9053923e-05
3,973 RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes 2024 VLDB 6.8876964e-05
4,142 AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries 2025 CIDR 6.7814795e-05
5,410 ELEET: Efficient Learned Query Execution over Text and Tables 2024 VLDB 6.1408676e-05
5,911 On Efficient Approximate Queries over Machine Learning Models 2023 VLDB 5.9501384e-05
6,993 Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach 2023 SIGMOD 5.626765e-05
7,801 Accelerating Aggregation Queries on Unstructured Streams of Data 2023 VLDB 5.4507311e-05
7,885 Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems 2025 VLDB 5.433531e-05
8,316 ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries 2025 VLDB 5.3561732e-05
Previous Page 1 / 1 Next

Semantically Similar Papers