SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint
Summary: SpareLLM selects task-specific, minimum-cost LLMs with an output-equivalence constraint. Profiling-first approach and heterogeneous model cascades yield Pareto-optimal cost-accuracy tradeoffs, up to 8.6x savings and 90% equivalence to GPT-4-Turbo. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Saehan Jo (Cornell University)
- 2. Immanuel Trummer (Cornell University)
BibTeX Citation
@inproceedings{jo_sigmod25,
title = {{SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint}},
author = {Jo, Saehan and Trummer, Immanuel},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3725356},
url = {https://dl.acm.org/doi/10.1145/3725356},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,432 | Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management | 2026 | SIGMOD | 5.093636e-05 |
| 10,483 | PRISM: Navigating Cost–Accuracy Trade-offs for NL2SQL | 2026 | SIGMOD | 5.093636e-05 |
| 10,504 | Task Cascades for Efficient Unstructured Data Processing | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 772 | VerdictDB: Universalizing Approximate Query Processing | 2018 | SIGMOD | 0.00014147905 |
| 819 | Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters | 2016 | SIGMOD | 0.00013815639 |
| 1,108 | Approximate Query Processing: No Silver Bullet | 2017 | SIGMOD | 0.00012145154 |
| 1,227 | Blink and It's Done: Interactive Queries on Very Large Data | 2012 | VLDB | 0.00011582387 |
| 1,634 | Rapid Sampling for Visualizations with Ordering Guarantees | 2015 | VLDB | 0.00010163938 |
| 2,413 | A Sampling Algebra for Aggregate Estimation | 2013 | VLDB | 8.6116764e-05 |
| 11,749 | BitGourmet: Deterministic Approximation via Optimized Bit Selection | 2020 | CIDR | 5.093636e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,504 | Task Cascades for Efficient Unstructured Data Processing | 2026 | SIGMOD |
| 2 | 7,075 | Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity | 2024 | VLDB |
| 3 | 713 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB |
| 4 | 10,527 | BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs | 2026 | VLDB |
| 5 | 13,343 | Database Perspective on LLM Inference Systems | 2025 | VLDB |
| 6 | 10,276 | OctoSelector: Efficient and Effective Batch-Aware Model Selection for Large Language Models | 2026 | SIGMOD |
| 7 | 10,856 | Optimized Batch Prompting for Cost-effective LLMs | 2025 | VLDB |
| 8 | 10,733 | ScaleLLM: A Technique for Scalable LLM-augmented Data Systems | 2025 | SIGMOD |
| 9 | 8,829 | Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees | 2026 | SIGMOD |
| 10 | 8,627 | ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries | 2025 | VLDB |