SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint
Summary: SpareLLM selects task-specific, minimum-cost LLMs with an output-equivalence constraint. Profiling-first approach and heterogeneous model cascades yield Pareto-optimal cost-accuracy tradeoffs, up to 8.6x savings and 90% equivalence to GPT-4-Turbo. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Saehan Jo (Cornell University)
- 2. Immanuel Trummer (Cornell University)
BibTeX Citation
@inproceedings{jo_sigmod25,
title = {{SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint}},
author = {Jo, Saehan and Trummer, Immanuel},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3725356},
url = {https://dl.acm.org/doi/10.1145/3725356},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,886 | Multi-Objective Agentic Rewrites for Unstructured Data Processing | 2026 | VLDB | 5.6551172e-05 |
| 10,621 | Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management | 2026 | SIGMOD | 4.9793485e-05 |
| 10,670 | PRISM: Navigating Cost–Accuracy Trade-offs for NL2SQL | 2026 | SIGMOD | 4.9793485e-05 |
| 10,690 | Task Cascades for Efficient Unstructured Data Processing | 2026 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 784 | VerdictDB: Universalizing Approximate Query Processing | 2018 | SIGMOD | 0.00014012614 |
| 840 | Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters | 2016 | SIGMOD | 0.0001354605 |
| 1,082 | Approximate Query Processing: No Silver Bullet | 2017 | SIGMOD | 0.00012122749 |
| 1,243 | Blink and It's Done: Interactive Queries on Very Large Data | 2012 | VLDB | 0.0001135375 |
| 1,661 | Rapid Sampling for Visualizations with Ordering Guarantees | 2015 | VLDB | 9.9535453e-05 |
| 2,456 | A Sampling Algebra for Aggregate Estimation | 2013 | VLDB | 8.4377192e-05 |
| 12,052 | BitGourmet: Deterministic Approximation via Optimized Bit Selection | 2020 | CIDR | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,711 | BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs | 2026 | VLDB |
| 2 | 9,158 | Database Perspective on LLM Inference Systems | 2025 | VLDB |
| 3 | 10,988 | Exploring the Benefits of Just-in-time Model Replacement | 2026 | VLDB |
| 4 | 10,488 | OctoSelector: Efficient and Effective Batch-Aware Model Selection for Large Language Models | 2026 | SIGMOD |
| 5 | 8,888 | Optimized Batch Prompting for Cost-effective LLMs | 2025 | VLDB |
| 6 | 10,838 | Unified Static–Dynamic Pruning for Efficient LLM Inference | 2026 | VLDB |
| 7 | 10,757 | Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs | 2026 | VLDB |
| 8 | 11,161 | ScaleLLM: A Technique for Scalable LLM-augmented Data Systems | 2025 | SIGMOD |
| 9 | 8,997 | Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees | 2026 | SIGMOD |
| 10 | 8,316 | ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries | 2025 | VLDB |