Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs
Summary: Mil formulates offline multi-LLM inference as an NP-hard minimum-makespan problem coupling GPU allocation, parallelism, and orchestration under relaxed precedences. Its cost-guided rate estimation, theoretically grounded greedy scheduling, and runtime adjustment deliver up to 3.4× speedups. (summarized by gpt-5.6-luna on Aug 17 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Jingzhi Fang (Hong Kong University of Science and Technology)
- 2. Yanyan Shen (Shanghai Jiao Tong University)
- 3. Yue Wang (Shenzhen University)
- 4. Lei Chen (Hong Kong University of Science and Technology)
BibTeX Citation
@article{fang_vldb26,
title = {{Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs}},
author = {Fang, Jingzhi and Shen, Yanyan and Wang, Yue and Chen, Lei},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {9},
pages = {1893--1906},
doi = {10.14778/3819518.3819522},
url = {https://doi.org/10.14778/3819518.3819522},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 329 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00020858443 |
| 3,973 | RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes | 2024 | VLDB | 6.8876964e-05 |
| 7,069 | Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads | 2024 | VLDB | 5.6083188e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,161 | ScaleLLM: A Technique for Scalable LLM-augmented Data Systems | 2025 | SIGMOD |
| 2 | 13,578 | Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking | 2026 | SIGMOD |
| 3 | 6,166 | Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving | 2025 | SIGMOD |
| 4 | 8,316 | ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries | 2025 | VLDB |
| 5 | 8,887 | mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs | 2025 | VLDB |
| 6 | 10,838 | Unified Static–Dynamic Pruning for Efficient LLM Inference | 2026 | VLDB |
| 7 | 10,435 | DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization | 2026 | SIGMOD |
| 8 | 8,888 | Optimized Batch Prompting for Cost-effective LLMs | 2025 | VLDB |
| 9 | 6,860 | SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint | 2025 | SIGMOD |
| 10 | 9,158 | Database Perspective on LLM Inference Systems | 2025 | VLDB |