Optimizing Inference Serving on Serverless Platforms
Summary: MBS optimizes heterogeneous ML-inference batching under padding overhead and serverless autoscaling, using analytical performance/cost models with Bayesian optimization. On AWS bursty workloads, it meets SLOs while cutting cost up to 8×, padding 37×, and invocations 3×. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Ahsan Ali (University of Nevada)
- 2. Riccardo Pinciroli (Gran Sasso Science Institute)
- 3. Feng Yan (University of Nevada)
- 4. Evgenia Smirni (William and Mary)
BibTeX Citation
@article{ali_vldb22,
title = {{Optimizing Inference Serving on Serverless Platforms}},
author = {Ali, Ahsan and Pinciroli, Riccardo and Yan, Feng and Smirni, Evgenia},
journal = {PVLDB},
series = {{VLDB} '22},
volume = {15},
number = {10},
pages = {2071--2084},
doi = {10.14778/3547305.3547313},
url = {https://doi.org/10.14778/3547305.3547313},
year = {2022}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,475 | BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach | 2023 | SIGMOD | 5.2634238e-05 |
| 10,623 | KEN: An Execution Engine for Unstructured Database Systems | 2026 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,334 | Improving Optimistic Concurrency Control Through Transaction Batching and Operation Reordering | 2019 | VLDB | 0.00011123567 |
| 1,784 | Lambada: Interactive Data Analytics on Cold Data Using Serverless Cloud Infrastructure | 2020 | SIGMOD | 9.7726335e-05 |
| 3,169 | Towards Demystifying Serverless Machine Learning Training | 2021 | SIGMOD | 7.6715222e-05 |
| 4,311 | Transactional Causal Consistency for Serverless Computing | 2020 | SIGMOD | 6.7676144e-05 |
| 4,974 | Stateful Functions as a Service in Action | 2019 | VLDB | 6.4178811e-05 |
| 5,508 | LMFAO: An Engine for Batches of Group-By Aggregates | 2020 | VLDB | 6.1922654e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,856 | Optimized Batch Prompting for Cost-effective LLMs | 2025 | VLDB |
| 2 | 9,664 | Cost-efficiency and Performance Robustness in Serverless Data Exchange | 2022 | SIGMOD |
| 3 | 4,987 | Releasing Cloud Databases from the Chains of Performance Prediction Models | 2017 | CIDR |
| 4 | 11,767 | Serverless Query Processing on a Budget | 2020 | SIGMOD |
| 5 | 4,436 | Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures | 2023 | VLDB |
| 6 | 13,288 | Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking | 2026 | SIGMOD |
| 7 | 6,900 | Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving | 2025 | SIGMOD |
| 8 | 3,169 | Towards Demystifying Serverless Machine Learning Training | 2021 | SIGMOD |
| 9 | 10,191 | AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System | 2026 | SIGMOD |
| 10 | 7,114 | Serverless Data Science - Are We There Yet? A Case Study of Model Serving | 2022 | SIGMOD |