DBScholar

Back to papers

Optimizing Inference Serving on Serverless Platforms

Summary: MBS optimizes heterogeneous ML-inference batching under padding overhead and serverless autoscaling, using analytical performance/cost models with Bayesian optimization. On AWS bursty workloads, it meets SLOs while cutting cost up to 8×, padding 37×, and invocations 3×. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
12891
Venue
VLDB
Year
2022
Pagerank
5.3350531e-05
Overall Rank
9,000 | 38.26%
DOI
10.14778/3547305.3547313

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{ali_vldb22,
        title = {{Optimizing Inference Serving on Serverless Platforms}},
        author = {Ali, Ahsan and Pinciroli, Riccardo and Yan, Feng and Smirni, Evgenia},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {10},
        pages = {2071--2084},
        doi = {10.14778/3547305.3547313},
        url = {https://doi.org/10.14778/3547305.3547313},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Rank Citing Paper Year Venue Pagerank
9,475 BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach 2023 SIGMOD 5.2634238e-05
10,623 KEN: An Execution Engine for Unstructured Database Systems 2026 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 6 of 6 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers