Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures
Summary: JellyBean: jointly selects AutoML-generated model variants and places them across tiered heterogeneous infrastructure (edge/hubs/edge-DC/cloud) to meet SLOs (throughput, accuracy) while minimizing serving cost. Yields up to 58% cost reduction on VQA, 36% on vehicle tracking, and up to 5x cost savings versus cloud-only serving. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yongji Wu (Duke University)
- 2. Matthew Lentz (Duke University)
- 3. Danyang Zhuo (Duke University)
- 4. Yao Lu (Microsoft)
BibTeX Citation
@article{wu_vldb23,
title = {{Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures}},
author = {Wu, Yongji and Lentz, Matthew and Zhuo, Danyang and Lu, Yao},
journal = {PVLDB},
series = {{VLDB} '23},
volume = {16},
number = {3},
pages = {406--419},
doi = {10.14778/3570690.3570692},
url = {https://doi.org/10.14778/3570690.3570692},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 7 of 7 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,007 | Optimizing Video Analytics with Declarative Model Relationships | 2023 | VLDB | 6.9632395e-05 |
| 7,066 | Aero: Adaptive Query Processing of ML Queries | 2025 | SIGMOD | 5.712204e-05 |
| 7,087 | DAHA: Accelerating GNN Training with Data and Hardware Aware Execution Planning | 2024 | VLDB | 5.7069166e-05 |
| 7,847 | Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines | 2024 | VLDB | 5.5330423e-05 |
| 10,623 | KEN: An Execution Engine for Unstructured Database Systems | 2026 | VLDB | 5.093636e-05 |
| 10,691 | Flux: Unifying Heterogeneous Infrastructure for Alibaba AnalyticDB | 2025 | SIGMOD | 5.093636e-05 |
| 11,077 | Algorithmic Data Minimization for Machine Learning over Internet-of-Things Data Streams | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 290 | An Overview of Query Optimization in Relational Systems | 1998 | PODS | 0.0002227038 |
| 295 | Accelerating Machine Learning Inference with Probabilistic Predicates | 2018 | SIGMOD | 0.00022238183 |
| 569 | BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics | 2020 | VLDB | 0.00016348191 |
| 865 | Natural language to SQL: Where are we today? | 2020 | VLDB | 0.00013521464 |
| 4,114 | Optimizing Machine Learning Inference Queries with Correlative Proxy Models | 2022 | VLDB | 6.8941194e-05 |
| 5,582 | Yugong: Geo-Distributed Data and Job Placement at Scale | 2019 | VLDB | 6.1619082e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,539 | Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data Applications | 2022 | SIGMOD |
| 2 | 6,708 | Serving Deep Learning Models with Deduplication from Relational Databases | 2022 | VLDB |
| 3 | 10,596 | NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud Environments | 2026 | VLDB |
| 4 | 295 | Accelerating Machine Learning Inference with Probabilistic Predicates | 2018 | SIGMOD |
| 5 | 4,987 | Releasing Cloud Databases from the Chains of Performance Prediction Models | 2017 | CIDR |
| 6 | 2,162 | Heterogeneity-aware Distributed Parameter Servers | 2017 | SIGMOD |
| 7 | 7,763 | MLBench: Benchmarking Machine Learning Services Against Human Experts | 2018 | VLDB |
| 8 | 7,114 | Serverless Data Science - Are We There Yet? A Case Study of Model Serving | 2022 | SIGMOD |
| 9 | 9,566 | Declarative Data Serving: The Future of Machine Learning Inference on the Edge | 2021 | VLDB |
| 10 | 9,000 | Optimizing Inference Serving on Serverless Platforms | 2022 | VLDB |