DBScholar

Back to papers

Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking

Summary: OmniServe co-serves latency-sensitive and best-effort LLMs by piggybacking BE Attention on CPUs, asynchronously overlapping CPU/GPU streams to reduce interference. Dynamic batching sustains GPU utilization, improving LS SLO attainment 1.48× and BE throughput 9.85×. (summarized by gpt-5.6-luna on Jul 26 2026)

Paper ID
7480
Venue
SIGMOD
Year
2026
Pagerank
-
Overall Rank
13,288 | 8.84%
DOI
10.1145/3802107

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{mo_sigmod26,
        title = {{Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking}},
        author = {Mo, Zizhao and Chen, Junlin and Xu, Huanle and Xu, Chengzhong},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3802107},
        url = {https://dl.acm.org/doi/10.1145/3802107},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 0 of 0 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Semantically Similar Papers