Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
Summary: Beluga exploits CXL switches to expose a shared, large-scale memory pool with native load/store access for GPU/CPU KVCache, avoiding RDMA’s latency/protocol overhead. Beluga-KVCache uses this architecture to scale long-context LLM inference, cutting TTFT 89.6% and boosting vLLM throughput 7.35x. (summarized by gpt-5-mini on Apr 11 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Xinjun Yang (Alibaba)
- 2. Qingda Hu (Alibaba)
- 3. Junru Li (Alibaba)
- 4. Feifei Li (Alibaba)
- 5. Yicong Zhu (Alibaba)
- 6. Yuqi Zhou (Alibaba)
- 7. Qiuru Lin (Alibaba)
- 8. Jian Dai (Alibaba)
- 9. Yang Kong (Alibaba)
- 10. Jiayu Zhang (Alibaba)
- 11. Guoqiang Xu (Alibaba)
- 12. Qiang Liu (Alibaba)
BibTeX Citation
@inproceedings{yang_sigmod26,
title = {{Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management}},
author = {Yang, Xinjun and Hu, Qingda and Li, Junru and Li, Feifei and Zhu, Yicong and Zhou, Yuqi and Lin, Qiuru and Dai, Jian and Kong, Yang and Zhang, Jiayu and Xu, Guoqiang and Liu, Qiang},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3786627},
url = {https://dl.acm.org/doi/10.1145/3786627},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 17 of 17 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next