DBScholar

Back to papers

Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management

Summary: Beluga exploits CXL switches to expose a shared, large-scale memory pool with native load/store access for GPU/CPU KVCache, avoiding RDMA’s latency/protocol overhead. Beluga-KVCache uses this architecture to scale long-context LLM inference, cutting TTFT 89.6% and boosting vLLM throughput 7.35x. (summarized by gpt-5-mini on Apr 11 2026)

Paper ID
h30a74d14fb460e64
Venue
SIGMOD
Year
2026
Pagerank
4.9793485e-05
Overall Rank
10,621 | 28.60%
DOI
10.1145/3786627

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{yang_sigmod26,
        title = {{Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management}},
        author = {Yang, Xinjun and Hu, Qingda and Li, Junru and Li, Feifei and Zhu, Yicong and Zhou, Yuqi and Lin, Qiuru and Dai, Jian and Kong, Yang and Zhang, Jiayu and Xu, Guoqiang and Liu, Qiang},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3786627},
        url = {https://dl.acm.org/doi/10.1145/3786627},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 17 of 17 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1,217 PolarFS: An Ultra-low Latency and Failure Resilient Distributed File System for Shared Storage Cloud Database 2018 VLDB 0.00011489607
1,683 ReAcTable: Enhancing ReAct for Table Question Answering 2024 VLDB 9.8822754e-05
2,025 OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency Alignment 2025 SIGMOD 9.163622e-05
2,030 Efficient Distributed Memory Management with RDMA and Caching 2018 VLDB 9.1579506e-05
3,270 Rethinking Database High Availability with RDMA Networks 2019 VLDB 7.475162e-05
3,826 ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMA 2022 SIGMOD 6.9984204e-05
4,387 Design Guidelines for Correct, Efficient, and Scalable Synchronization using One-Sided RDMA 2023 SIGMOD 6.6233353e-05
5,272 InferDB: In-Database Machine Learning Inference Using Indexes 2024 VLDB 6.2010954e-05
5,371 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.1584802e-05
5,949 DEX: Scalable Range Indexing on Disaggregated Memory 2024 VLDB 5.9368754e-05
6,844 Serving Deep Learning Models with Deduplication from Relational Databases 2022 VLDB 5.6664512e-05
6,860 SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint 2025 SIGMOD 5.6613762e-05
7,945 Unlocking the Potential of CXL for Disaggregated Memory in Cloud-Native Databases 2025 SIGMOD 5.4212704e-05
7,976 Rethinking Stateful Stream Processing with RDMA 2022 SIGMOD 5.4146123e-05
8,230 Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents 2025 SIGMOD 5.372368e-05
8,752 Zero-sided RDMA: Network-driven Data Shuffling for Disaggregated Heterogeneous Cloud DBMSs 2024 SIGMOD 5.2847178e-05
10,259 From Scale-Up to Scale-Out: PolarDB’s Journey to Achieving 2 Billion tpmC 2025 VLDB 5.050482e-05
Previous Page 1 / 1 Next

Semantically Similar Papers