DBScholar

Back to papers

Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management

Summary: Beluga exploits CXL switches to expose a shared, large-scale memory pool with native load/store access for GPU/CPU KVCache, avoiding RDMA’s latency/protocol overhead. Beluga-KVCache uses this architecture to scale long-context LLM inference, cutting TTFT 89.6% and boosting vLLM throughput 7.35x. (summarized by gpt-5-mini on Apr 11 2026)

Paper ID
7644
Venue
SIGMOD
Year
2026
Pagerank
5.093636e-05
Overall Rank
10,432 | 28.43%
DOI
10.1145/3786627

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{yang_sigmod26,
        title = {{Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management}},
        author = {Yang, Xinjun and Hu, Qingda and Li, Junru and Li, Feifei and Zhu, Yicong and Zhou, Yuqi and Lin, Qiuru and Dai, Jian and Kong, Yang and Zhang, Jiayu and Xu, Guoqiang and Liu, Qiang},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3786627},
        url = {https://dl.acm.org/doi/10.1145/3786627},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 17 of 17 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1,283 PolarFS: An Ultra-low Latency and Failure Resilient Distributed File System for Shared Storage Cloud Database 2018 VLDB 0.00011337934
1,981 ReAcTable: Enhancing ReAct for Table Question Answering 2024 VLDB 9.3579557e-05
2,034 Efficient Distributed Memory Management with RDMA and Caching 2018 VLDB 9.2788175e-05
2,710 OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency Alignment 2025 SIGMOD 8.2206448e-05
3,243 Rethinking Database High Availability with RDMA Networks 2019 VLDB 7.6051655e-05
3,761 ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMA 2022 SIGMOD 7.1450129e-05
4,317 Design Guidelines for Correct, Efficient, and Scalable Synchronization using One-Sided RDMA 2023 SIGMOD 6.7646416e-05
5,456 InferDB: In-Database Machine Learning Inference Using Indexes 2024 VLDB 6.2131252e-05
5,587 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.1596139e-05
5,852 DEX: Scalable Range Indexing on Disaggregated Memory 2024 VLDB 6.0658952e-05
6,708 Serving Deep Learning Models with Deduplication from Relational Databases 2022 VLDB 5.7964983e-05
7,316 SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint 2025 SIGMOD 5.646695e-05
7,814 Rethinking Stateful Stream Processing with RDMA 2022 SIGMOD 5.5389009e-05
8,596 Zero-sided RDMA: Network-driven Data Shuffling for Disaggregated Heterogeneous Cloud DBMSs 2024 SIGMOD 5.4051515e-05
8,978 Unlocking the Potential of CXL for Disaggregated Memory in Cloud-Native Databases 2025 SIGMOD 5.3406001e-05
9,155 Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents 2025 SIGMOD 5.3122817e-05
11,011 From Scale-Up to Scale-Out: PolarDB’s Journey to Achieving 2 Billion tpmC 2025 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Semantically Similar Papers