Database Paper Browser

Back to papers

Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management

Summary: Beluga exploits CXL switches to expose a shared, large-scale memory pool with native load/store access for GPU/CPU KVCache, avoiding RDMA’s latency/protocol overhead. Beluga-KVCache uses this architecture to scale long-context LLM inference, cutting TTFT 89.6% and boosting vLLM throughput 7.35x. (summarized by gpt-5-mini on Apr 11 2026)

Paper ID
7454
Venue
SIGMOD
Year
2026
Pagerank
5.1725247e-05
Overall Rank
10,143 | 29.51%
DOI
10.1145/3786627

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 17 of 17 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1,266 PolarFS: An Ultra-low Latency and Failure Resilient Distributed File System for Shared Storage Cloud Database 2018 VLDB 0.00011494173
1,960 ReAcTable: Enhancing ReAct for Table Question Answering 2024 VLDB 9.4913606e-05
2,007 Efficient Distributed Memory Management with RDMA and Caching 2018 VLDB 9.3992618e-05
3,200 Rethinking Database High Availability with RDMA Networks 2019 VLDB 7.7172767e-05
3,362 OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency Alignment 2025 SIGMOD 7.547454e-05
3,727 ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMA 2022 SIGMOD 7.2310205e-05
4,260 Design Guidelines for Correct, Efficient, and Scalable Synchronization using One-Sided RDMA 2023 SIGMOD 6.8645333e-05
5,518 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.2532393e-05
6,291 DEX: Scalable Range Indexing on Disaggregated Memory 2024 VLDB 5.9874246e-05
6,397 InferDB: In-Database Machine Learning Inference Using Indexes 2024 VLDB 5.9545897e-05
6,600 Serving Deep Learning Models with Deduplication from Relational Databases 2022 VLDB 5.8861676e-05
7,190 SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint 2025 SIGMOD 5.7341493e-05
7,687 Rethinking Stateful Stream Processing with RDMA 2022 SIGMOD 5.6249337e-05
8,473 Zero-sided RDMA: Network-driven Data Shuffling for Disaggregated Heterogeneous Cloud DBMSs 2024 SIGMOD 5.4888649e-05
8,849 Unlocking the Potential of CXL for Disaggregated Memory in Cloud-Native Databases 2025 SIGMOD 5.4233138e-05
9,472 Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents 2025 SIGMOD 5.3246578e-05
10,788 From Scale-Up to Scale-Out: PolarDB's Journey to Achieving 2 Billion tpmC 2025 VLDB 5.1725247e-05
Previous Page 1 / 1 Next

Semantically Similar Papers