Back to papers
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
Summary: Beluga exploits CXL switches to expose a shared, large-scale memory pool with native load/store access for GPU/CPU KVCache, avoiding RDMA’s latency/protocol overhead. Beluga-KVCache uses this architecture to scale long-context LLM inference, cutting TTFT 89.6% and boosting vLLM throughput 7.35x.
(summarized by gpt-5-mini on Apr 11 2026)
- Paper ID
- 7454
- Venue
- SIGMOD
- Year
- 2026
- Pagerank
- 5.1725247e-05
- Overall Rank
- 10,143 | 29.51%
- DOI
-
10.1145/3786627
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
Outgoing Citations (Sorted by Pagerank)
Showing 17 of 17 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 1,266 |
PolarFS: An Ultra-low Latency and Failure Resilient Distributed File System for Shared Storage Cloud Database |
2018 |
VLDB |
0.00011494173 |
| 1,960 |
ReAcTable: Enhancing ReAct for Table Question Answering |
2024 |
VLDB |
9.4913606e-05 |
| 2,007 |
Efficient Distributed Memory Management with RDMA and Caching |
2018 |
VLDB |
9.3992618e-05 |
| 3,200 |
Rethinking Database High Availability with RDMA Networks |
2019 |
VLDB |
7.7172767e-05 |
| 3,362 |
OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency Alignment |
2025 |
SIGMOD |
7.547454e-05 |
| 3,727 |
ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMA |
2022 |
SIGMOD |
7.2310205e-05 |
| 4,260 |
Design Guidelines for Correct, Efficient, and Scalable Synchronization using One-Sided RDMA |
2023 |
SIGMOD |
6.8645333e-05 |
| 5,518 |
Distributed GPU Joins on Fast RDMA-capable Networks |
2023 |
SIGMOD |
6.2532393e-05 |
| 6,291 |
DEX: Scalable Range Indexing on Disaggregated Memory |
2024 |
VLDB |
5.9874246e-05 |
| 6,397 |
InferDB: In-Database Machine Learning Inference Using Indexes |
2024 |
VLDB |
5.9545897e-05 |
| 6,600 |
Serving Deep Learning Models with Deduplication from Relational Databases |
2022 |
VLDB |
5.8861676e-05 |
| 7,190 |
SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint |
2025 |
SIGMOD |
5.7341493e-05 |
| 7,687 |
Rethinking Stateful Stream Processing with RDMA |
2022 |
SIGMOD |
5.6249337e-05 |
| 8,473 |
Zero-sided RDMA: Network-driven Data Shuffling for Disaggregated Heterogeneous Cloud DBMSs |
2024 |
SIGMOD |
5.4888649e-05 |
| 8,849 |
Unlocking the Potential of CXL for Disaggregated Memory in Cloud-Native Databases |
2025 |
SIGMOD |
5.4233138e-05 |
| 9,472 |
Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents |
2025 |
SIGMOD |
5.3246578e-05 |
| 10,788 |
From Scale-Up to Scale-Out: PolarDB's Journey to Achieving 2 Billion tpmC |
2025 |
VLDB |
5.1725247e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 10,222 |
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference |
2026 |
VLDB |
5.1725247e-05 |
| 10,066 |
DepCache: A KV Cache Management Framework for GraphRAG with Dependency Attention |
2026 |
SIGMOD |
5.1725247e-05 |
| 13,152 |
Database Perspective on LLM Inference Systems |
2025 |
VLDB |
- |
| 8,497 |
CXL Memory Performance for In-Memory Data Processing |
2025 |
VLDB |
5.4860111e-05 |
| 13,101 |
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration |
2026 |
VLDB |
- |
| 4,681 |
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation |
2025 |
SIGMOD |
6.6286556e-05 |
| 10,020 |
HotPrefix: Hotness-Aware KV Cache Scheduling for Efficient Prefix Sharing in LLM Inference Systems |
2026 |
SIGMOD |
5.1725247e-05 |
| 9,777 |
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training |
2025 |
SIGMOD |
5.2743647e-05 |
| 8,849 |
Unlocking the Potential of CXL for Disaggregated Memory in Cloud-Native Databases |
2025 |
SIGMOD |
5.4233138e-05 |
| 6,185 |
PQCache: Product Quantization-based KVCache for Long Context LLM Inference |
2025 |
SIGMOD |
6.0213899e-05 |