CARINA: An Efficient CXL-Oriented Embedding Serving System for Recommendation Models
Summary: CARINA optimizes ERM serving on CXL by using heterogeneous memory: hot embeddings on DRAM and NUMA-aware placement of tables. Bandwidth-aware decomposition and scheduling prevent CXL saturation, yielding 5x throughput and 4x latency gains on real devices. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Peiqi Yin (Chinese University of Hong Kong)
- 2. Qihui Zhou (Chinese University of Hong Kong)
- 3. Xiao Yan (Centre for Perceptual and Interactive Intelligence)
- 4. Chao Wang (Chinese University of Hong Kong)
- 5. Eric Lo (Chinese University of Hong Kong)
- 6. Changji Li (Chinese University of Hong Kong)
- 7. Lan Lu (University of Pennsylvania)
- 8. Hua Fan (Alibaba)
- 9. Wenchao Zhou (Alibaba)
- 10. Ming-Chang Yang (Chinese University of Hong Kong)
- 11. James Cheng (Chinese University of Hong Kong)
BibTeX Citation
@inproceedings{yin_sigmod25,
title = {{CARINA: An Efficient CXL-Oriented Embedding Serving System for Recommendation Models}},
author = {Yin, Peiqi and Zhou, Qihui and Yan, Xiao and Wang, Chao and Lo, Eric and Li, Changji and Lu, Lan and Fan, Hua and Zhou, Wenchao and Yang, Ming-Chang and Cheng, James},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3725274},
url = {https://dl.acm.org/doi/10.1145/3725274},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,251 | GPS: Revisiting the Data Layout for Disk-based High-Dimensional Vector Search | 2026 | SIGMOD | 5.093636e-05 |
| 10,420 | SG-Serve: Efficient Model Serving for Subgraph-based Graph Representation Learning | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,631 | High-Throughput Vector Similarity Search in Knowledge Graphs | 2023 | SIGMOD | 0.00010174628 |
| 2,688 | Accelerating Recommendation System Training by Leveraging Popular Choices | 2022 | VLDB | 8.2564305e-05 |
| 7,190 | PetPS: Supporting Huge Embedding Models with Persistent Memory | 2023 | VLDB | 5.6772818e-05 |
| 7,191 | WiscSort: External Sorting For Byte-Addressable Storage | 2023 | VLDB | 5.6772818e-05 |
| 9,164 | FEC: Efficient Deep Recommendation Model Training with Flexible Embedding Communication | 2023 | SIGMOD | 5.3101785e-05 |
| 11,187 | GE2: A General and Efficient Knowledge Graph Embedding Learning System | 2024 | SIGMOD | 5.093636e-05 |
| 11,192 | Atom: An Efficient Query Serving System for Embedding-based Knowledge Graph Reasoning with Operator-level Batching | 2024 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,233 | FINEdex: A Fine-grained Learned Index Scheme for Scalable and Concurrent Memory Systems | 2022 | VLDB |
| 2 | 10,209 | CMANNS: GPU-Accelerated Graph Index Construction for ANNS via Compute-Memory Disaggregation | 2026 | SIGMOD |
| 3 | 10,115 | Hash Joins Meet CXL: A Fresh Look | 2026 | CIDR |
| 4 | 10,035 | CoTra: Towards Efficient and Scalable Distributed Vector Search with RDMA | 2026 | SIGMOD |
| 5 | 11,464 | EmbedX: A Versatile, Efficient and Scalable Platform to Embed Both Graphs and High-Dimensional Sparse Data | 2023 | VLDB |
| 6 | 9,520 | Experimental Analysis of Large-scale Learnable Vector Storage Compression | 2024 | VLDB |
| 7 | 8,609 | CXL Memory Performance for In-Memory Data Processing | 2025 | VLDB |
| 8 | 2,688 | Accelerating Recommendation System Training by Leveraging Popular Choices | 2022 | VLDB |
| 9 | 3,729 | CARMI: A Cache-Aware Learned Index with a Cost-based Construction Algorithm | 2022 | VLDB |
| 10 | 9,556 | CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models | 2024 | SIGMOD |