AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference
Summary: AlayaDB rearchitects LLM inference by decoupling KV cache and attention into a dedicated vector store. It models attention/cache as a query-processing task, with a native optimizer, delivering lower resource use and higher quality than previous approaches. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yangshen Deng (AlayaDB AI)
- 2. Zhengxin You (AlayaDB AI; Southern University of Science and Technology)
- 3. Long Xiang (AlayaDB AI; Southern University of Science and Technology)
- 4. Qilong Li (AlayaDB AI; Southern University of Science and Technology)
- 5. Peiqi Yuan (AlayaDB AI; Southern University of Science and Technology)
- 6. Zhaoyang Hong (AlayaDB AI; Southern University of Science and Technology)
- 7. Yitao Zheng (AlayaDB AI; Southern University of Science and Technology)
- 8. Wanting Li (AlayaDB AI; Southern University of Science and Technology)
- 9. Runzhong Li (AlayaDB AI; Southern University of Science and Technology)
- 10. Haotian Liu (AlayaDB AI; Southern University of Science and Technology)
- 11. Kyriakos Mouratidis (Singapore Management University)
- 12. Man Lung Yiu (Hong Kong Polytechnic University)
- 13. Huan Li (Zhejiang University)
- 14. Qiaomu Shen (Beijing Institute of Technology)
- 15. Rui Mao (Shenzhen University)
- 16. Bo Tang (AlayaDB AI; Southern University of Science and Technology)
BibTeX Citation
@inproceedings{deng_sigmod25,
title = {{AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference}},
author = {Deng, Yangshen and You, Zhengxin and Xiang, Long and Li, Qilong and Yuan, Peiqi and Hong, Zhaoyang and Zheng, Yitao and Li, Wanting and Li, Runzhong and Liu, Haotian and Mouratidis, Kyriakos and Yiu, Man Lung and Li, Huan and Shen, Qiaomu and Mao, Rui and Tang, Bo},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3722212.3724428},
url = {https://dl.acm.org/doi/10.1145/3722212.3724428},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,432 | RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference | 2026 | VLDB | 5.8801533e-05 |
| 8,898 | Cracking Vector Search Indexes | 2025 | VLDB | 5.3495662e-05 |
| 10,227 | Efficient Index Layout and Search Strategy for Large-scale High-dimensional Vector Similarity Search | 2026 | SIGMOD | 5.093636e-05 |
| 10,290 | Skyline Retrieval meets Set-Cover Chunk Merging: A Cost-Effective RAG-Sketch for Long-Context LLM QA | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next