RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
Summary: RetroInfer turns sparse KV-cache retrieval into vector storage via a wave index combining tripartite approximation, accuracy-bounded estimation, and segmented clustering. A GPU–CPU wave buffer enables up to 12.2× faster million-token decoding with full-attention accuracy. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yaoqi Chen (Microsoft; University of Science and Technology Beijing)
- 2. Jinkai Zhang (Microsoft; Wuhan University)
- 3. Baotong Lu (Microsoft)
- 4. Qianxi Zhang (Microsoft)
- 5. Chengruidong Zhang (Microsoft)
- 6. Jing Liu (Microsoft)
- 7. Jingjia Luo (Microsoft; Tsinghua University)
- 8. Di Liu (Shanghai Jiao Tong University)
- 9. Huiqiang Jiang (Microsoft)
- 10. Qi Chen (Microsoft)
- 11. Bailu Ding (Microsoft)
- 12. Xiao Yan (Institute for Mathematics and Artificial Intelligence, Wuhan; Wuhan University)
- 13. Jiawei Jiang (Wuhan University)
- 14. Chen Chen (Shanghai Jiao Tong University)
- 15. Mingxing Zhang (Tsinghua University)
- 16. Cheng Li (Institute of Artificial Intelligence, Hefei Comprehensive National Science Center; University of Science and Technology Beijing)
- 17. Yuqing Yang (Microsoft)
- 18. Fan Yang (Microsoft)
- 19. Mao Yang (Microsoft)
BibTeX Citation
@article{chen_vldb26,
title = {{RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference}},
author = {Chen, Yaoqi and Zhang, Jinkai and Lu, Baotong and Zhang, Qianxi and Zhang, Chengruidong and Liu, Jing and Luo, Jingjia and Liu, Di and Jiang, Huiqiang and Chen, Qi and Ding, Bailu and Yan, Xiao and Jiang, Jiawei and Chen, Chen and Zhang, Mingxing and Li, Cheng and Yang, Yuqing and Yang, Fan and Yang, Mao},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {5},
pages = {1016--1031},
doi = {10.14778/3796195.3796212},
url = {https://doi.org/10.14778/3796195.3796212},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,227 | Efficient Index Layout and Search Strategy for Large-scale High-dimensional Vector Similarity Search | 2026 | SIGMOD | 5.093636e-05 |
| 10,259 | KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next