DBScholar

Back to papers

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

Summary: RetroInfer turns sparse KV-cache retrieval into vector storage via a wave index combining tripartite approximation, accuracy-bounded estimation, and segmented clustering. A GPU–CPU wave buffer enables up to 12.2× faster million-token decoding with full-attention accuracy. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
14444
Venue
VLDB
Year
2026
Pagerank
5.8801533e-05
Overall Rank
6,432 | 55.88%
DOI
10.14778/3796195.3796212

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{chen_vldb26,
        title = {{RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference}},
        author = {Chen, Yaoqi and Zhang, Jinkai and Lu, Baotong and Zhang, Qianxi and Zhang, Chengruidong and Liu, Jing and Luo, Jingjia and Liu, Di and Jiang, Huiqiang and Chen, Qi and Ding, Bailu and Yan, Xiao and Jiang, Jiawei and Chen, Chen and Zhang, Mingxing and Li, Cheng and Yang, Yuqing and Yang, Fan and Yang, Mao},
        journal = {PVLDB},
        series = {{VLDB} '26},
        volume = {19},
        number = {5},
        pages = {1016--1031},
        doi = {10.14778/3796195.3796212},
        url = {https://doi.org/10.14778/3796195.3796212},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
93 Fast Approximate Nearest Neighbor Search With The Navigating Spreading-out Graph 2019 VLDB 0.00034701237
286 Milvus: A Purpose-Built Vector Data Management System 2021 SIGMOD 0.00022357911
1,065 PASE: PostgreSQL Ultra-High-Dimensional Approximate Nearest Neighbor Search Extension 2020 SIGMOD 0.00012335063
1,515 ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data 2024 SIGMOD 0.00010521317
1,589 Manu: A Cloud Native Vector Database Management System 2022 VLDB 0.00010264469
1,631 High-Throughput Vector Similarity Search in Knowledge Graphs 2023 SIGMOD 0.00010174628
1,802 HVS: Hierarchical Graph Structure Based on Voronoi Diagrams for Solving Approximate Nearest Neighbor Search 2022 VLDB 9.7284341e-05
1,934 Towards Efficient Index Construction and Approximate Nearest Neighbor Search in High-Dimensional Spaces 2023 VLDB 9.4561907e-05
2,572 DeltaPQ: Lossless Product Quantization Code Compression for High Dimensional Similarity Search 2020 VLDB 8.4027322e-05
2,920 SingleStore-V: An Integrated Vector Database System in SingleStore 2024 VLDB 7.9628307e-05
3,761 ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMA 2022 SIGMOD 7.1450129e-05
4,018 Virtual-Memory Assisted Buffer Management 2023 SIGMOD 6.952278e-05
4,116 RoarGraph: A Projected Bipartite Graph for Efficient Cross-Modal Approximate Nearest Neighbor Search 2024 VLDB 6.892012e-05
4,430 PQCache: Product Quantization-based KVCache for Long Context LLM Inference 2025 SIGMOD 6.7091071e-05
4,886 Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs 2024 VLDB 6.4614955e-05
5,800 DET-LSH: A Locality-Sensitive Hashing Scheme with Dynamic Encoding Tree for Approximate Nearest Neighbor Search 2024 VLDB 6.0850924e-05
6,498 LeanStore: A High-Performance Storage Engine for NVMe SSDs 2024 VLDB 5.861854e-05
6,900 Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving 2025 SIGMOD 5.7430032e-05
7,826 AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference 2025 SIGMOD 5.53654e-05
8,466 TigerVector: Supporting Vector Search in Graph Databases for Advanced RAGs 2025 SIGMOD 5.4188627e-05
Previous Page 1 / 1 Next

Semantically Similar Papers