DBScholar

Back to papers

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

Summary: RetroInfer turns sparse KV-cache retrieval into vector storage via a wave index combining tripartite approximation, accuracy-bounded estimation, and segmented clustering. A GPU–CPU wave buffer enables up to 12.2× faster million-token decoding with full-attention accuracy. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h77ccaef2a0eb0f9c
Venue
VLDB
Year
2026
Pagerank
5.7482184e-05
Overall Rank
6,560 | 55.90%
DOI
10.14778/3796195.3796212

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{chen_vldb26,
        title = {{RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference}},
        author = {Chen, Yaoqi and Zhang, Jinkai and Lu, Baotong and Zhang, Qianxi and Zhang, Chengruidong and Liu, Jing and Luo, Jingjia and Liu, Di and Jiang, Huiqiang and Chen, Qi and Ding, Bailu and Yan, Xiao and Jiang, Jiawei and Chen, Chen and Zhang, Mingxing and Li, Cheng and Yang, Yuqing and Yang, Fan and Yang, Mao},
        journal = {PVLDB},
        series = {{VLDB} '26},
        volume = {19},
        number = {5},
        pages = {1016--1031},
        doi = {10.14778/3796195.3796212},
        url = {https://doi.org/10.14778/3796195.3796212},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
74 Fast Approximate Nearest Neighbor Search With The Navigating Spreading-out Graph 2019 VLDB 0.00037091678
194 Milvus: A Purpose-Built Vector Data Management System 2021 SIGMOD 0.00025636725
916 PASE: PostgreSQL Ultra-High-Dimensional Approximate Nearest Neighbor Search Extension 2020 SIGMOD 0.00013094482
1,225 ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data 2024 SIGMOD 0.00011444398
1,373 High-Throughput Vector Similarity Search in Knowledge Graphs 2023 SIGMOD 0.0001088854
1,462 Manu: A Cloud Native Vector Database Management System 2022 VLDB 0.0001058099
1,625 Towards Efficient Index Construction and Approximate Nearest Neighbor Search in High-Dimensional Spaces 2023 VLDB 0.0001004502
1,684 HVS: Hierarchical Graph Structure Based on Voronoi Diagrams for Solving Approximate Nearest Neighbor Search 2022 VLDB 9.8801753e-05
2,085 SingleStore-V: An Integrated Vector Database System in SingleStore 2024 VLDB 9.0709364e-05
2,448 DeltaPQ: Lossless Product Quantization Code Compression for High Dimensional Similarity Search 2020 VLDB 8.4494625e-05
3,826 ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMA 2022 SIGMOD 6.9984204e-05
3,881 RoarGraph: A Projected Bipartite Graph for Efficient Cross-Modal Approximate Nearest Neighbor Search 2024 VLDB 6.9509289e-05
4,094 Virtual-Memory Assisted Buffer Management 2023 SIGMOD 6.8103269e-05
4,498 PQCache: Product Quantization-based KVCache for Long Context LLM Inference 2025 SIGMOD 6.5762771e-05
4,718 Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs 2024 VLDB 6.4551989e-05
5,360 LeanStore: A High-Performance Storage Engine for NVMe SSDs 2024 VLDB 6.1614422e-05
5,527 DET-LSH: A Locality-Sensitive Hashing Scheme with Dynamic Encoding Tree for Approximate Nearest Neighbor Search 2024 VLDB 6.0936219e-05
6,166 Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving 2025 SIGMOD 5.8631131e-05
6,473 AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference 2025 SIGMOD 5.7739763e-05
7,721 TigerVector: Supporting Vector Search in Graph Databases for Advanced RAGs 2025 SIGMOD 5.4677994e-05
Previous Page 1 / 1 Next

Semantically Similar Papers