DBScholar

Back to papers

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference

Summary: KVDrive manages LLM KV caches holistically across GPU memory, DRAM, and SSD, rather than relying on increased sparsity. Attention-aware placement, tier coordination, and overlapped pipeline scheduling reduce data movement and stalls, delivering up to 1.74× higher throughput. (summarized by gpt-5.6-luna on Jul 26 2026)

Paper ID
7450
Venue
SIGMOD
Year
2026
Pagerank
5.093636e-05
Overall Rank
10,259 | 29.62%
DOI
10.1145/3802077

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{lin_sigmod26,
        title = {{KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference}},
        author = {Lin, Jian and Mi, Jiazhi and Hong, Zicong and Wang, Haodong and Liu, Qianli and Zhang, Haoyue and Li, Peng and Guo, Song},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3802077},
        url = {https://dl.acm.org/doi/10.1145/3802077},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers