Efficient Cooperation-Aware Key and Value Management for LLM Inference
Summary: CoKV treats attention-head KV caching as a cooperative game, accounting for inter-head synergies rather than valuing heads independently. Precomputed contribution attribution enables globally effective cache-budget allocation for eviction and quantization, improving long-context and reasoning inference. (summarized by gpt-5.6-luna on Aug 17 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Qiheng Sun (Hong Kong Polytechnic University; Zhejiang University)
- 2. Hongwei Zhang (Zhejiang University)
- 3. Junxu Liu (Hong Kong Polytechnic University)
- 4. Haocheng Xia (University of Illinois Urbana-Champaign)
- 5. Jinfei Liu (Zhejiang University)
- 6. Kui Ren (Zhejiang University)
- 7. Haibo Hu (Hong Kong Polytechnic University)
BibTeX Citation
@article{sun_vldb26,
title = {{Efficient Cooperation-Aware Key and Value Management for LLM Inference}},
author = {Sun, Qiheng and Zhang, Hongwei and Liu, Junxu and Xia, Haocheng and Liu, Jinfei and Ren, Kui and Hu, Haibo},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {9},
pages = {2154--2167},
doi = {10.14778/3819518.3819541},
url = {https://doi.org/10.14778/3819518.3819541},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 185 | CockroachDB: The Resilient Geo-Distributed SQL Database | 2020 | SIGMOD | 0.00026072278 |
| 3,405 | CoroGraph: Bridging Cache Efficiency and Work Efficiency for Graph Algorithm Execution | 2024 | VLDB | 7.3286366e-05 |
| 4,296 | Redy: Remote Dynamic Memory Cache | 2022 | VLDB | 6.6805238e-05 |
| 4,928 | Efficient Sampling Approaches to Shapley Value Approximation | 2023 | SIGMOD | 6.3502007e-05 |
| 5,032 | Put an Elephant into a Fridge: Optimizing Cache Efficiency for In-memory Key-value Stores | 2020 | VLDB | 6.3053972e-05 |
| 5,161 | HotPrefix: Hotness-Aware KV Cache Scheduling for Efficient Prefix Sharing in LLM Inference Systems | 2026 | SIGMOD | 6.2478968e-05 |
| 6,166 | Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving | 2025 | SIGMOD | 5.8631131e-05 |
| 6,399 | Catalyst: Optimizing Cache Management for Large In-memory Key-value Systems | 2023 | VLDB | 5.798136e-05 |
| 7,935 | Cache-aware load balancing of data center applications | 2019 | VLDB | 5.4227879e-05 |
Previous
Page 1 / 1
Next