CoDec: Prefix-Shared Decoding Kernel for LLMs
Summary: CoDec is a dedicated decode-stage attention kernel for prefix-sharing LLM prompts, coalescing shared-prefix KV-cache accesses. It addresses tree-induced dependencies and irregular workloads by improving memory-hierarchy utilization, mitigating the context-length bottleneck. (summarized by gpt-5.6-luna on Jul 26 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Zhibin Wang (Nanjing University)
- 2. Rui Ning (Nanjing University)
- 3. Chao Fang (Nanjing University)
- 4. Zhonghui Zhang (Nanjing University)
- 5. Xi Lin (Nanjing University)
- 6. Shaobo Ma (Nanjing University)
- 7. Mo Zhou (Nanjing University)
- 8. Xue Li (Alibaba)
- 9. Zhongfeng Wang (Nanjing University)
- 10. Chengying Huan (Nanjing University)
- 11. Rong Gu (Nanjing University)
- 12. Kun Yang (Nanjing University)
- 13. Guihai Chen (Nanjing University)
- 14. Sheng Zhong (Nanjing University)
- 15. Chen Tian (Nanjing University)
BibTeX Citation
@inproceedings{wang_sigmod26,
title = {{CoDec: Prefix-Shared Decoding Kernel for LLMs}},
author = {Wang, Zhibin and Ning, Rui and Fang, Chao and Zhang, Zhonghui and Lin, Xi and Ma, Shaobo and Zhou, Mo and Li, Xue and Wang, Zhongfeng and Huan, Chengying and Gu, Rong and Yang, Kun and Chen, Guihai and Zhong, Sheng and Tian, Chen},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3802028},
url = {https://dl.acm.org/doi/10.1145/3802028},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,816 | Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation | 2025 | SIGMOD | 8.0959781e-05 |
| 6,788 | HotPrefix: Hotness-Aware KV Cache Scheduling for Efficient Prefix Sharing in LLM Inference Systems | 2026 | SIGMOD | 5.7727874e-05 |
Previous
Page 1 / 1
Next