DepCache: A KV Cache Management Framework for GraphRAG with Dependency Attention
Summary: Dependency attention: graph-aware attention that prunes token-pair interactions to structural dependencies and reuses computations along relational paths to reduce inference cost. DepCache: KV-cache reuse aligned across graph-augmented prompts with a locality-aware replacement policy, yielding 1.5–5× throughput and up to 3.2× time-to-first-token reduction without accuracy loss. (summarized by gpt-5-mini on Feb 11 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Hao Yuan (Northeastern University)
- 2. Xin Ai (Northeastern University)
- 3. Qiange Wang (Northeastern University)
- 4. Peizheng Li (Northeastern University)
- 5. Jiayang Yu (Northeastern University)
- 6. Chaoyi Chen (Northeastern University)
- 7. Xinbo Yang (Northeastern University)
- 8. Yanfeng Zhang (Northeastern University)
- 9. Zhenbo Fu (Northeastern University)
- 10. Yingyou Wen (Neusoft AI Magic Technology Research)
- 11. Ge Yu (Northeastern University)
BibTeX Citation
@inproceedings{yuan_sigmod26,
title = {{DepCache: A KV Cache Management Framework for GraphRAG with Dependency Attention}},
author = {Yuan, Hao and Ai, Xin and Wang, Qiange and Li, Peizheng and Yu, Jiayang and Chen, Chaoyi and Yang, Xinbo and Zhang, Yanfeng and Fu, Zhenbo and Wen, Yingyou and Yu, Ge},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3769778},
url = {https://dl.acm.org/doi/10.1145/3769778},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,233 | FINEdex: A Fine-grained Learned Index Scheme for Scalable and Concurrent Memory Systems | 2022 | VLDB | 8.8968964e-05 |
| 2,695 | NeutronStar: Distributed GNN Training with Hybrid Dependency Management | 2022 | SIGMOD | 8.2468134e-05 |
| 2,816 | Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation | 2025 | SIGMOD | 8.0959781e-05 |
| 4,430 | PQCache: Product Quantization-based KVCache for Long Context LLM Inference | 2025 | SIGMOD | 6.7091071e-05 |
| 5,316 | Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective | 2024 | VLDB | 6.2687017e-05 |
| 6,806 | HongTu: Scalable Full-Graph GNN Training on Multiple GPUs | 2023 | SIGMOD | 5.7673207e-05 |
| 7,906 | Buffered Persistence in B+ Trees | 2024 | SIGMOD | 5.5181056e-05 |
| 9,546 | NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism | 2025 | VLDB | 5.2528121e-05 |
| 13,309 | NeutronRAG: Towards Understanding the Effectiveness of RAG from a Data Retrieval Perspective | 2025 | SIGMOD | - |
Previous
Page 1 / 1
Next