KV-Cache Reuse

Momentum

15 papers in the last four weeks, up 400% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 50

All topics
CardsList
  1. RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation

    Oct 8, 2026Sreetama Sarkar, Saptarshi Mitra, Sitao Huang +2KV-Cache ReuseLLM Inference Acceleration

  2. Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs

    Sep 28, 2026Zhengxiang Huang, Shengheng Chen, Chaoyue Niu +6AI Accelerator InferenceOn-Device Language Model Inference

  3. KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems

    Sep 28, 2026Hyesung Jeon, Hyeongju Ha, Seoyoung Lee +2Multi-Agent LLM SystemsKV-Cache Management

  4. PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

    Sep 28, 2026Hyesung Jeon, Hyeongju Ha, Jae-Joon KimMulti-Agent LLM SystemsMemory-Efficient Inference

  5. RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

    Sep 27, 2026Ruoling Qi, Yirui Liu, Xuaner Wu +5Retrieval-Augmented GenerationKV-Cache Management

  6. Budgeted Cache Repair for Cross-Context KV-Cache Reuse

    Sep 27, 2026Haeyong Kang, Chang D. YooKV-Cache ManagementKV-Cache Reuse

  7. In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

    Sep 26, 2026Yikai Wang, Xiao Han, Mengmeng Xu +9Long-Horizon Video GenerationDiffusion Model Caching

  8. Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits

    Sep 16, 2026Mingyang Mao, Wyatt Mackey, Xiaomin LinRetrieval-Augmented GenerationKV Caching

  9. Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving

    Sep 9, 2026Hongjian Fan, Kevin Zhang, David Habinsky +1LLM Inference EfficiencyKV Caching

  10. KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

    Sep 9, 2026Xi Shi, Mengxin Zheng, Qian LouKV CachingLLM Inference Acceleration

  11. RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving

    Sep 7, 2026Yang Liu, Zhaokai Luo, Huayi Jin +10KV CachingLLM Inference Acceleration

  12. CacheBridge: Efficient Cross-Model KV Cache Transfer

    Sep 1, 2026Xingyu Qu, Siyuan Lu, Zhiyu Chen +2Transformer InferenceKV Caching

  13. A Universal Context-Reuse Layer for Cross-Model KV Sharing

    Aug 31, 2026Yi Li, Dongming Jiang, Yi Zhao +1LLM InferenceKV Caching

  14. CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

    Aug 7, 2026Gyuwan Kim, Cheoneum Park, Tao YangLong-Context RetrievalRetrieval-Augmented Generation

  15. SemPIC: Learning Semantic Position-Independent KV Caches

    Jul 30, 2026Hui Xie, Peng Xiao, Yutong Deng +3KV CachingMemory-Efficient Inference

  16. InferScale: GPU-Native KV Injection for Personalized LLM Serving

    Jul 29, 2026Peter Li, Prashant PandeyLLM ServingMemory-Augmented Language Models

  17. Kalypso: Relational LLM Serving

    Jul 26, 2026Hojae Son, Md Ashraful Islam, Huy Gia Cao +2LLM ServingLLM Inference Scheduling

  18. HijackKV: New Threat in Position-Independent KV Cache Reuse

    Jul 22, 2026Yichi Zhang, Zhiqi Wang, Huan Zhang +1KV CachingLLM Security

  19. C2^2KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference

    Jul 20, 2026Chuheng Du, Junyi Chen, Hanlin Tang +7KV CachingLLM Inference Acceleration