KV Caching

KV: Key-Value

Momentum

29 papers in the last four weeks, up 16% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 276

All topics
CardsList
  1. PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving

    Sep 17, 2026Omkar Shewale, Deepak Kumar, Divakar Kumar YadavLLM Inference EfficiencyLLM Inference

  2. Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

    Sep 16, 2026Yipeng Liu, Yingqiang Zhang, Feifei Li +1LLM Inference EfficiencyKV Caching

  3. Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits

    Sep 16, 2026Mingyang Mao, Wyatt Mackey, Xiaomin LinRetrieval-Augmented GenerationKV Caching

  4. T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning With Dynamic Routing

    Sep 14, 2026Mingqian Yu, Wenpeng Zhang, Peilin ZhaoKV CachingRecurrent Transformers

  5. Pixel Decodability Is Not a Compression Signal: Causally Evaluating Importance Proxies for Visual KV-Cache Eviction

    Sep 14, 2026Chenyu Zhou, Qiliang Jiang, Shuning Wu +1KV CachingKV-Cache Eviction

  6. AgentKV: Phase-Aware KV Eviction for Agentic LLMs

    Sep 14, 2026Taowen Tony Liu, Jeffrey T. H. Wong, Can Xiao +3KV CachingKV-Cache Eviction

  7. OmniKVQuant: KV Cache Quantization for Omni-LLMs

    Sep 11, 2026Suho Yoo, Hyunjong Ok, Jongmin Choi +2Efficient Multimodal InferenceKV Caching

  8. Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs

    Sep 11, 2026Joseph Kanichai, Tiziano De Matteis, Animesh TrivediKV-Cache OffloadingKV Caching

  9. Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving

    Sep 9, 2026Hongjian Fan, Kevin Zhang, David Habinsky +1LLM Inference EfficiencyKV Caching

  10. KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

    Sep 9, 2026Xi Shi, Mengxin Zheng, Qian LouKV CachingLLM Inference Acceleration

  11. BIO-MEMART: Biometric-Aware KV Cache Memory for Multi-User LLM Agents

    Sep 8, 2026Yanhong Qian, Xuanying He, Qingguo Meng +3Agent MemoryPrivacy-Preserving Language Models

  12. Jacap: Robust KV Cache Eviction via Jacobian-Based Nonlinear Information Capacity Preservation

    Sep 8, 2026Jiaming Yang, Chenwei Tang, Liangli Zhen +2KV CachingKV-Cache Eviction

  13. MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

    Sep 7, 2026Michael Wang, Keith Li, Roozbeh BostandoostLLM Inference EfficiencyLLM Inference

  14. RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving

    Sep 7, 2026Yang Liu, Zhaokai Luo, Huayi Jin +10KV CachingLLM Inference Acceleration