KV Caching

KV: Key-Value

Momentum

29 papers in the last four weeks, up 16% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 276

All topics
CardsList
  1. ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference

    Feb 8, 2026Jitai Hao, Qiang Huang, Yaowei Wang +2Self-AttentionKV Caching

  2. Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation

    Dec 11, 2025Duo Zheng, Shijia Huang, Yanyang Li +1KV CachingEfficient VLM Inference

  3. NOSA: Native and Offloadable Sparse Attention

    Oct 15, 2025Yuxiang Huang, Pengjie Wang, Jicheng Han +9KV-Cache OffloadingGPU Acceleration

  4. ReFreeKV: Towards Threshold-Free KV Cache Compression

    Feb 24, 2025Xuanfan Ni, Liyan Xu, Chenyang Lyu +6KV CachingMemory-Efficient Optimization

  5. UltraQuant: 4-bit KV Caching for Context-Heavy Agents

    Date pendingInesh Chakrabarti, David Limpus, Aditi Ghai Rana +4LLM Inference EfficiencyGPU Kernel Optimization