KV-Cache Pruning

Momentum

3 papers in the last four weeks, against 2 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 25

All topics
CardsList
  1. DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency

    Sep 30, 2026Zeqi Xiao, Qingle Liu, Kaiwen Zhang +3KV-Cache ManagementAutoregressive Video Generation

  2. Pixel Decodability Is Not a Compression Signal: Causally Evaluating Importance Proxies for Visual KV-Cache Eviction

    Sep 14, 2026Chenyu Zhou, Qiliang Jiang, Shuning Wu +1KV CachingKV-Cache Eviction

  3. Small Frequency Corrections Can Change What Survives KV Cache Compression

    Aug 27, 2026Hong Chen, Yudong Zeng, Yongwei Huang +5KV CachingKV-Cache Eviction

  4. StepKV: Step-Aware KV Cache Compression for LLM Agents

    Aug 26, 2026Boyu Feng, Jiahong Liu, Yifan Li +7KV CachingKV-Cache Compression

  5. CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents

    Aug 8, 2026Weizhong Huang, Jinchao Zhang, Xiawu ZhengKV CachingLLM Agent Memory

  6. QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding

    Aug 5, 2026Ayushman Garg, Akshita Gupta, Shaswata Bhattacharya +3KV CachingLong-Context Language Model Inference

  7. TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning

    Aug 4, 2026Wonpyo Park, Seung-won HwangLLM PruningKV Caching

  8. Back from the Future: Key-Value Cache Management by Counter-Causal Surprise

    Jul 30, 2026Stephen Gould, Anton van den HengelKV CachingKV-Cache Eviction

  9. HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

    Jul 24, 2026Chao Fang, Jun Yin, Man Shi +1KV CachingMemory-Efficient Optimization

  10. KVpop -- Key-Value Cache Compression with Predictive Online Pruning

    Jul 6, 2026Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl +5KV CachingKV-Cache Eviction

  11. EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix

    Jun 25, 2026Steven Kolawole, Virginia SmithKV CachingKV-Cache Eviction

  12. IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference

    Jun 6, 2026Junjie Li, Jiong Lou, Jie LiKV CachingKV-Cache Management

  13. Adaptive Mass-Segmented KV Compression for Long-Context Reasoning

    May 22, 2026Junzhe Yang, Xiaoyu ShenKV CachingLong-Context Language Model Inference

  14. HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling

    May 14, 2026Jonathan Cederlund, Axel Berg, William Isaksson +3KV CachingVisual Autoregressive Models

  15. Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility

    May 13, 2026Gergely Szilvasy, Manuel Faysse, Maria Lomeli +5Self-AttentionKV Caching

  16. Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

    May 10, 2026Ngoc Bui, Hieu Trung Nguyen, Arman Cohan +1Self-AttentionKV Caching

  17. ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference

    May 9, 2026Junjie Li, Jiong Lou, Jie LiLLM PruningKV Caching

  18. DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference

    Apr 27, 2026Zahra Dehghanighobadi, Asja FischerLLM PruningKV Caching

  19. ReFreeKV: Towards Threshold-Free KV Cache Compression

    Feb 24, 2025Xuanfan Ni, Liyan Xu, Chenyang Lyu +6KV CachingMemory-Efficient Optimization