KV-Cache Eviction

Momentum

9 papers in the last four weeks, up 80% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 65

All topics
CardsList
  1. MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference

    Jul 12, 2026Venkatesha Matam, Keon KimKV CachingKV-Cache Eviction

  2. KVpop -- Key-Value Cache Compression with Predictive Online Pruning

    Jul 6, 2026Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl +5KV CachingKV-Cache Eviction

  3. EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix

    Jun 25, 2026Steven Kolawole, Virginia SmithKV CachingKV-Cache Eviction

  4. CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

    Jun 23, 2026Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3LLM Inference EfficiencyKV Caching

  5. Forget Without Compromise: Nexus Sampling for Streaming KV-Cache Eviction Under Fixed Budgets

    Jun 22, 2026Duc Duong, Hoang Anh Duy Le, Jianwen Xie +2KV CachingLong-Context Language Model Inference

  6. Recency/Frequency Adaptive KV Caching for Large Language Model Serving

    Jun 19, 2026Yang Shen, Meghana Madhyastha, Robert Underwood +2LLM ServingKV Caching

  7. Exploring a Layer-Wise Design Space for KV Cache Eviction

    Jun 13, 2026Chao Fei, Kaihua Liang, Hanzhi Hu +4KV CachingKV-Cache Eviction

  8. Value-Aware Stochastic KV Cache Eviction for Reasoning Models

    Jun 2, 2026Ting-Yun Chang, Harvey Yiyun Fu, Deqing Fu +3LLM Inference EfficiencyKV Caching

  9. TGV-KV: Text-Grounded KV Eviction for Vision-Language Models

    Jun 2, 2026Jizhihui Liu, Ruizi Han, Miao Zhang +4Vision-Language ModelsEfficient VLM Inference

  10. MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference

    Jun 1, 2026Yu Li, Binxu Li, Tian LanKV CachingLong-Context Language Model Inference

  11. Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation

    May 28, 2026Soumyadeep Jana, Sagar Nishad, Sanasam Ranbir SinghKV CachingLong-Context Language Modeling

  12. IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference

    May 25, 2026Xintong Yang, Hao Gu, Binxing Xu +6Memory-Augmented Language ModelsSelf-Attention

  13. A Simple Plug-in for Improving Eviction-Based KV Cache Compression

    May 22, 2026Yuping Lin, Jiayuan Ding, Yue Xing +3KV CachingMemory-Efficient Optimization

  14. Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression

    May 21, 2026Wei Luo, Yi Huang, Songchen Ma +3KV CachingKV-Cache Eviction

  15. ArborKV: Structure-Aware KV Cache Management for Scaling Tree-based LLM Reasoning

    May 21, 2026Yeqiu Chen, Ziyan Liu, Zhenxin Huang +3Inference-Time SearchTree of Thoughts

  16. Tensor Cache: Eviction-conditioned Associative Memory for Transformers

    May 21, 2026Kabir Swain, Sijie Han, Daniel Karl I. Weidele +2Memory-Augmented Neural NetworksSelf-Attention

  17. Minimal-Intervention KV Retention via Set-Conditioned Diversity

    May 14, 2026Libo Sun, Po-wei Harn, Peixiong He +1KV CachingKV-Cache Eviction

  18. Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

    May 12, 2026Shaoke Fang, Ziang Li, Wenfei Wu +3LLM Inference EfficiencyKV Caching

  19. Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

    May 10, 2026Ngoc Bui, Hieu Trung Nguyen, Arman Cohan +1Self-AttentionKV Caching

  20. PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving

    May 10, 2026Bing Xie, Zhipeng Wang, Masahiro Tanaka +1LLM ServingKV Caching

  21. ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference

    May 9, 2026Junjie Li, Jiong Lou, Jie LiLLM PruningKV Caching

  22. ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing

    May 9, 2026Yongqi An, Chang Lu, Kuan Zhu +5KV CachingLong-Context Language Model Inference

  23. Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective

    Apr 28, 2026Jiaming Yang, Chenwei Tang, Liangli Zhen +1KV CachingLong-Context Language Model Inference