KV-Cache Compression

Momentum

24 papers in the last four weeks, up 71% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 166

All topics
CardsList
  1. Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression

    May 21, 2026Wei Luo, Yi Huang, Songchen Ma +3KV CachingKV-Cache Eviction

  2. MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering

    May 21, 2026Junbin Xiao, Jiajun Chen, Tianxiang Sun +2KV CachingVideo QA

  3. EntmaxKV: Support-Aware Decoding for Entmax Attention

    May 20, 2026Gonçalo Duarte, Miguel Couceiro, Marcos V. TrevisoSelf-AttentionKV Caching

  4. Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation

    May 20, 2026Guotao Liang, Baoquan Zhang, Zhiyuan Wen +1KV CachingAutoregressive Image Generation

  5. Latent Cache Flow: Model-to-Model Communication Without Text

    May 19, 2026Maximillian Rossi, Prajwal Raghunath, Eugene WuKV CachingKV-Cache Compression

  6. VeriCache: Turning Lossy KV Cache into Lossless LLM Inference

    May 17, 2026Jiayi Yao, Samuel Shen, Kuntai Du +7KV-Cache OffloadingKV Caching

  7. KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy

    May 14, 2026Yingbing Huang, Tharun Adithya Srikrishnan, Steven K. Reinhardt +1Vision-Language ModelsEfficient VLM Inference

  8. HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling

    May 14, 2026Jonathan Cederlund, Axel Berg, William Isaksson +3KV CachingVisual Autoregressive Models

  9. CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding

    May 14, 2026Ailar Mahdizadeh, Puria Azadi, Muchen Li +2Streaming Video UnderstandingKV Caching

  10. Minimal-Intervention KV Retention via Set-Conditioned Diversity

    May 14, 2026Libo Sun, Po-wei Harn, Peixiong He +1KV CachingKV-Cache Eviction

  11. Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility

    May 13, 2026Gergely Szilvasy, Manuel Faysse, Maria Lomeli +5Self-AttentionKV Caching

  12. SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference

    May 13, 2026Anay Chauhan, Gurucharan Marthi Krishna Kumar, Arion Das +4KV CachingLong-Context Language Model Inference

  13. KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving

    May 13, 2026Zedong Liu, Xinyang Ma, Dejun Luo +9LLM ServingKV Caching

  14. How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment

    May 7, 2026Rui Zhu, Weiheng Bai, Qiushi Wu +3LLM AlignmentKV-Cache Compression

  15. Training Transformers for KV Cache Compressibility

    May 7, 2026Yoav Gelberg, Yam Eitan, Michael Bronstein +2Long-Context Language ModelingKV-Cache Compression

  16. SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving

    May 3, 2026Yipin Guo, Siddharth JoshiLLM ServingKV Caching

  17. Make Your LVLM KV Cache More Lightweight

    May 1, 2026Xihao Chen, Yangyang Guo, Roger ZimmermannKV CachingEfficient VLM Inference