KV Caching

KV: Key-Value

Momentum

29 papers in the last four weeks, up 16% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 276

All topics
CardsList
  1. Stochastic Sparse Attention for Memory-Bound Inference

    May 3, 2026Kyle Lee, Corentin Delacour, Kevin Callahan-Coray +5GPU AccelerationSelf-Attention

  2. Decouple and Cache: KV Cache Construction for Streaming Video Understanding

    May 3, 2026Zhanzhong Pang, Dibyadip Chatterjee, Fadime Sener +1Streaming Video UnderstandingKV Caching

  3. SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving

    May 3, 2026Yipin Guo, Siddharth JoshiLLM ServingKV Caching

  4. Make Your LVLM KV Cache More Lightweight

    May 1, 2026Xihao Chen, Yangyang Guo, Roger ZimmermannKV CachingEfficient VLM Inference

  5. Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving

    Apr 29, 2026Zihan Zhao, Baotong Lu, Shengjie Lin +8LLM ServingKV-Cache Offloading

  6. Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective

    Apr 28, 2026Jiaming Yang, Chenwei Tang, Liangli Zhen +1KV CachingLong-Context Language Model Inference

  7. DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference

    Apr 27, 2026Zahra Dehghanighobadi, Asja FischerLLM PruningKV Caching

  8. Kwai Summary Attention Technical Report

    Apr 27, 2026Chenglong Chu, Guorui Zhou, Guowang Zhang +35Self-AttentionKV Caching

  9. MTServe: Efficient Serving for Generative Recommendation Models with Hierarchical Caches

    Apr 24, 2026Xin Wang, Chi Ma, Shaobin Chen +14KV-Cache OffloadingKV Caching

  10. Sub-Token Routing for KV Cache Compression

    Apr 23, 2026Wei Jiang, Wei WangLLM Inference EfficiencyKV Caching

  11. LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction

    Apr 22, 2026Enshuai Zhou, Yifan Hao, Chao Wang +7KV CachingLong-Context Language Model Inference

  12. RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory

    Apr 22, 2026Fei Zuo, Zikang Zhou, Hao Cong +2LLM Inference EfficiencyMixed-Precision Quantization

  13. MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression

    Apr 20, 2026Libo Sun, Peixiong He, Po-Wei Harn +1LLM InferenceKV Caching

  14. StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression

    Apr 16, 2026Xuanyi Liu, Chunan Yu, Deyi Ji +63D ViTsStreaming 3D Reconstruction

  15. MemoSight: Unifying Context Compression and Multi Token Prediction for Reasoning Acceleration

    Apr 16, 2026Xinyu Liu, Xin Liu, Bo Jin +8Multi-Token PredictionKV Caching

  16. When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration

    Apr 14, 2026Yiping Li, Zhiyu An, Wan DuLLM CompressionKV Caching

  17. KV Cache Offloading for Context-Intensive Tasks

    Apr 9, 2026Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev +2KV-Cache OffloadingKV Caching

  18. InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context

    Mar 5, 2026Xin Teng, Canyu Zhang, Shaoyi Zheng +3Retrieval-Augmented GenerationKV Caching

  19. Learning to Evict from Key-Value Cache

    Feb 10, 2026Luca Moschella, Laura Manduchi, Ozan SenerKV CachingLong-Context Language Model Inference

  20. Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference

    Feb 9, 2026Yifei Gao, Lei Wang, Rong-Cheng Tu +3Self-AttentionKV Caching