KV Caching

KV: Key-Value

Momentum

29 papers in the last four weeks, up 16% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 276

All topics
CardsList
  1. CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding

    May 14, 2026Ailar Mahdizadeh, Puria Azadi, Muchen Li +2Streaming Video UnderstandingKV Caching

  2. Minimal-Intervention KV Retention via Set-Conditioned Diversity

    May 14, 2026Libo Sun, Po-wei Harn, Peixiong He +1KV CachingKV-Cache Eviction

  3. Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility

    May 13, 2026Gergely Szilvasy, Manuel Faysse, Maria Lomeli +5Self-AttentionKV Caching

  4. SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference

    May 13, 2026Anay Chauhan, Gurucharan Marthi Krishna Kumar, Arion Das +4KV CachingLong-Context Language Model Inference

  5. KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving

    May 13, 2026Zedong Liu, Xinyang Ma, Dejun Luo +9LLM ServingKV Caching

  6. Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

    May 12, 2026Shaoke Fang, Ziang Li, Wenfei Wu +3LLM Inference EfficiencyKV Caching

  7. KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference

    May 12, 2026Alireza Nadali, Patrick Cooper, Ashutosh Trivedi +1Transformer InferenceKV Caching

  8. KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving

    May 10, 2026Zhiqing Zhong, Zhijing Ye, Jian Zhang +3KV CachingKV-Cache Management

  9. Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

    May 10, 2026Ngoc Bui, Hieu Trung Nguyen, Arman Cohan +1Self-AttentionKV Caching

  10. Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning

    May 10, 2026Aojie Yuan, Tianqi Shen, Dajun ZhangKV-Cache OffloadingKV Caching

  11. PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving

    May 10, 2026Bing Xie, Zhipeng Wang, Masahiro Tanaka +1LLM ServingKV Caching

  12. ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference

    May 9, 2026Junjie Li, Jiong Lou, Jie LiLLM PruningKV Caching

  13. Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes

    May 9, 2026Willy Fitra HendriaLLM InferenceKV Caching

  14. ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing

    May 9, 2026Yongqi An, Chang Lu, Kuan Zhu +5KV CachingLong-Context Language Model Inference

  15. PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design

    May 9, 2026Xingyu Qu, Tianhao Lin, Yiqi Li +2LLM ServingKV Caching

  16. Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache

    May 7, 2026Mohsen Dehghankar, Abolfazl AsudehSelf-AttentionKV Caching

  17. Efficient Serving for Dynamic Agent Workflows with Prediction-based KV-Cache Management

    May 7, 2026Haoyu Zheng, Fangcheng Fu, Jia Wu +6LLM Inference EfficiencyKV Caching

  18. Towards Distributed Inference of LLMs on a P2P Network

    May 7, 2026Shabari S Nair, Krishanu SainiLLM InferenceKV Caching

  19. Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs

    May 7, 2026Andy Zeyi Liu, Michael Zhang, Ilana Greenberg +3Language Model SteeringMemory-Augmented Language Models

  20. Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving

    May 7, 2026Bole Ma, Jan Eitzinger, Harald KöstlerKV CachingLLM Inference Serving

  21. One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving

    May 6, 2026Wenjun Yu, Shuguang Han, Amelie Chi ZhouLLM Inference EfficiencyGPU Acceleration