Key-Value Cache Compression

Momentum

15 papers in the last four weeks, up 150% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 110

All topics
CardsList
  1. Behavior-Preserving KV Cache Compression

    Oct 5, 2026Doo Hwan Hwang, Junyoung Jang, Junho Na +2Key-Value Cache Compression

  2. DeferKV: Rethinking Eviction Timing for One-Shot KV Cache Compression

    Oct 5, 2026Zhe Wang, Jiakai Li, Yujia Sun +2Key-Value Cache Compression

  3. PatchKV: Weight-Space Compensation of KV Cache

    Sep 30, 2026Chanryeol Lee, Chanhyuk Lee, Yeonwoo Choi +2Key-Value Cache CompressionKey-Value Cache

  4. ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction

    Sep 29, 2026Zheyu Shen, Guanhua Wang, Dezhan Tu +5Key-Value Cache CompressionCache

  5. Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

    Sep 28, 2026Xingyu Zhu, Pu, Yi +6Key-Value Cache CompressionData Compression Methods

  6. TORQUE: Optimizing What (not) to Quantize Before and After Rotation

    Sep 28, 2026Ran Ben Basat, Michael Mitzenmacher, Shay VargaftikKey-Value Cache CompressionTop-K

  7. Tensor Decomposition of Transformer Key-Value Caches: Spectral Structure and Format Comparison

    Sep 23, 2026Rahul Krishnan, Volker SchulzKey-Value Cache CompressionTensor Decomposition

  8. Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference

    Sep 22, 2026Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia +2LLM Inference OptimizationLarge Language Model Memory

  9. KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation

    Sep 21, 2026Sihyeon Ha, Jaeho Lee, Yo-Seb JeonKey-Value Cache CompressionCache

  10. LLM Inference in a Flash!

    Sep 14, 2026Sebastian Zhao, Minseo Kim, Coleman Hooper +5LLM Inference OptimizationInference-Time

  11. MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

    Sep 7, 2026Michael Wang, Keith Li, Roozbeh BostandoostKey-Value Cache CompressionLLM Inference Optimization

  12. Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

    Sep 3, 2026Heng Wang, Jielin Qiu, Wenting Zhao +8Key-Value Cache EvictionKey-Value Cache

  13. SGD-KV: Summarization Guided KV Cache Compression

    Sep 3, 2026Zeyu Liu, Woomin Song, Xuandi Fu +5Key-Value Cache CompressionEfficient Long-Context Inference

  14. StepKV: Step-Aware KV Cache Compression for LLM Agents

    Aug 26, 2026Boyu Feng, Jiahong Liu, Yifan Li +7Key-Value Cache CompressionCache

  15. RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation

    Aug 9, 2026Dongjie Xu, Kai Qian, Julius +6Key-Value Cache CompressionKey-Value Cache

  16. RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

    Aug 8, 2026Anthony. Lui, Mohamed. Elsaied, N. P. SavaniMixture-Of-ExpertsKey-Value Cache Compression

  17. CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents

    Aug 8, 2026Weizhong Huang, Jinchao Zhang, Xiawu ZhengKey-Value Cache CompressionDepthweave-Kv

  18. AnchorKV: Anchor-Residual KV Cache Compression

    Aug 3, 2026Malik Khalaf, Yara Shamshoum, Nitzan Hodos +2Key-Value Cache CompressionKey-Value Cache

  19. RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

    Aug 2, 2026Changwoo Baek, Seungjun Shin, Kyeongbo KongKey-Value Cache CompressionKey-Value Cache Eviction

  20. Practical Online KV Cache Compaction for LLM Agents: An Empirical Study

    Aug 2, 2026Yujian Liu, Jiabao Ji, Li An +4Key-Value Cache CompressionCompaction

  21. S4^4R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

    Aug 1, 2026Jialong Han, You Wu, Kewei TuKey-Value Cache CompressionKey-Value Cache

  22. ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression

    Jul 31, 2026Yuhang Zhan, Lisi Chen, Shuo ShangKey-Value Cache CompressionCache

  23. DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation

    Jul 27, 2026Tan T. Nguyen, Quan V. DangKey-Value Cache CompressionCache

  24. A JoLT for the KV cache: Near-Lossless KV Cache Compression via Joint Rank-bit Allocation

    Jul 14, 2026Rahul Krishnan, Volker SchulzKey-Value Cache CompressionKey-Value Cache

  25. DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

    Jul 7, 2026Anna Cordoba, Adam Puente Tercero, Nerea Angulo Hijo +4Key-Value Cache CompressionDepthweave-Kv

  26. FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

    Jul 7, 2026Anna Córdoba, Adam Puente Tercero, Nerea Angulo Hijo +4Key-Value Cache CompressionDepthweave-Kv

  27. KVpop -- Key-Value Cache Compression with Predictive Online Pruning

    Jul 6, 2026Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl +5Key-Value Cache EvictionKey-Value Cache Compression

  28. The risk of KV cache compression

    Jul 1, 2026Lukas Haverbeck, Carmen Amo Alonso, Andres Felipe Posada-Moreno +2Key-Value Cache CompressionCache

  29. MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression

    Jul 1, 2026Sheng Qiang, Ruiwei Chen, Yinpeng Wu +5Key-Value Cache CompressionLarge Language Model Compression

  30. Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption

    Jul 1, 2026Xiaomeng Fu, Jia Li, Yiming Hu +5Autoregressive Video GenerationKey-Value Cache Compression

  31. SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference

    Jun 30, 2026Amirhossein Abaskohi, Giuseppe Carenini, Peter West +1Key-Value Cache CompressionKey-Value Cache

  32. HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression

    Jun 27, 2026Yuxuan Yang, Feiyang Ren, Bowen Zeng +4Key-Value Cache CompressionLLM Inference Optimization

  33. Information-Aware KV Cache Compression for Long Reasoning

    Jun 25, 2026Jushi Kai, Zhuiri Xiao, Alexandra Birch +1Key-Value Cache CompressionEfficient Long-Context Inference

  34. CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

    Jun 23, 2026Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3Key-Value Cache CompressionLLM Inference Optimization

  35. ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference

    Jun 18, 2026Yang Tan, Junlong Tong, Linan Yue +3Streaming VideoStreaming

  36. AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor

    Jun 16, 2026Ning Ni, Yingjie LaoKey-Value Cache CompressionLarge Language Model Safety

  37. Ablation, Statistical Inference, and Validation for KV-Cache Compression

    Jun 14, 2026Paolo D'Alberto, Ashish Siarasao, Elliott Delaye +1Key-Value Cache CompressionResidual Vector Quantization

  38. Exploring a Layer-Wise Design Space for KV Cache Eviction

    Jun 13, 2026Chao Fei, Kaihua Liang, Hanzhi Hu +4Key-Value Cache CompressionKey-Value Cache

  39. Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression

    Jun 10, 2026Tianhao Chen, Yuheng Wu, Kelu Yao +3Key-Value Cache CompressionMultimodal Large Language Models

  40. From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs

    Jun 8, 2026Zhanchao Xu, Haoyang Li, Qingfa Xiao +4LLM Inference OptimizationDynamic Sparse Attention

  41. EinSort: Sorting is All We Need for Tensorizing LLM

    Jun 7, 2026Toshiaki Koike-Akino, Jing Liu, Ye WangTensor NetworksLow-Rank Structure

  42. STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control

    Jun 7, 2026Priyansh Bhatnagar, Ashkan Moradifirouzabadi, Se-Hyun Yang +3Key-Value Cache CompressionKey-Value Cache