cs.CLOct 5, 2026

DeferKV: Rethinking Eviction Timing for One-Shot KV Cache Compression

Authors: Zhe Wang, Jiakai Li, Yujia Sun, Rongzheng Wang, Shuang Liang

Organizations: University of Electronic Science and Technology of China, Chengdu, China · Ubiquitous Intelligence and Trusted Services Key Laboratory of Sichuan Province, Chengdu, China

Abstract

Long-context large language models (LLMs) have demonstrated strong capabilities across a wide range of tasks, but the growing KV cache introduces substantial memory and inference overhead. Existing one-shot KV cache compression methods typically commit to irreversible eviction immediately after prefill, before any signal from actual generation becomes available. Our quantitative analysis shows that early queries from the actual generation stage provide attention signals that are more consistent with subsequent decode attention, with the largest single-step gain occurring at the prefill-decode boundary. Based on this observation, we propose DeferKV, which moves the eviction decision from the end of prefill to the first real decoding step and temporally combines prompt-side and decode-side observations, thereby better aligning KV importance estimation with subsequent generation requirements. DeferKV requires no additional training, draft model, or future-query prediction module, making it simple and easy to deploy. Experiments on LongBench, RULER, and Needle-in-a-Haystack demonstrate that DeferKV consistently improves model performance under KV cache compression while maintaining low inference latency.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation

    May 28, 2026Soumyadeep Jana, Sagar Nishad, Sanasam Ranbir SinghKey-Value Cache CompressionLarge Language Model Decoding

  2. CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

    Jun 23, 2026Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3Key-Value Cache CompressionLLM Inference Optimization

  3. Behavior-Preserving KV Cache Compression

    Oct 5, 2026Doo Hwan Hwang, Junyoung Jang, Junho Na +2Key-Value Cache Compression