cs.LGSep 30, 2026

EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction

Authors: Jiawei Lin, Saibo Geng, Thomas Bourgeat

Organizations: EPFL

Abstract

KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieves strong compression quality at the cost of additional forward passes. Learned approximations reduce this cost but require model-specific training. We analyze how KVzip identifies important cached information and show how to approximate its reconstruction scores using information already computed during prefill. These findings motivate EchoPress, a training-free method that approximates reconstruction attention using queries and keys from standard prefill. For each request, it reconstructs only the first chunk to calibrate importance scores for the remaining context. Experiments on LongBench and RULER with Qwen3-8B and Llama-3.1-8B-Instruct show that EchoPress matches KVzip in task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by a factor of 1.7-19.6 and total prefill time by a factor of up to 2.9. Code is available at https://github.com/ljwljwljwljw/kvpress/tree/echo-press.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Behavior-Preserving KV Cache Compression

    Oct 5, 2026Doo Hwan Hwang, Junyoung Jang, Junho Na +2Key-Value Cache Compression

  2. AnchorKV: Anchor-Residual KV Cache Compression

    Aug 3, 2026Malik Khalaf, Yara Shamshoum, Nitzan Hodos +2Key-Value Cache CompressionKey-Value Cache