EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction
Organizations: EPFL
Abstract
KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieves strong compression quality at the cost of additional forward passes. Learned approximations reduce this cost but require model-specific training. We analyze how KVzip identifies important cached information and show how to approximate its reconstruction scores using information already computed during prefill. These findings motivate EchoPress, a training-free method that approximates reconstruction attention using queries and keys from standard prefill. For each request, it reconstructs only the first chunk to calibrate importance scores for the remaining context. Experiments on LongBench and RULER with Qwen3-8B and Llama-3.1-8B-Instruct show that EchoPress matches KVzip in task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by a factor of 1.7-19.6 and total prefill time by a factor of up to 2.9. Code is available at https://github.com/ljwljwljwljw/kvpress/tree/echo-press.
Figures & tables
| Scoring and calibration | RULER Avg. | NIAH (8) |
|---|---|---|
| Exact first chunk, no calibration | 71.0 | 72.4 |
| Per-request, per-layer/head exact virtual | 50.8 | 43.8 |
| Per-request, global virtual exact | 77.3 | 83.5 |
| EchoPress: per-request, per-layer/head virtual exact | 87.6 | 95.7 |
| KVzip | 87.5 | 92.5 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Qwen3-8B | ||||||||
|---|---|---|---|---|---|---|---|---|
| Method | Single QA | Multi QA | Summ. | Few- shot | Retr. | Code | Avg. | |
| Uncompressed | 0 | 42.4 | 49.2 | 29.1 | 40.7 | 91.4 | 64.5 | 47.5 |
| KVzip | 0.5 | 42.4 | 48.4 | 29.2 | 46.4 | 94.0 | 64.7 | 48.5 |
| EchoPress | 0.5 | 42.4 | 48.7 | 29.1 | 44.0 | 95.0 | 63.4 | 48.2 |
| KVzip | 0.75 | 41.7 | 49.0 | 28.7 | 56.5 | 92.0 | 63.1 | 49.8 |
| EchoPress | 0.75 | 42.1 | 50.0 | 29.0 | 56.6 | 93.6 | 54.0 | 49.6 |
| Qwen3-8B | |||||||
|---|---|---|---|---|---|---|---|
| Method | NIAH (8) | VT | CWE | FWE | QA (2) | Avg. | |
| Uncompressed | 0 | 100.0 | 100.0 | 98.9 | 95.4 | 72.9 | 95.4 |
| KVzip | 0.5 | 100.0 | 100.0 | 99.0 | 95.9 | 71.4 | 95.2 |
| EchoPress | 0.5 | 100.0 | 100.0 | 99.1 | 95.7 | 71.7 | 95.2 |
| KVzip | 0.75 | 99.9 | 100.0 | 98.9 | 96.4 | 70.8 | 95.1 |
| EchoPress | 0.75 | 99.8 | 100.0 | 98.0 | 95.9 | 70.5 | 94.9 |
| Question answering | |||||||
|---|---|---|---|---|---|---|---|
| Method | Qasper | MultiFQA | NrtvQA | HotpotQA | 2WikiMQA | MuSiQue | |
| Uncompressed | 0 | 44.8 | 55.6 | 26.7 | 62.9 | 49.2 | 35.4 |
| KVzip | 0.5 | 44.5 | 55.8 | 26.8 | 62.8 | 47.8 | 34.7 |
| EchoPress | 0.5 | 44.6 | 56.4 | 26.1 | 62.9 | 48.8 | 34.5 |
| KVzip | 0.75 | 42.6 | 55.0 | 27.6 | 63.8 | 47.8 | 35.3 |
| EchoPress | 0.75 | 43.4 | 54.4 | 28.4 | 63.7 | 48.9 | 37.3 |
| Question answering | |||||||
|---|---|---|---|---|---|---|---|
| Method | Qasper | MultiFQA | NrtvQA | HotpotQA | 2WikiMQA | MuSiQue | |
| Uncompressed | 0 | 47.1 | 55.3 | 28.9 | 59.5 | 51.8 | 32.7 |
| KVzip | 0.5 | 47.3 | 56.8 | 29.1 | 55.8 | 48.2 | 28.6 |
| EchoPress | 0.5 | 47.3 | 57.0 | 28.7 | 57.4 | 49.7 | 29.9 |
| KVzip | 0.75 | 44.6 | 54.8 | 27.9 | 56.1 | 46.4 | 27.4 |
| EchoPress | 0.75 | 46.0 | 55.6 | 29.7 | 54.2 | 46.3 | 28.0 |
| Needle retrieval | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | S1 | S2 | S3 | MK1 | MK2 | MK3 | MQ | MV | |
| Uncompressed | 0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.9 | 100.0 |
| KVzip | 0.5 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 |
| EchoPress | 0.5 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 |
| KVzip | 0.75 | 100.0 | 100.0 | 100.0 | 99.8 | 100.0 | 99.8 | 100.0 | 99.9 |
| EchoPress | 0.75 | 100.0 | 100.0 | 100.0 | 99.6 | 100.0 | 99.2 | 100.0 | 100.0 |
| Needle retrieval | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | S1 | S2 | S3 | MK1 | MK2 | MK3 | MQ | MV | |
| Uncompressed | 0 | 100.0 | 100.0 | 99.8 | 99.8 | 100.0 | 99.8 | 99.9 | 99.9 |
| KVzip | 0.5 | 99.8 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 | 100.0 | 100.0 |
| EchoPress | 0.5 | 99.8 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 | 100.0 | 100.0 |
| KVzip | 0.75 | 100.0 | 100.0 | 99.8 | 100.0 | 100.0 | 100.0 | 99.9 | 99.9 |
| EchoPress | 0.75 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 | 100.0 | 100.0 |
| Qwen3-8B | Llama-3.1-8B | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Budget | Full | KVzip | EchoPress | Full | KVzip | EchoPress | |||
| 32K | 0.75 | 87.88 | 86.98 | 86.19 | 86.38 | 87.47 | 86.31 | ||
| 32K | 0.9 | 87.88 | 73.99 | 71.83 | 86.38 | 74.12 | 76.65 | ||
| 64K | 0.75 | 77.72 | 77.75 | 78.08 | 85.22 | 81.13 | 83.85 | ||
| 64K | 0.9 | 77.72 | 69.11 | 69.87 | 85.22 | 65.79 | 65.94 | ||
| Setting | Configuration |
|---|---|
| Data | NVIDIA/RULER v1 ( c3f5e3b4f87f ), base template, seed 42, model-specific tokenizers; no truncation. |
| Lengths | Budgets of 32,768 / 65,536 tokens include output allowances; actual prefill prefixes span 28,057–32,718 / 60,224–65,478 tokens. |
| Runtime | bf16, SDPA, batch size one; PyTorch 2.8.0 (CUDA 12.8), transformers 5.2.0, Triton 3.4.0. |
| Positions | Qwen3: fixed YaRN factor four [ Peng et al., 2024 ] , original context 32,768, checkpoint RoPE base. Llama: released RoPE configuration. Both use a maximum position limit of 131,072. |
| Decoding | Greedy with official per-task output limits; Qwen thinking disabled. |
| Compression | The 2,048-token accuracy configuration, global KV-pair allocation, and compression before the question. Question and generated tokens remain uncompressed. EchoPress uses exact first-chunk scores and per-request, per-layer/KV-head calibration, without an offline map. |
| Setting | Accuracy | Timing |
|---|---|---|
| Engine / attention | KVPress / SDPA | KVzip / FlashAttention-2 |
| Chunk budget | First chunk: 2,048 tokens including prompt; later chunks: up to 2,048 | 2,000 context tokens plus repeat prompt |
| Prompts / prefix | KVPress chat/prefix settings | KVzip prompts and prefix/sinks |
| Context coverage | RULER 4K/32K/64K; LongBench 32K | 4K–128K, excluding protected prefix |
| Output | Task scores | Compacted cache; no generation |
| Qwen3-8B | |||||||
|---|---|---|---|---|---|---|---|
| Context | Plain 16K-chunk | Plain Full K | Plain Full E | KVzip 16K-chunk | KVzip full prefill | EchoPress | Speedup |
| 4K | 0.36 | 0.36 | 0.36 | 0.77 | 0.78 | 0.46 | 1.7 |
| 8K | 0.73 | 0.73 | 0.72 | 1.63 | 1.62 | 0.57 | 2.9 |
| 16K | 1.67 | 1.63 | 1.61 | 3.68 | 3.68 | 0.79 | 4.7 |
| 32K | 4.11 | 3.96 | 3.91 | 9.03 | 9.04 | 1.23 | 7.3 |
| 64K | 11.23 | 10.90 | 10.78 | 24.53 | 24.47 | 2.14 | 11.5 |
| Qwen3-8B | Llama-3.1-8B | |||||
|---|---|---|---|---|---|---|
| Context | KVzip 16K-chunk | KVzip full | EchoPress | KVzip 16K-chunk | KVzip full | EchoPress |
| 4K | 1.13 | 1.14 | 0.81 | 1.01 | 1.01 | 0.74 |
| 8K | 2.36 | 2.36 | 1.29 | 2.09 | 2.09 | 1.17 |
| 16K | 5.34 | 5.31 | 2.39 | 4.73 | 4.69 | 2.17 |
| 32K | 13.14 | 12.99 | 5.15 | 11.60 | 11.47 | 4.64 |
| 64K | 35.76 | 35.37 | 12.92 | 31.61 | 31.35 | 11.57 |
| Setting | Configuration |
|---|---|
| Input / prefill | Saved concatenated LongBench gov_report contexts; three-token chat prefix; 131,075 total tokens; one full-context prefill. |
| Software | PyTorch 2.8.0 (CUDA 12.8), transformers 5.2.0, Triton 3.4.0; SDPA. |
| RoPE / precision | Fixed YaRN factor four, original size 32,768, original RoPE base. bf16 model/KVzap weights; released mixed bf16/fp32 Fast KVzip gates. |
| KVzap-MLP | nvidia/KVzap-mlp-Qwen3-8B , revision bd5c59178466 . Replace fixed threshold with global fixed-budget selection; protect the final 128 tokens. |
| Fast KVzip | Jang-Hyun/Fast-KVzip , revision 67ce1265f510 , q4_dim16_sink16.pt gates. Global allocation; promote first four and last 4,096 token scores. KVPress full prefill replaces the released implementation’s eviction between prefill chunks. |
| Repetitions | Five seeded, interleaved repetitions after one warm-up per method; plain-prefill baseline in each repetition; CUDA synchronized at both boundaries. |
| Measurement | Qwen3-8B | Llama-3.1-8B |
|---|---|---|
| Query difference, first later | ||
| Query difference, prefix Full | ||
| Mean within-chunk Spearman | ||
| Median within-chunk Spearman |