Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework that compensates KV cache compression methods by carrying part of the context in the model's weights. PatchKV pairs an off-the-shelf compressed KV cache with a context-specific weight patch, which is computed once at context-loading time and served for downstream queries for the context. The weight patch is derived in closed form via ridge regression, by aligning the block-wise activations of context-derived reference query tokens under the full cache and the compressed cache. Once merged into the model, the patch leaves the forward graph and per-query inference cost unchanged in the single-context, multi-query setting. Across long-context QA (SCBench with up to 170K tokens, SQuAD, NIAH) and math (GSM8K) benchmarks on three model architectures, PatchKV consistently improves cache compression methods, suggesting an alternative direction to compensate them at aggressive budgets.
Figures & tables
Figure 1: Overview of PatchKV. (a) Standard KV cache compression stores context information only in a compressed cache. (b) PatchKV pairs the compressed cache with a context-specific patch to the MLP output projections, compensating for part of the full-cache behavior lost during compression.
Figure 2: Benchmark results using Qwen2.5-7B-1M with PatchKV applied on top of (a) an eviction based method (KVzip) and (b) an approximation based method (Attention Matching). In (b) , † marks shortened (tiny) task variants for SCBench retrieval tasks.
Ratio
Full KV
None
repeat-context
self-study
joint
ref.
test
ref.
test
ref.
test
ref.
test
ref.
test
0.20
1.00
8.79
1.00
18.16
1.00
12.24
1.00
17.01
1.00
17.07
0.10
1.00
8.79
1.10
17.16
1.00
14.03
1.03
16.51
1.00
16.60
0.05
1.00
8.79
1.81
123.36
1.00
21.08
1.10
18.90
1.00
18.84
0.02
1.00
8.79
6.18
180.63
1.00
32.82
1.39
28.34
1.00
30.74
Table 1: Perplexity ( ↓ ) on SQuAD across compression ratios. ref. stands for reference query.
Figure 3: Qualitative example of PatchKV over KVzip on SQuAD. Evicted tokens shown in red.
Eval cache
Patch
CA
CB
None
54.95
57.74
ΔWA
85.39
66.60
ΔWB
65.63
88.09
Table 2: Cross-evaluation of patches built from disjoint cache subsets. Rows: weight patch applied at inference. Columns: compressed cache used at inference.
Figure 4: Multi-task results on RepoQA+KV and Summary+NIAH using Qwen2.5-7B-1M.
Chunk
1K
2K
5K
8K
10K
Rel. Acc.
.696
.695
.696
.672
.679
Time (s)
114
98
85
86
81
Mem. (GB)
30.8
31.6
33.0
34.1
35.9
Table 3: Effect of chunk size.
Ratio
Trace
Weighted
0.2
0.92
0.93
0.1
0.88
0.91
0.05
0.72
0.72
0.02
0.45
0.46
0.01
0.32
0.36
Table 4: Ridge scaling results.
Figure 5: Averaged score across datasets, normalized by the full-KV score, for different models.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
ti←Wℓ(hi−hi′)+(zi−zi′)
Appendix
Algorithm 1 PatchKV weight patch construction
Figure 6: FP32 vs. TF32 patch construction.
Figure 7: Additional benchmark results across KV cache budgets using Qwen2.5-7B-1M and KVzip.
Task
Method
0.20
0.10
0.05
0.02
0.01
Retr.KV
Attention Matching
0.0
0.0
0.0
0.0
0.0
+ PatchKV
0.0
0.0
0.0
0.0
0.0
Retr.Prefix-Suffix
Attention Matching
0.0
0.0
0.0
0.0
0.0
+ PatchKV
0.0
0.0
0.0
0.0
0.0
Code.RepoQA
Attention Matching
15.23
3.18
1.59
0.23
0.23
+ PatchKV
33.64
8.64
3.41
0.91
0.68
Appendix
Table 5: Score (%) on the original SCBench retrieval tasks using Qwen2.5-7B-1M. Both methods achieve zero accuracy on Retr.KV and Retr.Prefix-Suffix under the evaluated cache budgets.
Stage
Time (s)
Peak Memory (GB)
(i) Cache compression (KVzip)
67.82
24.89
(ii) Teacher forward
41.41
30.92
(iii) Student forward
16.38
30.99
(iv) Patch computation
11.40
32.98
(v) Student-state propagation
16.16
31.11
Total
153.17
32.98
Appendix
Table 6: Stagewise wall-clock time and peak GPU memory for PatchKV construction.
Figure 8: Performance across reference context lengths.
Figure 9: Performance when restricting the rank of weight patch of PatchKV.
KV cache size
Method
4K Acc.
8K Acc.
Full KV
Full
9/30
14/30
2048→1034
AM
9/30
13/30
AM + PatchKV
9/30
12/30
2048→527
AM
9/30
11/30
AM + PatchKV
8/30
13/30
2048→222
AM
2/30
4/30
Appendix
Table 7: Online evaluation on AIME 2025 using Qwen3-4B. A compaction event is triggered whenever the physical KV cache reaches 2048 entries, with the most recent 20 entries protected during compaction. 2048→m denotes the post-compaction KV cache size.
Figure 10: Performance of PatchKV on INT4/INT8 quantized KV cache.
Figure 11: Qualitative example of PatchKV over KVzip on SQuAD. Evicted tokens shown in red.
Figure 12: Qualitative example of PatchKV over KVzip on SQuAD. Evicted tokens shown in red.
Figure 13: Qualitative example of PatchKV over KVzip on GSM8K. Evicted tokens shown in red.
Figure 14: Qualitative example of PatchKV over KVzip on GSM8K. Evicted tokens shown in red.
Figure 15: Benchmark results using LLaMA3.1-8B with PatchKV applied on top of an eviction based method (KVzip).
Figure 16: Benchmark results using Qwen3-4B with PatchKV applied on top of an eviction based method (KVzip).
The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction methods permanently discard tokens, degrading performance whenever a discarded token later proves essential, while quantization methods retain all tokens at low precision but offer limited compression. We propose AnchorKV, a compression scheme that shrinks the cache by 20× without discarding a single token. AnchorKV represents the cache using a small set of anchors stored exactly, expresses every other token through its most similar anchor, and refines only those whose approximation most affects the model's output. AnchorKV consistently preserves accuracy across models and datasets, retaining 99% of the full-cache score at the 70B scale, while keeping the entire context at a fraction of its cost.
Malik Khalaf, Yara Shamshoum, Nitzan Hodos +2
Department of Computer Science, Technion - Israel Institute of Technology
Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware. Existing KV cache eviction methods typically apply heuristic token scoring over all heads in GQA-based LLMs. These methods ignore the different functionalities of attention heads, leading to the eviction of critical tokens and thus degrading the performance of LLMs. To address this issue, we propose CompressKV, a resource-efficient KV-cache compression framework for GQA-based LLMs. Instead of aggregating attention scores from all heads, CompressKV identifies Semantic Retrieval Heads (SRHs) that capture both the initial and final tokens of a prompt and semantically important mid-context evidence, and uses them to select tokens whose KV pairs should be retained. Furthermore, CompressKV allocates cache budgets across layers according to offline estimates of layer-wise eviction error. Experiments on LongBench and Needle-in-a-Haystack show that CompressKV consistently outperforms existing KV-cache eviction methods across memory budgets. Notably, it preserves over 97% of full-cache performance using only 3% of the KV cache on LongBench question-answering tasks and achieves 90% accuracy with just 0.7% KV storage on Needle-in-a-Haystack. These results demonstrate an improved resource--performance trade-off for long-context LLM inference. Our code is publicly available at: https://github.com/TUDa-HWAI/CompressKV
Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3
Technical University of Darmstadt, Darmstadt, Germany · University of Notre Dame, Notre Dame, IN, USA · Technical University of Ilmenau, Ilmenau, Germany
Large language models increasingly operate over long contexts, where the KV cache becomes a dominant memory bottleneck: its size grows linearly with sequence length and must be retained throughout decoding, making full GPU caching prohibitively expensive without compression. Existing KV cache compression methods struggle to balance efficiency with faithful context preservation. Token eviction discards information, while semantic grouping fixes compression decisions at prefill time; neither can recover token-level detail from a compressed span once it becomes relevant during generation. As a solution, we propose SeKV, a resolution-adaptive semantic KV cache that organizes context into entropy-guided semantic spans and stores them across a GPU-CPU memory hierarchy without discarding information. Each span keeps a lightweight summary vector on GPU for coarse routing and a low-rank SVD basis on CPU for on-demand token-level reconstruction. A trained zoom-in mechanism selectively expands query-relevant spans during decoding, enabling precise retrieval without materializing the full KV cache on GPU. SeKV enables adaptive token-level reconstruction while keeping the base LLM fully frozen and adding fewer than 0.05% trainable parameters. Across four benchmarks, SeKV improves over the strongest semantic compression baseline by 5.9% on average while reducing GPU memory by 53.3% versus full KV caching at 128K context. Code is available on https://github.com/AmirAbaskohi/SeKV.
Amirhossein Abaskohi, Giuseppe Carenini, Peter West +1
University of British Columbia · Microsoft Research