Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework that compensates KV cache compression methods by carrying part of the context in the model's weights. PatchKV pairs an off-the-shelf compressed KV cache with a context-specific weight patch, which is computed once at context-loading time and served for downstream queries for the context. The weight patch is derived in closed form via ridge regression, by aligning the block-wise activations of context-derived reference query tokens under the full cache and the compressed cache. Once merged into the model, the patch leaves the forward graph and per-query inference cost unchanged in the single-context, multi-query setting. Across long-context QA (SCBench with up to 170K tokens, SQuAD, NIAH) and math (GSM8K) benchmarks on three model architectures, PatchKV consistently improves cache compression methods, suggesting an alternative direction to compensate them at aggressive budgets.
Figures & tables
Figure 1: Overview of PatchKV. (a) Standard KV cache compression stores context information only in a compressed cache. (b) PatchKV pairs the compressed cache with a context-specific patch to the MLP output projections, compensating for part of the full-cache behavior lost during compression.
Figure 2: Benchmark results using Qwen2.5-7B-1M with PatchKV applied on top of (a) an eviction based method (KVzip) and (b) an approximation based method (Attention Matching). In (b) , † marks shortened (tiny) task variants for SCBench retrieval tasks.
Ratio
Full KV
None
repeat-context
self-study
joint
ref.
test
ref.
test
ref.
test
ref.
test
ref.
test
0.20
1.00
8.79
1.00
18.16
1.00
12.24
1.00
17.01
1.00
17.07
0.10
1.00
8.79
1.10
17.16
1.00
14.03
1.03
16.51
1.00
16.60
0.05
1.00
8.79
1.81
123.36
1.00
21.08
1.10
18.90
1.00
18.84
0.02
1.00
8.79
6.18
180.63
1.00
32.82
1.39
28.34
1.00
30.74
Table 1: Perplexity ( ↓ ) on SQuAD across compression ratios. ref. stands for reference query.
Figure 3: Qualitative example of PatchKV over KVzip on SQuAD. Evicted tokens shown in red.
Eval cache
Patch
CA
CB
None
54.95
57.74
ΔWA
85.39
66.60
ΔWB
65.63
88.09
Table 2: Cross-evaluation of patches built from disjoint cache subsets. Rows: weight patch applied at inference. Columns: compressed cache used at inference.
Figure 4: Multi-task results on RepoQA+KV and Summary+NIAH using Qwen2.5-7B-1M.
Chunk
1K
2K
5K
8K
10K
Rel. Acc.
.696
.695
.696
.672
.679
Time (s)
114
98
85
86
81
Mem. (GB)
30.8
31.6
33.0
34.1
35.9
Table 3: Effect of chunk size.
Ratio
Trace
Weighted
0.2
0.92
0.93
0.1
0.88
0.91
0.05
0.72
0.72
0.02
0.45
0.46
0.01
0.32
0.36
Table 4: Ridge scaling results.
Figure 5: Averaged score across datasets, normalized by the full-KV score, for different models.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
ti←Wℓ(hi−hi′)+(zi−zi′)
Appendix
Algorithm 1 PatchKV weight patch construction
Figure 6: FP32 vs. TF32 patch construction.
Figure 7: Additional benchmark results across KV cache budgets using Qwen2.5-7B-1M and KVzip.
Task
Method
0.20
0.10
0.05
0.02
0.01
Retr.KV
Attention Matching
0.0
0.0
0.0
0.0
0.0
+ PatchKV
0.0
0.0
0.0
0.0
0.0
Retr.Prefix-Suffix
Attention Matching
0.0
0.0
0.0
0.0
0.0
+ PatchKV
0.0
0.0
0.0
0.0
0.0
Code.RepoQA
Attention Matching
15.23
3.18
1.59
0.23
0.23
+ PatchKV
33.64
8.64
3.41
0.91
0.68
Appendix
Table 5: Score (%) on the original SCBench retrieval tasks using Qwen2.5-7B-1M. Both methods achieve zero accuracy on Retr.KV and Retr.Prefix-Suffix under the evaluated cache budgets.
Stage
Time (s)
Peak Memory (GB)
(i) Cache compression (KVzip)
67.82
24.89
(ii) Teacher forward
41.41
30.92
(iii) Student forward
16.38
30.99
(iv) Patch computation
11.40
32.98
(v) Student-state propagation
16.16
31.11
Total
153.17
32.98
Appendix
Table 6: Stagewise wall-clock time and peak GPU memory for PatchKV construction.
Figure 8: Performance across reference context lengths.
Figure 9: Performance when restricting the rank of weight patch of PatchKV.
KV cache size
Method
4K Acc.
8K Acc.
Full KV
Full
9/30
14/30
2048→1034
AM
9/30
13/30
AM + PatchKV
9/30
12/30
2048→527
AM
9/30
11/30
AM + PatchKV
8/30
13/30
2048→222
AM
2/30
4/30
Appendix
Table 7: Online evaluation on AIME 2025 using Qwen3-4B. A compaction event is triggered whenever the physical KV cache reaches 2048 entries, with the most recent 20 entries protected during compaction. 2048→m denotes the post-compaction KV cache size.
Figure 10: Performance of PatchKV on INT4/INT8 quantized KV cache.
Figure 11: Qualitative example of PatchKV over KVzip on SQuAD. Evicted tokens shown in red.
Figure 12: Qualitative example of PatchKV over KVzip on SQuAD. Evicted tokens shown in red.
Figure 13: Qualitative example of PatchKV over KVzip on GSM8K. Evicted tokens shown in red.
Figure 14: Qualitative example of PatchKV over KVzip on GSM8K. Evicted tokens shown in red.
Figure 15: Benchmark results using LLaMA3.1-8B with PatchKV applied on top of an eviction based method (KVzip).
Figure 16: Benchmark results using Qwen3-4B with PatchKV applied on top of an eviction based method (KVzip).
Technical University of Darmstadt, Darmstadt, Germany · University of Notre Dame, Notre Dame, IN, USA · Technical University of Ilmenau, Ilmenau, Germany