Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely. We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion. Motivated by this value-side dispersion, we present ValueDiff, a value-geometric eviction that ranks tokens by the L2 deviation of their value vectors from the cache mean. The same score arises as the minimal-disturbance eviction under a max-entropy assumption about future attention. We evaluate under fixed cache budgets, with eviction at every block boundary during prefill and at every decoding step during generation. On RULER at a tight 2k token budget, ValueDiff retains 88-99% of dense across seven sink-suppressed models (best on 6 out of 7). On LongBench at the 4k budget, ValueDiff averages 92% retention across sink-suppressed models versus 83% for the strongest prior baseline. On MATH-500, ValueDiff is the strongest non-dense method on every sink-suppressed model tested at the 25% cache budget, outperforming prior methods by up to ~20 points on gated-attention models. Across all three benchmarks, value geometry emerges as the more reliable query-invariant eviction signal for sink-suppressed models.
Figures & tables
Figure 1 : Sink rate versus σV/σK . Color and marker encode suppression category; shaded band = 95% conditional-mean CI; region above σV/σK=1 marks value-dominant geometry
Sink
Attn
Key Geometry
Attn × Val.
Value Geometry
Model
Bud.
Streaming
TOVA
SnapKV
KNorm
KeyDiff
ManifKV
F-CAOTE
V-Norm
V-Dir
ValueDiff
Standard
Llama 3.1-8B
2k
40.1
62.0
75.5
91.8
89.6
60.9
78.8
53.1
88.7
88.1
4k
60.4
85.6
85.0
95.5
95.2
84.9
90.0
82.9
94.1
95.0
Llama 3.2-3B
2k
40.2
59.7
72.2
79.6
92.9
40.5
72.3
44.5
89.3
88.8
4k
60.6
82.4
82.3
94.1
97.4
89.9
87.9
78.7
96.5
94.7
Table 1 : RULER accuracy retention (% of dense, ↑ ; ctx=8192). Each model has two budget rows (2048, 4096). Bold marks the best non-dense strategy per budget, and underline marks the second best.
Sink
Attn
Key Geometry
Attn × Val.
Value Geometry
Model
Streaming
TOVA
SnapKV
KeyNorm
KeyDiff
ManifoldKV
F-CAOTE
V-Norm
V-Dir
ValueDiff
Standard
Llama3.1-8B
55.0 / 74.4
69.0 / 88.4
70.6 / 91.0
73.5 / 90.3
82.6 / 93.6
71.7 / 91.9
71.4 / 90.0
65.1 / 85.2
79.7 / 91.5
73.1 / 88.0
Llama3.2-3B
70.0 / 79.4
74.8 / 85.5
75.9 / 85.6
76.6 / 86.3
85.4 / 89.2
76.9 / 86.8
76.4 / 85.3
69.6 / 82.2
79.7 / 87.5
75.9 / 85.1
QK-normalization
Qwen3-4B
78.0 / 89.4
81.3 / 92.0
83.8 / 93.7
71.9 / 87.4
85.0 / 94.6
89.5 / 97.1
84.0 / 94.9
95.2 / 98.8
91.2 / 97.4
95.3 / 98.7
Table 2: LongBench accuracy retention (% of dense, ↑ ) at 4k / 8k budgets. Bold marks the best non-dense strategy by average retention over the two reported budgets for each model.
Sink
Attn
Key Geometry
Attn × Val.
Value Geometry
Model
Streaming
TOVA
SnapKV
KeyNorm
KeyDiff
ManifoldKV
F-CAOTE
V-Norm
V-Dir
ValueDiff
Standard
DS-R1-Qwen-1.5B
74.2 / 89.7
70.3 / 88.8
81.1 / 91.1
45.7 / 75.4
89.2 / 99.0
95.9 / 97.1
75.1 / 93.1
85.4 / 94.5
79.2 / 93.1
89.0 / 97.6
DS-R1-Llama-8B
63.1 / 73.9
69.8 / 79.1
76.0 / 83.7
64.3 / 79.1
79.4 / 84.9
77.0 / 84.9
74.1 / 84.9
77.7 / 83.2
75.8 / 84.7
77.9 / 85.1
QK-normalization
Gemma3-4B
81.1 / 94.6
94.3 / 104.9
88.1 / 98.1
75.5 / 90.0
87.1 / 97.8
87.3 / 100.0
95.4 / 103.2
92.7 / 101.3
91.9 / 100.5
96.8 / 102.2
Table 3: MATH-500 reasoning retention (% of dense flex@5 except Qwen3.5; pass@1 greedy; ↑ ; n=100 ) at 25% / 50% cache budgets, set per model as fractions of average sequence length (see Appendix F ). Bold marks the best non-dense strategy by average retention over the two budgets. † Sink on Qwen3.5-4B collapses output formatting; see Appendix G .
Figure 2 : bos-cosim vs. σV/σK for all models, colored by ValueDiff − KeyDiff gap, across RULER (left), LongBench (center), and MATH-500 (right).
Figure 3 : Peak GPU memory vs. context length on Qwen3.5-4B and GPT-OSS-20B (H100 80 GB, budget = 2048). Dense scales linearly; eviction caps at O(budget) . All eviction methods overlap within 0.1%; ValueDiff is representative eviction curve.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Mechanism
Sink rate
σK
σV
σV/σK
Llama 3.1-8B
None
0.801
1.783
0.583
0.33
Llama 3.2-3B
None
0.835
1.478
0.584
0.40
Qwen2.5-3B
None
0.446
1.561
1.851
1.19
Qwen2.5-7B
None
0.392
1.597
2.133
1.34
Gemma2-2B
Softcapping
0.464
1.807
1.590
0.88
Gemma2-9B
Softcapping
0.542
1.808
1.808
1.00
Appendix
Table 4: Sink rate and K/V L2 diversity across 11 models. σV/σK>1 indicates value-dominant geometry.
Model
Mechanism
αˉBOS
BOS key cos. sim.
Llama 3.1-8B
None
0.591
− 0.604
Llama 3.2-3B
None
0.612
− 0.586
Qwen2.5-3B
None
0.317
− 0.091
Qwen2.5-7B
None
0.279
− 0.337
Gemma2-2B
Softcapping
0.333
− 0.350
Gemma2-9B
Softcapping
0.396
− 0.424
Appendix
Table 5: BOS token Q-K alignment and key-space geometric distinctiveness across 11 models. High αˉBOS indicates strong persistent Q-K alignment (attentional anchor). Negative BOS key cosim indicates the BOS key is anti-aligned with the content key mean (geometric outlier).
Model
Mechanism
∥WV∥2/∥WK∥2
σ1(V)/σ1(K)
σV/σK
Llama 3.1-8B
None
0.192
0.057
0.33
Llama 3.2-3B
None
0.195
0.071
0.40
Qwen2.5-3B
None
0.668
0.216
1.19
Qwen2.5-7B
None
0.743
0.194
1.34
Gemma2-2B
Softcapping
0.550
0.257
0.88
Gemma2-9B
Softcapping
0.632
0.288
1.00
Appendix
Table 6: Three geometric metrics across architectures. The weight ratio is data-independent (projection matrix structure only); the other two require forward passes on calibration text. All three preserve the same architectural ordering.
Figure 4 : Per-layer sink rate ( τ=0.3 ) for all 11 models. Llama maintains high sink rates throughout; Qwen3.5 and GPT-OSS show near-zero; mixed-regime models show intermediate rates. Qwen2.5-3B uniquely exhibits sharp oscillations attributable to its 2 KV heads.
Figure 5 : Per-layer key and value L2 diversity ( σK(ℓ) , grey; σV(ℓ) , mechanism color) for all 11 models. Error bars show per-head standard deviation. Qwen3.5 shows only 8 full-attention layers.
Figure 6 : Per-layer BOS key cosine similarity for all 11 models. Values below zero indicate a geometrically distinct BOS key (attention sink); values above zero indicate sink suppression. The dashed line marks bos-cossim=0 .
Figure 7 : Per-layer σV/σK ratio for all 11 models. The dashed horizontal line marks σV=σK . All models except Qwen2.5 maintain a coherent regime across layers.
Model
Mechanism
Q heads
KV heads
GQA ratio
Attention scope
Llama 3.1-8B
None
32
8
4:1
Full
Llama 3.2-3B
None
24
8
3:1
Full
Qwen2.5-3B
None
16
2
8:1
Full
Qwen2.5-7B
None
28
4
7:1
Full
Gemma2-2B
Softcapping
8
4
2:1
Local 4096 (alt.)
Gemma2-9B
Softcapping
16
8
2:1
Local 4096 (alt.)
Appendix
Table 7: GQA compression ratios and effective attention scope across 11 models. Qwen2.5-3B and 7B have the highest GQA ratios among all-full-attention models; GPT-OSS shares an 8:1 ratio but limits the context burden via a 128-token sliding window in half its layers.
Figure 8 : Per-layer instability of sink rate (x-axis) vs. σV/σK (y-axis) across 11 models. Each point is one model. Qwen2.5 uniquely occupies the top-right quadrant where both signals are simultaneously unstable across layers.
Model
Mechanism
Mean Jaccard
Std (across layers)
Qwen2.5-7B
None
0.334
0.023
Qwen2.5-3B
None
0.338
0.030
Llama 3.2-3B
None
0.339
0.022
Llama 3.1-8B
None
0.347
0.031
Qwen3-4B
QK-norm
0.347
0.040
Gemma2-9B
Logit-cap
0.347
0.026
Appendix
Table 8: Mean top-256 attention Jaccard across query positions (model-level aggregate). Higher = more consistent rankings. Seven of nine models cluster in [0.33,0.40] ; Qwen3.5-4B falls slightly below (0.319) and GPT-OSS-20B is the sole high outlier (0.477).
Dense
Sink
TOVA
SnapKV
KeyDiff
ValueDiff
V-Dir
Qwen3.5-4B (gated attention), ctx = 128k
Peak memory (GB)
15.0
8.6
8.6
8.6
8.6
8.6
8.6
TTFT (s)
70.6
59.7
61.9
61.9
62.6
61.5
62.2
GPT-OSS-20B (learned sink), ctx = 65k
Peak memory (GB)
19.8
14.0
14.0
14.0
14.0
14.0
14.0
TTFT (s)
50.2
32.7
33.3
34.4
34.8
34.0
34.6
Appendix
Table 9: Long-context inference efficiency on H100 80 GB (budget = 2048). All eviction methods keep peak memory at O(budget) regardless of context length; ValueDiff variants match the cheapest eviction baseline (Sink) within 3–4% on TTFT.
Category
Model
Budget
flex@ k
exact@ k
score
Avg tok
Empty%
Trunc%
Logit softcapping
Gemma2-2B
16k
0.216
0.176
0.430
572
45.2
0.0
Gemma2-9B
16k
0.526
0.446
0.680
430
16.6
0.0
QK-normalization
Gemma3-4B
16k
0.742
0.682
0.820
932
11.2
0.0
Qwen3-4B
16k
0.928
0.792
0.960
1 947
1.4
0.0
Qwen3-4B-Thinking-2507
32k
0.888
0.704
0.940
6 974
6.6
0.0
Gated hybrid
Qwen3.5-4B ⋆
32k
0.890
0.780
0.890
1 124
0.0
0.0
Appendix
Table 10 : MATH-500 dense baselines ( n=100 , first 100 problems). For all models except Qwen3.5, k=5 generations per problem, flex@ k = mean grade_answer accuracy, and score = pass@ k ( ≥ 1 correct). For Qwen3.5 ⋆ , greedy k=1 with thinking disabled; flex@ k and score both equal pass@1. Tokens and empty-response rate are per-generation. Budget is max_new_tokens ; Trunc% is the fraction of generations that hit it. ‡ deepseek-r1-distill-llama-8b aggregated from dense_decoded/ after a BPE-marker post-processing fix. ⋆ Qwen3.5 uses enable_thinking=False , greedy ( do_sample=False ). Average token counts for Qwen3.5 set the reference lengths for eviction budgets b=256 and b=512 in Table 3 .
Category
b256
%
b512
%
Format (no \boxed{} )
59
60%
40
56%
Truncated mid-reasoning
37
37%
–
–
Context lost (input fidelity)
25
25%
< 5
–
Wrong answer ( \boxed{} present)
2
2%
1
1%
Appendix
Table 11: Failure categories for Sink eviction on Qwen3.5-4B, MATH-500 ( n=100 ). “Format (no box)” = correct reasoning, answer present but not in \boxed{} . “Truncated” = generation ends mid-reasoning without conclusion. “Context lost” = model states the problem was not provided (input fidelity failure). “Wrong answer” = \boxed{} present but mathematically incorrect.