Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing methods selectively recompute token states to recover these missing interactions, but primarily allocate the recomputation budget to selecting which states to recompute, while fixing the recomputation context to the full causal prefix. We introduce RelaxKV, which formulates selective cache repair as a joint allocation problem over repair targets and recomputation context. Guided by the user query, RelaxKV identifies layer-specific repair targets and restricts their recomputation to a query-relevant context, reducing attention computation. Across four decoder models, RelaxKV at a 15% anchor ratio improves aggregate LongBench performance over ProphetKV on all models. On Qwen3-14B, RelaxKV provides a stronger quality-TTFT trade-off than ProphetKV across a 5%-30% anchor-ratio sweep, and achieves the best selective results on RULER-MV and LV-Eval at 16K and 32K context lengths. Controlled ablations further demonstrate the importance of recomputation context selection.
Figures & tables
Figure 1: Layer-wise query-attention scores over prefix tokens. High-attention score tokens positions shift across layers.
Qwen3-14B
Method
WQA
TQA
HQA
NQA
MQue
PR-en
PR-zh
LB Avg
Full Recompute
8.50
88.41
54.95
24.15
25.57
68.00
100.00
52.80
Full Reuse
7.22
66.60
28.01
12.45
5.47
24.00
72.14
30.84
CacheBlend
8.46
85.11
42.57
21.57
20.57
66.00
77.00
45.90
EPIC
8.52
87.23
47.63
23.31
21.11
53.00
88.00
46.97
KVShare
7.46
86.47
33.46
13.64
7.03
36.50
69.04
36.23
Table 1: LongBench results. Full Recompute serves as the dense reference. Directly comparable selective baselines and RelaxKV G use a nominal 20% repair ratio, while RelaxKV uses a fixed 15% anchor ratio. Bold and underline denote the best and second-best selective results, respectively.
Method
RULER-MV ↑
LV-Eval ↑
4K
8K
16K
32K
16K
32K
Full Recompute
100.00
100.00
96.67
95.83
32.37
23.27
Full Reuse
52.00
41.17
38.67
13.83
0.64
0.29
CacheBlend
69.33
55.50
41.50
20.33
1.66
0.74
EPIC
57.17
49.00
52.17
71.33
14.37
10.11
KVShare
53.67
47.17
44.33
24.00
0.98
0.38
Table 2: Long-context results on Qwen3-14B. RULER-MV reports accuracy and LV-Eval reports token F1.
Variant Context Source Scope Set size Global Sparse Across-Layer Mean Shared ∣U∣ Layerwise Sparse Per-Layer Score At Each Layer ∣U∣ RelaxKV .15 Anchor union U Shared ∣U∣
Table 3: Effect of recomputation context construction on Qwen3-14B at a 15% anchor ratio. All variants share the same repair anchors, dependency-closed active repair targets, and pre-truncation context-set size ∣U∣ ; only context construction varies. Panel (a) defines the variants, and panel (b) reports task quality / mean recomputation latency (s).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Active Repair Targets
Recomputation Context
Role
RelaxKV
Rl
Cl(p)
Main method
RelaxKV G
Global P
Cl(p)
Matched-target diagnostic control
Full-Prefix Control
Rl
Full causal prefix
Full-prefix quality–cost control
Global Sparse
Rl
Global Top- ∣U∣
Global context construction
Layerwise Sparse
Rl
Layerwise Top- ∣U∣
Layerwise context construction
Global Repair Target Sweep
Global Prg
Cl(p)
Global target capacity
Appendix
Table 4: Controlled configurations. Every active repair target updates its corresponding KV entry; configurations vary in active repair targets or recomputation context columns. Global repair targets P are shared by all layers.
MuSiQue
RULER-MV 32K
LV-Eval 16K
LV-Eval 32K
Variant
Recomputation Context
F1 ↑
Recomp. (s) ↓
Score ↑
Recomp. (s) ↓
F1 ↑
Recomp. (s) ↓
F1 ↑
Recomp. (s) ↓
Full-Prefix Control
Full causal prefix
39.63
10.11
96.00
55.05
29.32
38.78
23.32
115.97
RelaxKV .20
Cl(p)
39.01
8.51
95.00
42.30
28.87
32.88
23.86
104.45
Appendix
Table 5: Full-prefix quality–cost control on Qwen3-14B with a 20% anchor ratio. Both configurations share repair anchors, dependency-closed active targets, global position IDs, and cache updates. Recomp. is mean recomputation latency per example.
Chunk size
Recompute
Reuse
CacheBlend
EPIC
KVShare
ProphetKV
RelaxKV G .20
RelaxKV .15
64
41.89
4.45
5.54
22.54
6.40
26.77
24.94
32.66
128
41.89
5.27
7.33
25.49
7.81
31.75
30.08
34.54
256
41.89
7.81
9.64
30.49
10.46
33.97
31.62
36.39
512
41.89
12.62
13.64
30.19
13.41
35.57
36.28
38.08
1024
41.89
18.46
22.47
34.05
19.00
39.27
37.62
38.93
Appendix
Table 6: Sensitivity to reusable chunk size on Qwen3-14B MuSiQue. Selective baselines and RelaxKV G use a nominal 20% repair target ratio; RelaxKV uses a 15% anchor ratio. Full Recompute is independent of reusable chunk size and is repeated as a common dense reference. Bold and underline mark the best and second-best selective results.
Repair Target Ratio
RULER ↓
MuSiQue ↓
LB Avg ↑
0.2
361.58
267.86
57.84
0.3
486.23
332.68
57.24
0.4
609.55
396.31
57.57
0.5
753.35
458.32
57.37
0.6
871.97
539.33
58.06
0.7
941.26
608.62
58.10
Appendix
Table 7: Global repair target capacity sweep on Llama-3.1-8B-Instruct. The shared context is fixed as the union of layerwise Top-20% anchors. TTFT is in milliseconds.
Repair Target
RULER Ratio
MuSiQue Ratio
WQA
TQA
HQA
NQA
MQue
PR-en
PR-zh
Avg
0.2
.200/.651
.200/.704
45.86
90.80
48.62
25.30
25.40
74.50
94.43
57.84
0.3
.300/.651
.300/.704
44.96
90.30
47.61
24.90
25.02
75.50
92.38
57.24
0.4
.400/.651
.400/.704
45.14
90.30
48.57
25.37
25.71
76.00
91.88
57.57
0.5
.500/.651
.500/.704
45.56
91.13
48.48
24.51
25.48
76.00
90.46
57.37
0.6
.600/.651
.600/.704
46.40
91.46
49.29
25.86
25.99
75.50
91.90
58.06
0.7
.651/.651
.694/.704
45.70
91.46
48.07
26.19
28.13
76.00
91.13
58.10
Appendix
Table 8: Effective repair-target/context-set ratios and complete LongBench scores for the global target sweep. Requested target ratios are capped per example at ∣U∣ .