Cross-context KV-cache reuse predicts a shared segment's keys and values under a new prefix instead of recomputing them, and has been reported to do so without quality loss. We find otherwise, and identify two problems. (1) A hidden cost: on MMLU and GSM8K, reuse costs substantial accuracy. (2) A decision at the wrong unit: no rule for deciding whether to reuse a cache removes that cost. What does help is choosing which parts of the cache to recompute, and the value of choosing well falls as the unit of choice grows: informed selection removes 49.5% of the cache error beyond chance at single rows (one token's keys and values), 10.6% at 64-token chunks, and nothing at the level of whole calls. Budgeted Cache Repair (BCR) acts at the unit where selection still pays. It drafts two tokens from the assembled cache, ranks cache rows by the attention those tokens pay them, and recomputes a fixed budget of rows exactly, in one of three layouts. The cost is paid rather than predicted away, and the draft that fails as a gate succeeds as a selector. BCR restores GSM8K to dense-prefill accuracy while still serving most calls from cache, and its best layout outperforms every reuse baseline's mean in the reference grid. The draft also beats a coin-flip selector at the same budget - a control prior evaluations lack.
Figures & tables
Figure 1: Three ways to serve a shared segment. Rows show three agents sharing one segment in execution order: A1 is the first caller and computes it exactly; A2 and A3 are later callers and can reuse it from cache. The horizontal axis denotes token position.
ρ
mean
spread
0.05
76.5
3.3
0.20
77.5
11.3
0.40
78.8
2.7
KVComm
77.3
–
Dense
82.7
–
Table 1: The budget: how large, in what layout, at what price. (a) two agents, 150 matched GSM8K problems, 64 -token chunks drawn by coin at four seeds so ρ is the only variable; mean and spread are the accuracy over those seeds and its range, max−min . KVComm is ρ=0 , Dense ρ=1 . (b) rows recomputed per agent call at four agents. (c) time to the first token on a 4096 -symbol synthetic segment at ρ=0.20 , × the speedup against Dense . Full tables in Tables 11 , 15 and 10 .
Accuracy vs. Dense
Reuse
Method
reuse when
acc. (%)
95% CI
cost
b/c
disc.
p
rate
Dense
never
85.7
[81.2, 89.2]
–
–
–
–
0.000
KVComm
H(s)≤γ=0.3log2n
81.3
[76.5, 85.3]
−4.3
23/10
33
0.035 ∗
0.638
OpenGate
conf. ≥θ=0.5
79.0
[74.0, 83.2]
−6.7
37/17
54
0.009 ∗
0.902
KVComm
H(s)≤γ=1.0log2n
78.0
[73.0, 82.3]
−7.7
35/12
47
0.001 ∗
0.959
Table 2: What the gate can and cannot do. The first 300 GSM8K problems in the reference configuration. b counts problems Dense solves and the row misses, c the reverse, disc. their sum; ∗ marks p<0.05 . cost is points against Dense , taken from the discordant counts rather than the rounded accuracies. rate is the fraction of the 1200 agent calls served without a dense prefill.
Figure 2: What reuse costs, and what the gate can do about it. GSM8K, n=300 ( Tables 2 and 6 ).
MMLU
GSM8K
HumanEval
# Agents
2
3
4
5
avg
2
3
4
5
avg
2
3
4
5
avg
Accuracy (%), HumanEval pass@1
Dense
74.5
75.2
75.2
75.2
75.0
82.7
86.0
84.7
84.7
84.5
87.5
86.2
85.6
85.0
86.1
KVComm
71.9
73.2
72.5
72.5
72.5
77.3
82.0
80.0
84.0
80.8
86.2
85.6
85.6
85.6
85.8
CacheBlend, reuse-matched
69.3
69.3
69.3
68.6
69.1
82.0
86.7
84.7
81.3
83.7
88.8
85.6
85.0
86.2
86.4
CacheBlend, cost-matched
72.5
71.2
69.3
73.9
71.7
85.3
83.3
83.3
82.0
83.5
85.0
84.4
85.6
85.6
85.2
Table 3: What reuse costs across agent counts. Reference configuration, γ=0.3 ( 1.0 on HumanEval and on GSM8K’s CacheBlend rows); 153 , 150 and 160 items, matched per problem. Bold is the best score in a column excluding Dense , shading the best BCR mean; avg carries no test. Dense reuses nothing and the cost-matched pairing matches rows, so neither has a reuse row. Splits in Tables 13 and 14 .
Policy
reuse
resid.
Protect-the-Root
0.456
+0.9
Random
0.454
+1.6
Protect-the-Leaf
0.504
+0.4
Dense
0.000
+0.0
KVComm
0.600
+0.0
Table 4: The evidence behind the cost, the two nulls and the layout ordering. (a) five agents, 150 matched GSM8K problems, roles pinned; resid. is the points earned beyond Equation 3 . (b) six prompts; gap is the cache error informed selection removes beyond chance, as a share of the unrepaired 0.1054 , won the prompts where it beat the coin. (c) the three layouts of § 3 , b/c splits pooled over four agent counts. (d) is Table 13 ’s cost rows, p in parentheses. Throughout b counts problems KVComm misses and the other solves. Full tables in Tables 7 , 8 and 13 .
GSM8K
MMLU
Method
acc.
b/c ( p )
acc.
b/c ( p )
KVComm
71.0
–
51.7
–
BCR -chunk , coin
78.9 †
–
50.0 †
–
BCR -chunk , draft
84.3
9/49
( 0.000∗ )
54.7
12/21
( 0.163 )
token-matched Dense
–
–
59.3
25/48
( 0.010∗ )
Dense
84.7
10/51
( 0.000∗ )
66.0
21/64
( 0.000∗ )
Table 5: Where repair works, and what it is worth against a published selector. Five agents pinned to one role, 300 matched problems. In (a) b counts problems KVComm solves at its shipped γ=0.3 and the row misses, c the reverse, p in parentheses, ∗ marking p<0.05 ; − marks a row with no split of its own— KVComm itself, the † five-seed coin mean, and the token-matched control, which only MMLU needs. Shading marks the repair row. (b) pools the reference grid over four agent counts, descriptively ( Table 14 ).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Problems
reuse rate
Dense
KVComm γ=0.3
Δ
1–60
0.537
80.0
81.7
+1.7
61–120
0.625
88.3
80.0
−8.3
121–180
0.679
90.0
85.0
−5.0
181–240
0.675
83.3
80.0
−3.3
241–300
0.675
86.7
80.0
−6.7
1–300
0.638
85.7
81.3
−4.3
Appendix
Table 6: Accuracy and reuse rate by block. Dense against the shipped default ( γ=0.3 ) in blocks of 60 GSM8K problems, with that block’s reuse rate. The shaded first block—what a 60 -problem evaluation would see—is the only one where reuse does not lose, and the one with the lowest reuse rate.
Policy
decides from
reuse
acc. (%)
95% CI
p vs Dense
pred.
resid.
Dense
no reuse
0.000
85.3
[79, 90]
–
85.3
+0.0
Protect-the-Root
structural
0.456
78.7
[71, 84]
0.013 ∗
77.7
+0.9
Random
information-free
0.454
79.3
[72, 85]
0.064
77.8
+1.6
Protect-the-Leaf
structural
0.504
77.3
[70, 83]
0.023 ∗
76.9
+0.4
KVComm γ=0.3
entropy gate
0.600
75.3
[68, 82]
0.006 ∗
75.3
+0.0
Appendix
Table 7: Allocation policies on one line. Five agents, 150 matched GSM8K problems, roles pinned; rows sorted by reuse rate. pred. is what Equation 3 predicts from the reuse rate alone, resid. the points earned beyond it. p is an exact McNemar test against Dense ; ∗ marks p<0.05 . Both columns use the unrounded scores, so redoing them from the printed ones can shift the last digit.
unit of choice
budget
attention
random
gap
gap / no-repair
prompts ( p )
1 (single row)
20.0%
0.0375
0.0896
+0.0521
49.5%
6/6
( 0.01 )
16
20.3%
0.0650
0.0851
+0.0201
19.0%
6/6
( 0.02 )
64
18.8%
0.0763
0.0875
+0.0112
10.6%
5/6
( 0.08 )
256
25.0%
0.0780
0.0880
+0.0100
9.5%
4/6
( 0.28 )
1024
50.0%
0.0691
0.0631
−0.0060
−5.7%
1/6
( 0.69 )
2048 (whole segment)
100.0%
both restore everything
0
0%
–
Appendix
Table 8: Selection value by unit of choice. One 2048 -token segment, six prompts, varying only how rows are grouped and picked. Error is the relative L2 distance of the attention output from a dense cache ( 0.1054 unrestored); gap is what informed selection removes beyond chance; the last column gives prompts where informed won and a paired t over the six. The budget keeps max(1,round(0.2nchunks)) chunks, so the last three rows are not budget-matched.
Accuracy
Reuse
Latency
Method
γ
acc. (%)
95% CI
p
all
late
calls
TTFT (ms)
E2E (s)
Dense
–
80.0
[68, 88]
–
0.000
0.000
0/240
101.5
28.33
KVComm
0.3
81.7
[70, 89]
–
0.537
0.662
129/111
79.2
32.00
KVComm
0.5
81.7
[70, 89]
1.00
0.671
0.850
161/79
82.7
36.78
KVComm
0.7
78.3
[66, 87]
0.73
0.787
0.875
189/51
79.6
36.70
KVComm
1.0
78.3
[66, 87]
0.62
0.838
0.950
201/39
78.8
36.33
Appendix
Table 9: The gate sweep at n=60 . The threshold γ over its meaningful range on 60 matched GSM8K problems; γ=1.0 disables the entropy test, so the last row is unconditional reuse. all / late are the fraction of calls served without a dense prefill, over all 60 problems and the last 20 ; calls counts reused versus dense calls out of 240 . p is an exact McNemar against the γ=0.3 default; brackets are Wilson 95% intervals.
Shared segment (symbols)
Metric
Method
ρ
64
512
2048
4096
TTFT (ms)
Dense
1.00
159.2
198.1
341.2
542.7
KVComm
0.00
60.3
56.2
64.7
72.9
BCR -chunk (SDPA)
0.20
102.4
112.1
155.2
229.8
BCR -chunk (FA2)
0.20
101.2
109.8
140.5
190.1
BCR -chunk (FA2)
0.10
100.8
99.0
111.1
137.0
Appendix
Table 10: Latency against segment length. Five CopyMachine agents on synthetic prompts. The harness runs each call twice in one process—assembled cache and full prefill—so every ratio is paired, and the speedup rows are means of those per-call ratios rather than ratios of the means above them; the two differ, 7.63× against 542.7/72.9 . ρ is the repaired fraction, 0 plain reuse and 1 a full prefill; the parenthetical names the repair kernel: PyTorch’s scaled dot-product attention, or FlashAttention-2. Caveat: 12 samples per length, so the anchor pool never saturates—the best case for reuse, not its steady state. Shading marks the ρ=0.20 rows, the setting used elsewhere.
Accuracy (%), by coin seed
ρ
0
1
2
3
mean
95% CI
0.05
78.0
74.7
78.0
75.3
76.5
± 2.8
0.10
76.7
76.0
82.0
75.3
77.5
± 4.9
0.20
70.0
78.0
80.7
81.3
77.5
± 8.3
0.30
81.3
78.0
75.3
82.7
79.3
± 5.3
0.40
78.7
77.3
79.3
80.0
78.8
± 1.8
Appendix
Table 11: The repair budget swept on GSM8K, at four coin seeds. Two agents, 150 matched problems, 64 -token chunks drawn by coin so that ρ is the only variable; KVComm scores 77.3% , Dense 82.7% . Each budget is run four times : the coin’s stream advances per repaired call, so one run is one realisation of the policy, not the policy.
Accuracy (%)
Rows/call
b/c vs. cap 64
Span cap
Mean len.
N=2
N=4
N=2
N=4
N=2
N=4
1
1.00
84.0
86.0
185
216
11/6
( 0.33 )
11/4
( 0.12 )
2
1.78
–
83.3
–
223
–
9/6
( 0.61 )
4
2.87
–
84.7
–
215
–
7/2
( 0.18 )
8
3.68
80.0
–
186
–
3/4
( 1.00 )
–
16
3.82
80.7
–
186
–
0/0
( 1.00 )
–
Appendix
Table 12: The repair’s layout, swept with its budget held fixed. GSM8K, n=150 , ρ=0.20 , draft-scored, two and four agents; the foot rows reproduce Table 3 ’s GSM8K columns. Only the cap on one contiguous span moves. Cells not run are dashed; b/c is against the 64 -token cap, whose answers are identical to the 16 -token one on all 150 problems.
# Agents
Dataset
Split
2
3
4
5
MMLU
Dense vs. KVComm
16/12
(0.572)
15/12
(0.701)
18/14
(0.597)
18/14
(0.597)
BCR -chunk vs. KVComm
4/0
(0.125)
4/3
(1.000)
5/5
(1.000)
8/4
(0.388)
BCR -grow vs. KVComm
17/9
(0.169)
8/9
(1.000)
15/10
(0.424)
14/9
(0.405)
BCR -single vs. KVComm
11/12
(1.000)
11/13
(0.839)
13/11
(0.839)
15/14
(1.000)
GSM8K
Dense vs. KVComm
17/9
(0.169)
11/5
(0.210)
15/8
(0.210)
9/8
(1.000)
Appendix
Table 13: Paired statistics for Table 3 . Same runs; the discordant split and exact p at each cell. Each BCR row is split against KVComm , the Dense row against KVComm . In every row b counts problems KVComm misses and the other condition solves, c the reverse—the opposite orientation to Table 5 . Condition against condition is in Table 14 . Shading marks the BCR rows.
# Agents
Dataset
Condition
2
3
4
5
MMLU
CacheBlend, reuse-matched
8/16
( 0.152 )
7/14
( 0.189 )
13/18
( 0.473 )
11/21
( 0.110 )
CacheBlend, cost-matched
12/15
( 0.701 )
10/14
( 0.541 )
12/17
( 0.458 )
13/15
( 0.851 )
BCR -grow
14/10
( 0.541 )
8/10
( 0.815 )
17/12
( 0.458 )
12/11
( 1.000 )
BCR -single
9/14
( 0.405 )
10/13
( 0.678 )
13/11
( 0.839 )
12/15
( 0.701 )
GSM8K
CacheBlend, reuse-matched
10/10
( 1.000 )
5/4
( 1.000 )
6/7
( 1.000 )
4/10
( 0.180 )
Appendix
Table 14: Each condition against BCR -chunk. Discordant split of an exact McNemar test between the named condition and BCR -chunk ( 64 -token chunks), per agent count; b counts problems the named condition gets right and BCR -chunk does not. Table 13 pairs against KVComm instead. MMLU’s CacheBlend rows were re-run at the matched gate offline; GSM8K’s still carry the mismatch of App. D .
# Agents
Dataset
Method
2
3
4
5
MMLU
CacheBlend, reuse-matched
249
342
417
521
CacheBlend, cost-matched
90
124
158
193
BCR -chunk
96
123
159
191
BCR -grow
91
136
174
209
BCR -single
90
132
167
202
Appendix
Table 15: Rows recomputed per agent call. What each repair actually recomputes, which Table 3 ’s reuse column cannot show, since a repaired call still counts as reused. The three layouts spend the same nominal ρ=0.20 , the fixed-chunk condition quantising to whole 64 -row chunks. CacheBlend’s cost-matched pairing uses the same ρ ; its reuse-matched pairing spends whatever that cell’s reuse rate requires. MMLU is measured offline. Shading marks the BCR rows.