Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure both on one three-tier agent architecture. Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines. The persistent tier does not: across eight controlled dataset pairs at n=100 per arm it costs +0.368 MiB [+0.167, +0.590] of peak cache and produces no detectable accuracy change (+0.015, 95% CI [-0.011, +0.046]). We argue the null is structural: single-question benchmarks supply each item with its own evidence and score it independently, and correctness requires resetting stored traces between conditions, so recall has nothing informative to retrieve. Reaching it took four measurement corrections -- three inflating the apparent benefit, the fourth making an effect that size look resolvable -- none visible in the results table. We give the conditions an agent-memory ablation must satisfy and detection procedures that need no knowledge of the specific defect.
Figures & tables
Figure 1: The measured system and the single axis the ablation varies; Appendix A.1 gives the block-level call sequence.
Flat ‡
RAG ( Lewis et al., 2020 )
CoA-PKV (ours)
vs Flat
vs RAG
peak KV per query (MiB)
35.5 [34.0, 37.0]
35.3 [30.9, 39.8]
14.3 [13.0, 15.4]
−59.9%
−59.7%
latency per query (s) †
67
52
29
−56.7%
−44.2%
Table 1: Peak KV and latency: Flat , RAG , and CoA-PKV (ours). Brackets: 95% CI.
effect
95% CI
peak KV cost of enabling recall
+0.368 MiB
[+0.167,+0.590]
accuracy change from enabling recall
+0.015
[−0.011,+0.046]
Table 2: Cost and accuracy effect of enabling persistent recall.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
cell
n
Δ peak KV (MiB)
95% CI
n
Δ accuracy
95% CI
HotpotQA
96
+0.321
[+0.145,+0.504]
100
+0.002
[−0.026,+0.029]
GPQA
89
+0.640
[+0.083,+1.217]
100
+0.090
[+0.010,+0.170]
MMLU-Pro
98
+0.748
[+0.314,+1.189]
100
+0.010
[−0.060,+0.080]
LB-HotpotQA
95
+0.028
[−0.161,+0.226]
100
+0.029
[−0.015,+0.077]
LB-2WikiMultihopQA
98
+0.221
[−0.018,+0.492]
100
+0.007
[−0.031,+0.047]
LB-MuSiQue
99
+0.283
[−0.049,+0.635]
100
+0.004
[−0.037,+0.045]
Appendix
Table 3: Per-cell ablation results.
correction
effect
output schema + aligned prompt
+0.145
scoped retrieval
+0.082/+0.032/+0.045
trace-reset protocol
−0.067
disabling reasoner CoT (reverted)
−0.209
lexical reranker (reverted)
−4 of 40 gold-in-top-5
Appendix
Table 4: What the corrections cost.
cell
as-sent
content-only
Δ reachable
Δ unreachable
HotpotQA
0%
0%
—
−0.006
GPQA
27%
24%
+0.087
+0.000
MMLU-Pro
86%
86%
−0.012
+0.000
LB-HotpotQA
2%
2%
+0.000
+0.017
LB-2WikiMultihopQA
10%
1%
−0.045
+0.019
LB-MuSiQue
16%
7%
−0.088
+0.022
Appendix
Table 5: Recall reachability by cell.
dataset
Flat
RAG
CoA-PKV
vs Flat
vs RAG
est/peak ‡
peak KV (MiB)
(CoA-PKV)
HotpotQA
35.00
30.18
13.59
−61.2%
−55.0%
12.2×
GPQA
42.13
n/a ∗
19.89
−52.8%
—
8.9×
MMLU-Pro
36.76
n/a ∗
21.92
−40.4%
—
6.3×
LB-HotpotQA
36.09
40.06
15.49
−57.1%
−61.3%
11.1×
LB-2WikiMultihopQA
n/a †
38.84
14.34
—
−63.1%
9.1×
Appendix
Table 6: Per-dataset KV measurement, apples to apples.