Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputation can be history-dependent: in our rolling-agent workload, an unchanged prompt can produce different answers depending on the requests processed before it. At a matched 5% recomputation budget, document-aligned recomputation reduces answer variation across request orders from 69.0% with CacheBlend's token top-k policy to 26.1%. When each prompt is evaluated after a different sequence of preceding requests, document-aligned recomputation improves fidelity to full prefill by 34.5-52.5 percentage points over token top-k, while both policies achieve approximately 5.7× median TTFT speedup. Our ablation study shows that, in our rolling-agent workload, contiguity is the main factor associated with robust selective recomputation.
Figures & tables
Figure 1: Overview of the rolling workload, persistent cache history, and equal-budget recomputation policies. Different request orders expose the same prompt to different cached chunk KV states; every policy recomputes 5% of reusable tokens.
Figure 2: (A) Answer variation, measured as mean pairwise exact-answer disagreement across the five request orders; error bars show 95% confidence intervals. (B) Fidelity to full prefill in each request order. Both recomputation policies use a 5% recomputation ratio.
Policy
Median TTFT
P95 TTFT
Median speedup vs. full prefill
Full prefill
2,408 ms
2,524 ms
1.00 ×
Token top- k , 5%
426 ms
459 ms
5.68 ×
Document-aligned, 5%
415 ms
477 ms
5.74 ×
Table 1: Request-level time to first token.
Figure 3: Fidelity of five equal-budget selection policies on the same prompts under chronological and one fixed-random order. Every policy uses a 5% recomputation ratio and matches the realized token count within each order. Error bars show 95% Wilson score intervals.
Policy
Median segments
Median captured deviation
Token top- k
199 / 180.5
98.8% / 99.7%
Single-window
1 / 1
7.2% / 6.1%
Fixed-chunk
3 / 5
6.9% / 7.1%
Random-document
6 / 6
4.8% / 4.8%
Document-aligned
3 / 6
7.2% / 6.9%
Table 2: Selection fragmentation and captured deviation. Each cell reports chronological / fixed-random results.