Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents
Organizations: Capital One · University of Washington
Abstract
Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputation can be history-dependent: in our rolling-agent workload, an unchanged prompt can produce different answers depending on the requests processed before it. At a matched 5% recomputation budget, document-aligned recomputation reduces answer variation across request orders from 69.0% with CacheBlend's token top- policy to 26.1%. When each prompt is evaluated after a different sequence of preceding requests, document-aligned recomputation improves fidelity to full prefill by 34.5-52.5 percentage points over token top-, while both policies achieve approximately 5.7 median TTFT speedup. Our ablation study shows that, in our rolling-agent workload, contiguity is the main factor associated with robust selective recomputation.
Figures & tables
| Policy | Median TTFT | P95 TTFT | Median speedup vs. full prefill |
|---|---|---|---|
| Full prefill | 2,408 ms | 2,524 ms | 1.00 |
| Token top- , 5% | 426 ms | 459 ms | 5.68 |
| Document-aligned, 5% | 415 ms | 477 ms | 5.74 |
| Policy | Median segments | Median captured deviation |
|---|---|---|
| Token top- | 199 / 180.5 | 98.8% / 99.7% |
| Single-window | 1 / 1 | 7.2% / 6.1% |
| Fixed-chunk | 3 / 5 | 6.9% / 7.1% |
| Random-document | 6 / 6 | 4.8% / 4.8% |
| Document-aligned | 3 / 6 | 7.2% / 6.9% |