cs.AIOct 5, 2026

Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents

Authors: Tiffany Gu, Annie Guan, Manshu Huang, Nitin Rao, Siddhant Shah, Margaret Capetz

Organizations: Capital One · University of Washington

Abstract

Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputation can be history-dependent: in our rolling-agent workload, an unchanged prompt can produce different answers depending on the requests processed before it. At a matched 5% recomputation budget, document-aligned recomputation reduces answer variation across request orders from 69.0% with CacheBlend's token top-kk policy to 26.1%. When each prompt is evaluated after a different sequence of preceding requests, document-aligned recomputation improves fidelity to full prefill by 34.5-52.5 percentage points over token top-kk, while both policies achieve approximately 5.7×\times median TTFT speedup. Our ablation study shows that, in our rolling-agent workload, contiguity is the main factor associated with robust selective recomputation.

Figures & tables

Explore similar work

CardsList
  1. RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

    Sep 27, 2026Ruoling Qi, Yirui Liu, Xuaner Wu +5CacheKey-Value Cache

  2. Leyline: KV Cache Directives for Agentic Inference

    May 31, 2026Bole Ma, Jan Eitzinger, Harald KoestlerCacheKey-Value Cache