cs.ARSep 27, 2026

Budgeted Cache Repair for Cross-Context KV-Cache Reuse

Authors: Haeyong Kang, Chang D. Yoo

Organizations: Duksung Women’s University · KAIST

Abstract

Cross-context KV-cache reuse predicts a shared segment's keys and values under a new prefix instead of recomputing them, and has been reported to do so without quality loss. We find otherwise, and identify two problems. (1) A hidden cost: on MMLU and GSM8K, reuse costs substantial accuracy. (2) A decision at the wrong unit: no rule for deciding whether to reuse a cache removes that cost. What does help is choosing which parts of the cache to recompute, and the value of choosing well falls as the unit of choice grows: informed selection removes 49.5% of the cache error beyond chance at single rows (one token's keys and values), 10.6% at 64-token chunks, and nothing at the level of whole calls. Budgeted Cache Repair (BCR) acts at the unit where selection still pays. It drafts two tokens from the assembled cache, ranks cache rows by the attention those tokens pay them, and recomputes a fixed budget of rows exactly, in one of three layouts. The cost is paid rather than predicted away, and the draft that fails as a gate succeeds as a selector. BCR restores GSM8K to dense-prefill accuracy while still serving most calls from cache, and its best layout outperforms every reuse baseline's mean in the reference grid. The draft also beats a coin-flip selector at the same budget - a control prior evaluations lack.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

    Sep 27, 2026Ruoling Qi, Yirui Liu, Xuaner Wu +5Retrieval-Augmented GenerationKV-Cache Management

  2. KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

    Sep 9, 2026Xi Shi, Mengxin Zheng, Qian LouKV CachingLLM Inference Acceleration