cs.AISep 27, 2026

RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

Authors: Ruoling Qi, Yirui Liu, Xuaner Wu, Yuxin Jin, Jian Chen, Jiayu Qin, Yin Chen, Jiawei Shao

Organizations: Shanghai Jiao Tong University · Institute of Artificial Intelligence, China Telecom (TeleAI) · State University of New York at Buffalo

Abstract

Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing methods selectively recompute token states to recover these missing interactions, but primarily allocate the recomputation budget to selecting which states to recompute, while fixing the recomputation context to the full causal prefix. We introduce RelaxKV, which formulates selective cache repair as a joint allocation problem over repair targets and recomputation context. Guided by the user query, RelaxKV identifies layer-specific repair targets and restricts their recomputation to a query-relevant context, reducing attention computation. Across four decoder models, RelaxKV at a 15% anchor ratio improves aggregate LongBench performance over ProphetKV on all models. On Qwen3-14B, RelaxKV provides a stronger quality-TTFT trade-off than ProphetKV across a 5%-30% anchor-ratio sweep, and achieves the best selective results on RULER-MV and LV-Eval at 16K and 32K context lengths. Controlled ablations further demonstrate the importance of recomputation context selection.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion

    Sep 28, 2026Genglin Wang, Wangsong Yin, Yeerzhati Abudunuer +3Key-Value CacheCross-Attention

  2. InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context

    Mar 5, 2026Xin Teng, Canyu Zhang, Shaoyi Zheng +3Long-Context ReasoningChunk

  3. RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

    Aug 2, 2026Changwoo Baek, Seungjun Shin, Kyeongbo KongKey-Value Cache CompressionKey-Value Cache Eviction