cs.LGSep 28, 2026

CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion

Authors: Genglin Wang, Wangsong Yin, Yeerzhati Abudunuer, Haoxuan Xu, Guoliang Xing, Zhenyu Yan

Organizations: The Chinese University of Hong Kong · Peking University · The Hong Kong University of Science and Technology

Abstract

Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns the difference between independently computed KV caches and those produced by processing the chunks together. The network combines compressed KV features with token embeddings and uses attention that is bidirectional within each chunk and flows from earlier to later chunks. Each repair block receives the compressed cache features, and the predicted residual is added to every document token's cache. Each repair network is trained for a specific frozen target LLM on a generic retrieval corpus and reused across downstream datasets. Our analysis shows that repair reduces KV errors both near chunk boundaries and throughout chunk interiors. Evaluation across three target LLMs and four downstream datasets places CacheRepair on the measured answer-quality-latency Pareto frontier in eleven of twelve model-dataset combinations. Reported time to first token (TTFT) includes online cache transfer and repair. Across all twelve combinations, the largest repairers achieve 1.69-4.61×\times speedups in median TTFT over full prefill and improve mean F1 by 2.1-26.1 percentage points over direct cache reuse.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

    Sep 27, 2026Ruoling Qi, Yirui Liu, Xuaner Wu +5CacheKey-Value Cache

  2. QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving

    Jun 4, 2026Jianxin Yan, Wangze Ni, Zhenxin Li +8Prefill

  3. CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

    Aug 7, 2026Gyuwan Kim, Cheoneum Park, Tao YangRetrieval-Augmented Generation PipelinesCache