cs.AISep 29, 2026

ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction

Authors: Zheyu Shen, Guanhua Wang, Dezhan Tu, Mengchi Zhang, Yanjia Li, Adnan Aziz, Chunqiang Tang, Ang Li

Organizations: University of Maryland, College Park · Meta

Abstract

Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific reconstruction. In this work, we propose ARC-KV, a novel reconstruction-based KV cache compaction method that follows this principle. To this end, we first train a value-aware indexer to select real-key anchors in a single scoring pass. ARC-KV then applies convex-hull-constrained key merging and fits an attention-mass bias and compact values against the full cache. At inference time, ARC-KV builds the compact cache once per context using the frozen indexer and reuses it for all subsequent queries. Extensive experiments demonstrate that ARC-KV outperforms reported compaction methods in most settings across QuALITY, RULER, and LongBench on Llama-3.1-8B-Instruct. In particular, at 10% KV retention on QuALITY, ARC-KV improves accuracy from 0.6409 to 0.6474 over Attention Matching while reducing compaction time by a factor of 25.73, from 959.8 s to 37.3 s.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AnchorKV: Anchor-Residual KV Cache Compression

    Aug 3, 2026Malik Khalaf, Yara Shamshoum, Nitzan Hodos +2Key-Value Cache CompressionKey-Value Cache

  2. CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

    Jun 23, 2026Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3Key-Value Cache CompressionLLM Inference Optimization

  3. PatchKV: Weight-Space Compensation of KV Cache

    Sep 30, 2026Chanryeol Lee, Chanhyuk Lee, Yeonwoo Choi +2Key-Value Cache CompressionKey-Value Cache