Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific reconstruction. In this work, we propose ARC-KV, a novel reconstruction-based KV cache compaction method that follows this principle. To this end, we first train a value-aware indexer to select real-key anchors in a single scoring pass. ARC-KV then applies convex-hull-constrained key merging and fits an attention-mass bias and compact values against the full cache. At inference time, ARC-KV builds the compact cache once per context using the frozen indexer and reuses it for all subsequent queries. Extensive experiments demonstrate that ARC-KV outperforms reported compaction methods in most settings across QuALITY, RULER, and LongBench on Llama-3.1-8B-Instruct. In particular, at 10% KV retention on QuALITY, ARC-KV improves accuracy from 0.6409 to 0.6474 over Attention Matching while reducing compaction time by a factor of 25.73, from 959.8 s to 37.3 s.
Figures & tables
Construction
Ck
β
Cv
Rel. ℓ2↓
Cosine ↑
Hard subset
KS
0
VS
0.8778±0.2973
0.7543±0.1055
Mass calibration
KS
fitted
VS
0.8782±0.3189
0.7503±0.1100
Key/value merging
Ckmrg
fitted
Cvmrg
0.5109±0.1481
0.8558±0.0738
Value fitting
KS
fitted
CvLS
0.3546±0.1265
0.9135±0.0586
Key merging + value fitting
Ckmrg
fitted
CvLS
0.3200±0.1291
0.9250±0.0566
Table 1: Held-out-query reconstruction scores (mean ± SD over 2,560 article–layer–KV-head cells) on QuALITY with Llama-3.1-8B-Instruct at 5% retention. All constructions use the same fixed anchors. The key/value-merging row is diagnostic; ARC-KV merges only keys and fits Cv globally.
Figure 1: Overview of ARC-KV. Full-attention supervision trains a reusable indexer with hard top- t selection and soft surrogate gradients. At inference, selected anchors undergo context-specific key merging and fitting of β and Cv .
References
Compacted methods
Benchmark@Model
Full ctx.
No ctx.
AM
EA
SnapKV
KVzip
ARC-KV (Ours)
QuALITY@Qwen3-8B
0.5537
0.2383
0.4933
0.2830
0.2562
0.2125
0.5034
RULER@Qwen3-8B
0.9764
0.0000
0.2527
0.0109
0.0818
0.4040
0.3815
QuALITY@Gemma-3-12B-IT
0.7114
0.4944
0.6812
0.5962
0.5403
0.4832
0.6846
RULER@Gemma-3-12B-IT
0.9727
0.0000
0.1982
0.0509
0.1018
0.2273
0.3954
Table 2: Cross-model results at ρ=0.05 (higher is better). AM denotes OMP-based Attention Matching; bold marks the best compacted method in each row.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Pairs
Total indexer
Base model
Ratio
Llama-3.1-8B
256
155M
8.03B
1.93%
Qwen3-8B
288
175M
8.19B
2.13%
Gemma-3-12B-IT
64
67M
11.77B
0.57%
Appendix
Table 3: Total trainable parameter count of the value-aware indexers across all compacted layer–KV-head pairs, with HI=8 and dI=dA=128 per pair. The Pairs column reports the number of independently parameterized layer–KV-head indexers, and ratios are relative to each model’s text backbone. Gemma includes only the eight full-attention layers used for compaction.
Context
K/V per layer (GB) Full → compact
Latency per layer ( μ s) FA2 → ARC-KV
Speedup
32-layer attention (ms/token)
4K
1.07→0.21
174.3→54.6
3.19×
5.58→1.75
8K
2.15→0.43
341.4→89.4
3.82×
10.92→2.86
16K
4.29→0.86
667.9→166.3
4.02×
21.37→5.32
32K
8.59→1.72
1472.5→332.2
4.43×
47.12→10.63
Appendix
Table 4: Decode-time attention efficiency with a full cache and an ARC-KV cache at ρ=0.20 . K/V memory and latency are reported per attention layer for a batch of 64. The latency column compares full-cache FlashAttention-2 with compact attention including the fitted bias β . The final column sums the per-layer latency over all 32 attention layers.
λm
ρ=0.20
ρ=0.10
ρ=0.05
ρ=0.02
ρ=0.01
0.0
0.791
1.019
1.279
1.505
1.539
0.2
0.534
0.625
0.745
0.887
0.948
0.3
0.460
0.523
0.632
0.792
0.879
0.4
0.417
0.470
0.584
0.764
0.864
0.5
0.395
0.446
0.571
0.764
0.868
0.6
0.386
0.439
0.573
0.774
0.877
Appendix
Table 5: Relative attention-mass ℓ2 error for different key-merging coefficients λm on 50 QuALITY articles with Llama-3.1-8B-Instruct (lower is better). Bold indicates the minimum at each retention ratio at the reported precision.
Aggregation
0.20
0.10
0.05
0.02
0.01
Root-mean-square
0.00
0.00
0.00
0.00
0.00
Mean
−1.12
−3.69∗
−6.60∗
−5.37∗
−4.59∗
Max
−1.45
−2.24
−7.27∗
−8.95∗
−7.05∗
Appendix
Table 6: Aggregation ablation on QuALITY with Llama-3.1-8B-Instruct. Entries are accuracy differences in percentage points relative to root-mean-square aggregation across KV-cache retention ratios (higher is better). An asterisk marks a statistically significant difference from root-mean-square aggregation.
Allocation
0.20
0.10
0.05
0.02
0.01
Per-head (default)
0.00
0.00
0.00
0.00
0.00
Uniform
−1.68
−4.59∗
−10.29∗
−16.22∗
−10.96∗
Appendix
Table 7: Per-head budget-allocation ablation on QuALITY with Llama-3.1-8B-Instruct. Entries are accuracy differences in percentage points relative to the default model-specific per-head allocation across KV-cache retention ratios (higher is better). Both variants use the same total nominal cache budget. An asterisk marks a statistically significant difference from the default allocation.
The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction methods permanently discard tokens, degrading performance whenever a discarded token later proves essential, while quantization methods retain all tokens at low precision but offer limited compression. We propose AnchorKV, a compression scheme that shrinks the cache by 20× without discarding a single token. AnchorKV represents the cache using a small set of anchors stored exactly, expresses every other token through its most similar anchor, and refines only those whose approximation most affects the model's output. AnchorKV consistently preserves accuracy across models and datasets, retaining 99% of the full-cache score at the 70B scale, while keeping the entire context at a fraction of its cost.
Malik Khalaf, Yara Shamshoum, Nitzan Hodos +2
Department of Computer Science, Technion - Israel Institute of Technology
Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware. Existing KV cache eviction methods typically apply heuristic token scoring over all heads in GQA-based LLMs. These methods ignore the different functionalities of attention heads, leading to the eviction of critical tokens and thus degrading the performance of LLMs. To address this issue, we propose CompressKV, a resource-efficient KV-cache compression framework for GQA-based LLMs. Instead of aggregating attention scores from all heads, CompressKV identifies Semantic Retrieval Heads (SRHs) that capture both the initial and final tokens of a prompt and semantically important mid-context evidence, and uses them to select tokens whose KV pairs should be retained. Furthermore, CompressKV allocates cache budgets across layers according to offline estimates of layer-wise eviction error. Experiments on LongBench and Needle-in-a-Haystack show that CompressKV consistently outperforms existing KV-cache eviction methods across memory budgets. Notably, it preserves over 97% of full-cache performance using only 3% of the KV cache on LongBench question-answering tasks and achieves 90% accuracy with just 0.7% KV storage on Needle-in-a-Haystack. These results demonstrate an improved resource--performance trade-off for long-context LLM inference. Our code is publicly available at: https://github.com/TUDa-HWAI/CompressKV
Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3
Technical University of Darmstadt, Darmstadt, Germany · University of Notre Dame, Notre Dame, IN, USA · Technical University of Ilmenau, Ilmenau, Germany
Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework that compensates KV cache compression methods by carrying part of the context in the model's weights. PatchKV pairs an off-the-shelf compressed KV cache with a context-specific weight patch, which is computed once at context-loading time and served for downstream queries for the context. The weight patch is derived in closed form via ridge regression, by aligning the block-wise activations of context-derived reference query tokens under the full cache and the compressed cache. Once merged into the model, the patch leaves the forward graph and per-query inference cost unchanged in the single-context, multi-query setting. Across long-context QA (SCBench with up to 170K tokens, SQuAD, NIAH) and math (GSM8K) benchmarks on three model architectures, PatchKV consistently improves cache compression methods, suggesting an alternative direction to compensate them at aggressive budgets.