Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific reconstruction. In this work, we propose ARC-KV, a novel reconstruction-based KV cache compaction method that follows this principle. To this end, we first train a value-aware indexer to select real-key anchors in a single scoring pass. ARC-KV then applies convex-hull-constrained key merging and fits an attention-mass bias and compact values against the full cache. At inference time, ARC-KV builds the compact cache once per context using the frozen indexer and reuses it for all subsequent queries. Extensive experiments demonstrate that ARC-KV outperforms reported compaction methods in most settings across QuALITY, RULER, and LongBench on Llama-3.1-8B-Instruct. In particular, at 10% KV retention on QuALITY, ARC-KV improves accuracy from 0.6409 to 0.6474 over Attention Matching while reducing compaction time by a factor of 25.73, from 959.8 s to 37.3 s.
Figures & tables
Construction
Ck
β
Cv
Rel. ℓ2↓
Cosine ↑
Hard subset
KS
0
VS
0.8778±0.2973
0.7543±0.1055
Mass calibration
KS
fitted
VS
0.8782±0.3189
0.7503±0.1100
Key/value merging
Ckmrg
fitted
Cvmrg
0.5109±0.1481
0.8558±0.0738
Value fitting
KS
fitted
CvLS
0.3546±0.1265
0.9135±0.0586
Key merging + value fitting
Ckmrg
fitted
CvLS
0.3200±0.1291
0.9250±0.0566
Table 1: Held-out-query reconstruction scores (mean ± SD over 2,560 article–layer–KV-head cells) on QuALITY with Llama-3.1-8B-Instruct at 5% retention. All constructions use the same fixed anchors. The key/value-merging row is diagnostic; ARC-KV merges only keys and fits Cv globally.
Figure 1: Overview of ARC-KV. Full-attention supervision trains a reusable indexer with hard top- t selection and soft surrogate gradients. At inference, selected anchors undergo context-specific key merging and fitting of β and Cv .
References
Compacted methods
Benchmark@Model
Full ctx.
No ctx.
AM
EA
SnapKV
KVzip
ARC-KV (Ours)
QuALITY@Qwen3-8B
0.5537
0.2383
0.4933
0.2830
0.2562
0.2125
0.5034
RULER@Qwen3-8B
0.9764
0.0000
0.2527
0.0109
0.0818
0.4040
0.3815
QuALITY@Gemma-3-12B-IT
0.7114
0.4944
0.6812
0.5962
0.5403
0.4832
0.6846
RULER@Gemma-3-12B-IT
0.9727
0.0000
0.1982
0.0509
0.1018
0.2273
0.3954
Table 2: Cross-model results at ρ=0.05 (higher is better). AM denotes OMP-based Attention Matching; bold marks the best compacted method in each row.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Pairs
Total indexer
Base model
Ratio
Llama-3.1-8B
256
155M
8.03B
1.93%
Qwen3-8B
288
175M
8.19B
2.13%
Gemma-3-12B-IT
64
67M
11.77B
0.57%
Appendix
Table 3: Total trainable parameter count of the value-aware indexers across all compacted layer–KV-head pairs, with HI=8 and dI=dA=128 per pair. The Pairs column reports the number of independently parameterized layer–KV-head indexers, and ratios are relative to each model’s text backbone. Gemma includes only the eight full-attention layers used for compaction.
Context
K/V per layer (GB) Full → compact
Latency per layer ( μ s) FA2 → ARC-KV
Speedup
32-layer attention (ms/token)
4K
1.07→0.21
174.3→54.6
3.19×
5.58→1.75
8K
2.15→0.43
341.4→89.4
3.82×
10.92→2.86
16K
4.29→0.86
667.9→166.3
4.02×
21.37→5.32
32K
8.59→1.72
1472.5→332.2
4.43×
47.12→10.63
Appendix
Table 4: Decode-time attention efficiency with a full cache and an ARC-KV cache at ρ=0.20 . K/V memory and latency are reported per attention layer for a batch of 64. The latency column compares full-cache FlashAttention-2 with compact attention including the fitted bias β . The final column sums the per-layer latency over all 32 attention layers.
λm
ρ=0.20
ρ=0.10
ρ=0.05
ρ=0.02
ρ=0.01
0.0
0.791
1.019
1.279
1.505
1.539
0.2
0.534
0.625
0.745
0.887
0.948
0.3
0.460
0.523
0.632
0.792
0.879
0.4
0.417
0.470
0.584
0.764
0.864
0.5
0.395
0.446
0.571
0.764
0.868
0.6
0.386
0.439
0.573
0.774
0.877
Appendix
Table 5: Relative attention-mass ℓ2 error for different key-merging coefficients λm on 50 QuALITY articles with Llama-3.1-8B-Instruct (lower is better). Bold indicates the minimum at each retention ratio at the reported precision.
Aggregation
0.20
0.10
0.05
0.02
0.01
Root-mean-square
0.00
0.00
0.00
0.00
0.00
Mean
−1.12
−3.69∗
−6.60∗
−5.37∗
−4.59∗
Max
−1.45
−2.24
−7.27∗
−8.95∗
−7.05∗
Appendix
Table 6: Aggregation ablation on QuALITY with Llama-3.1-8B-Instruct. Entries are accuracy differences in percentage points relative to root-mean-square aggregation across KV-cache retention ratios (higher is better). An asterisk marks a statistically significant difference from root-mean-square aggregation.
Allocation
0.20
0.10
0.05
0.02
0.01
Per-head (default)
0.00
0.00
0.00
0.00
0.00
Uniform
−1.68
−4.59∗
−10.29∗
−16.22∗
−10.96∗
Appendix
Table 7: Per-head budget-allocation ablation on QuALITY with Llama-3.1-8B-Instruct. Entries are accuracy differences in percentage points relative to the default model-specific per-head allocation across KV-cache retention ratios (higher is better). Both variants use the same total nominal cache budget. An asterisk marks a statistically significant difference from the default allocation.
Technical University of Darmstadt, Darmstadt, Germany · University of Notre Dame, Notre Dame, IN, USA · Technical University of Ilmenau, Ilmenau, Germany