Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computation grows quadratically. Existing approaches typically address one at the expense of irreversible token eviction, full-cache retention, or full-history reconstruction, limiting their effectiveness for multi-turn interaction and long-form reasoning. Motivated by two empirical properties, Long-Range Inter-Token Similarity and Smooth Residual Distribution, we propose ResidualKV, which factorizes the KV cache into a sparse set of globally retrieved references and compact, quantized residual codes for the remaining tokens. This representation preserves token-specific information without permanent eviction and, when combined with sparse attention, reconstructs only the selected states on demand. Dynamic-stride scheduling further reduces reference growth from linear to approximately logarithmic at ultra-long contexts. Across Llama, Qwen, LLaVA-OV, and Qwen3-VL backbones, ResidualKV maintains near-full-cache performance using only 13%-16% KV storage and 30% attention computation on LongBench, and 8%-10% storage and 10% computation in matched-budget multimodal evaluation. It also accelerates decoding by up to 1.5× with KV-cache quantization and 3.4× without it. These results show that global cross-token redundancy supports accurate, memory-efficient, and computation-efficient long-context inference. The source code is available at https://github.com/CURRENTF/ResidualKV.
Figures & tables
Fig. 1: Three-way trade-off in KV cache optimization. Existing method families typically balance only two of information fidelity, memory reduction, and computational efficiency, whereas ResidualKV jointly achieves all three.
Fig. 2: Two empirical properties underlying ResidualKV. KV states exhibit long-range inter-token similarity , with highly similar references often located beyond local neighborhoods. Subtracting these references removes shared components and yields a smooth residual distribution concentrated near zero, facilitating compression and low-bit quantization.
Fig. 3: Empirical analysis of cross-token KV redundancy and residualization. (a) The histogram of maximum historical cosine similarity for each token. (b) The histogram of log distance to the most similar historical token for each token. (c) Cumulative variance of retrieved-reference residuals, original KV states, and random residuals. (d) The histogram of L2 norm distributions of original and residual KV states. (e) and (f) Value distributions before and after reference subtraction, respectively; Residualization removes dominant shared components and produces low-magnitude, near-zero token-specific residuals.
Fig. 4: Overview of ResidualKV. This framework encompasses both inference and training workflows across four key components: (A) Residual Compression (§ IV-A ), which retrieves reference KV cache via flexible scheduling to generate highly quantizable compressed residuals; (B) Residual Reconstruction (§ IV-B ), which employs a lightweight linear decompressor during the memory-bound decoding phase to reconstruct original KV cache; (C) Joint Sparse Inference (§ IV-C ), which seamlessly integrates with sparse attention methods to perform selective KV reconstruction; and (D) Hybrid Objective Training (§ IV-D ), a low-cost training scheme (costing only $8) using a joint objective.
Fig. 5: Reference count and RULER-Variable Tracking (VT) performance under fixed- and dynamic-stride scheduling. The left and right axes report the number of reference tokens and variable-tracking scores, respectively, under the same compressor and sparse-attention configuration.
Fig. 6: Latency benefit of the heavy-compressor, lightweight-decompressor (heavy-light) design. With comparable parameter counts (7.37M for the symmetric baseline and 7.34M for ours), the heavier one-time compressor increases prefill cost but reduces per-step reconstruction latency from 0.93 to 0.47 ms and total latency from 480 to 243 ms over a 4,096-token prefill and 512 decode steps.
Fig. 7: Joint sparse inference with OmniKV and ResidualKV. An OmniKV filter layer selects historical token indices using full attention. ResidualKV then fetches and reconstructs only the corresponding residual codes through lightweight linear decompression and reference addition.
Model
2dk
dc
dff
Lfull
LLaVA-OV-7B
1,024
384
3,072
0,1,4,7,14
Qwen3-VL-8B-Ins
2,048
384
3,072
0,1,2,3,10,21,29
Llama-3.1-8B-Ins
2,048
512
3,072
0,1,2,8,18
Qwen2.5-7B-Inst
1,024
384
3,072
0,1,2,4,7,14
Qwen2.5-32B-Ins
2,048
512
4,096
0,1,2,3,4,5,17,29,40
Qwen3-4B-Ins
2,048
512
3,072
0,3,9,13,16,21,28
TABLE I: ResidualKV configuration by model. 2dk is the original per-token KV dimension, dc is the compressed residual dimension, dff is the SwiGLU intermediate width, and Lfull lists the full-attention layers used.
Model
Method
( KR , CR ) ↓
MMB ↑
MME ↑
MMMU ↑
POPE ↑
SQA ↑
V-MME ↑
MVB ↑
Average ↑
LLaVA-OV -7B
FastV [ 47 ]
( 10 , 10 )
80.3
81.7/1816.0
44.9
76.8
83.3
53.6
51.0
67.4
PACT [ 48 ]
( 10 , 10 )
84.9
84.7/2013.9
47.8
88.7
92.3
58.7
53.1
72.9
DivPrune [ 49 ]
( 10 , 10 )
80.4
79.6/1745.2
46.4
85.4
83.5
54.7
52.1
68.9
FastVID [ 50 ]
( 10 , 10 )
–
–
–
–
–
56.3
53.5
–
ResidualKV (Ours)
( 8 , 10 )
85.1
84.9/2032.3
48.6
89.2
95.2
58.6
53.7
73.6
Qwen3-VL -8B-Ins
FastV [ 47 ]
( 10 , 10 )
72.4
80.0/1641.6
47.7
74.2
78.7
47.4
48.4
64.1
TABLE II: Main results on multimodal benchmarks. We evaluate ResidualKV on LLaVA-OV-7B and Qwen3-VL-8B-Ins backbones across seven representative multimodal benchmarks.
Model
Method
( KR , CR ) ↓
Single-Doc ↑
Multi-Doc ↑
Summ. ↑
Few-Shot ↑
Synthetic ↑
Code ↑
Average ↑
Llama3.1 -8B-Ins
SnapKV [ 12 ]
( 30 , 30 )
43.6
46.3
26.4
68.4
55.4
60.2
50.1
PyramidKV [ 66 ]
( 30 , 30 )
43.7
46.6
26.2
68.8
55.6
59.7
50.1
Quest [ 15 ]
( 100 , 30 )
42.6
46.0
29.1
69.1
54.8
57.8
49.9
KIVI-2bit [ 18 ]
( 19 , 100 )
43.2
46.3
28.9
69.1
54.8
59.9
50.3
AdaKV [ 67 ]
( 30 , 30 )
43.1
46.0
26.4
69.5
55.0
58.2
49.7
OmniKV [ 16 ]
( 100 , 30 )
42.9
46.0
28.5
68.7
54.9
60.0
50.2
TABLE III: Main results on LongBench. We compare ResidualKV against state-of-the-art baselines across varying model scales. Variants marked † omit the neural compressor; all methods are matched by target ratio r and averaged over 16 datasets.
Model
Method
( KR , CR ) ↓
R.KV ↑
R.PS ↑
En.QA ↑
S+N ↑
MS ↑
Avg. ↑
Qwen2.5 7B-Ins
SnapKV
( 20 , 20 )
1.4
1.6
18.0
28.6
59.3
22.5
KIVI-2bit
( 19 , 100 )
0.0
0.0
25.6
58.9
59.3
24.4
KVzip
( 20 , 20 )
36.2
3.0
44.0
66.5
33.7
36.7
ResidualKV
( 19 , 20 )
52.6
23.4
29.9
47.6
61.5
40.7
Qwen3 4B-Ins
SnapKV
( 20 , 20 )
2.6
0.0
16.8
50.1
61.1
22.8
KIVI-2bit
( 19 , 100 )
0.2
0.2
16.5
55.4
53.0
20.7
TABLE IV: Main results on SCBench. R.KV, R.PS, En.QA, S+N, and MS denote key-value retrieval, prefix-suffix retrieval, English question answering, summarization plus needle-in-a-haystack retrieval, and many-shot in-context learning.
Model
Full
SnapKV
OmniKV
ResidualKV
Qwen3-4B-Think
80.0
60.0
76.7
80.0
TABLE V: Main results on the AIME reasoning benchmark.
Fig. 8: Batch-size-1 decoding throughput across context lengths. Comparison of full attention and ResidualKV with and without KV-cache quantization.
α
Nref @ 262K ↓
Ref. Fraction ↓
LongBench ↑
0.000
26,189
9.99%
47.4
0.001
3,354
1.28%
47.4
0.020
319
0.12%
47.3
0.050
149
0.06%
47.0
0.100
83
0.03%
47.1
TABLE VI: Dynamic stride trade-off on Qwen2.5-7B-Ins. Increasing α sharply reduces the reference set at 262K context with a modest change in the LongBench average score.
Fig. 9: Distribution of compressed KV representations. (a) Magnitudes remain smooth across tokens and channels without localized spikes. (b) Values concentrate near zero, facilitating low-bit quantization.
Method
( KR , CR ) ↓
Average ↑
ResidualKV Standalone
( 16 , 100 )
50.5
ResidualKV + OmniKV
( 13 , 30 )
50.3
ResidualKV + SnapKV
( 17.5 , 40 )
50.1
TABLE VII: Standalone and compositional use of ResidualKV on LongBench with Llama-3.1-8B-Ins. Standalone ResidualKV reconstructs all compressed KV pairs without an external sparsity mask, while ResidualKV+OmniKV and ResidualKV+SnapKV compose residual compression with dynamic and static token selection, respectively.
Fig. 10: Joint vs. separate K/V compression. Step-aligned NTP loss under matched trainable parameters, compression ratios, and training steps, smoothed over 200 steps.
Method
( KR , CR ) ↓
SD ↑
MD ↑
Sum ↑
FS ↑
Syn ↑
Code ↑
Avg. ↑
ResidualKV
( 45 , 30 )
43.2
47.1
27.9
69.2
54.4
59.8
50.3
w/o fc and fd
( 45 , 30 )
34.4
43.2
24.5
65.9
55.1
57.2
46.7
w/o Ref. Tokens T
( 47 , 30 )
39.3
45.3
22.1
62.3
50.8
55.4
45.9
TABLE VIII: Component-wise ablation study on Llama-3.1-8B-Ins. SD, MD, Sum, FS, Syn, and Code denote single-document QA, multi-document QA, summarization, few-shot learning, synthetic tasks, and code generation; KV quantization is disabled and dc=768 so that all KRs are comparable.
Fig. 11: Hyperparameter sensitivity. We vary the compressed dimension, number of references, compressor intermediate width, and reference stride while holding other settings fixed. Blue/red curves show NTP loss/relative decoding throughput, red circles denote the defaults, and red crosses indicate OOM.
Fig. 12: Ablation of the hybrid training objective. Comparison of NTP-only, MSE-only, and joint training using total, NTP, and MSE losses. Joint training preserves prediction quality while minimizing reconstruction error.
Fig. 13: Multimodal reference retrieval and residual distribution. Left and middle: target tokens T on object boundaries retrieve references from semantically and structurally similar boundaries elsewhere in the image. Right: reference subtraction transforms the long-tailed raw KV distribution into a compact, near-zero residual distribution.
Large language models increasingly operate over long contexts, where the KV cache becomes a dominant memory bottleneck: its size grows linearly with sequence length and must be retained throughout decoding, making full GPU caching prohibitively expensive without compression. Existing KV cache compression methods struggle to balance efficiency with faithful context preservation. Token eviction discards information, while semantic grouping fixes compression decisions at prefill time; neither can recover token-level detail from a compressed span once it becomes relevant during generation. As a solution, we propose SeKV, a resolution-adaptive semantic KV cache that organizes context into entropy-guided semantic spans and stores them across a GPU-CPU memory hierarchy without discarding information. Each span keeps a lightweight summary vector on GPU for coarse routing and a low-rank SVD basis on CPU for on-demand token-level reconstruction. A trained zoom-in mechanism selectively expands query-relevant spans during decoding, enabling precise retrieval without materializing the full KV cache on GPU. SeKV enables adaptive token-level reconstruction while keeping the base LLM fully frozen and adding fewer than 0.05% trainable parameters. Across four benchmarks, SeKV improves over the strongest semantic compression baseline by 5.9% on average while reducing GPU memory by 53.3% versus full KV caching at 128K context. Code is available on https://github.com/AmirAbaskohi/SeKV.
Amirhossein Abaskohi, Giuseppe Carenini, Peter West +1
University of British Columbia · Microsoft Research
Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that grows linearly with sequence length, quickly becoming the bottleneck for long context inference. A practical remedy is to evict less important KV entries; however, existing eviction policies are largely heuristic and struggle to capture the rich, input-dependent distribution of token importance. In this work, we introduce a learnable indexer that predicts KV importance, enabling more accurate retention of critical tokens. Meanwhile, naively evicting tokens permanently discards their information, leading to irreversible forgetting and degraded retrieval over long ranges. To address this, we propose a lightweight latent memory module that compresses evicted tokens into a compact, online-updated state and provides residual readouts to compensate for the attention contributions lost through KV eviction. Collectively, our method enables accurate long-context inference under a bounded KV budget, delivering consistent improvements on RULER (4K/16K) across Qwen, Mistral, and Llama models (up to 25 points under aggressive eviction), markedly more stable Needle-in-a-Haystack retrieval, and superior LongBench scores and compression curves compared to existing eviction policies.
Xintong Yang, Hao Gu, Binxing Xu +6
The Hong Kong University of Science and Technology · Zhejiang University
Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware. Existing KV cache eviction methods typically apply heuristic token scoring over all heads in GQA-based LLMs. These methods ignore the different functionalities of attention heads, leading to the eviction of critical tokens and thus degrading the performance of LLMs. To address this issue, we propose CompressKV, a resource-efficient KV-cache compression framework for GQA-based LLMs. Instead of aggregating attention scores from all heads, CompressKV identifies Semantic Retrieval Heads (SRHs) that capture both the initial and final tokens of a prompt and semantically important mid-context evidence, and uses them to select tokens whose KV pairs should be retained. Furthermore, CompressKV allocates cache budgets across layers according to offline estimates of layer-wise eviction error. Experiments on LongBench and Needle-in-a-Haystack show that CompressKV consistently outperforms existing KV-cache eviction methods across memory budgets. Notably, it preserves over 97% of full-cache performance using only 3% of the KV cache on LongBench question-answering tasks and achieves 90% accuracy with just 0.7% KV storage on Needle-in-a-Haystack. These results demonstrate an improved resource--performance trade-off for long-context LLM inference. Our code is publicly available at: https://github.com/TUDa-HWAI/CompressKV
Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3
Technical University of Darmstadt, Darmstadt, Germany · University of Notre Dame, Notre Dame, IN, USA · Technical University of Ilmenau, Ilmenau, Germany