Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computation grows quadratically. Existing approaches typically address one at the expense of irreversible token eviction, full-cache retention, or full-history reconstruction, limiting their effectiveness for multi-turn interaction and long-form reasoning. Motivated by two empirical properties, Long-Range Inter-Token Similarity and Smooth Residual Distribution, we propose ResidualKV, which factorizes the KV cache into a sparse set of globally retrieved references and compact, quantized residual codes for the remaining tokens. This representation preserves token-specific information without permanent eviction and, when combined with sparse attention, reconstructs only the selected states on demand. Dynamic-stride scheduling further reduces reference growth from linear to approximately logarithmic at ultra-long contexts. Across Llama, Qwen, LLaVA-OV, and Qwen3-VL backbones, ResidualKV maintains near-full-cache performance using only 13%-16% KV storage and 30% attention computation on LongBench, and 8%-10% storage and 10% computation in matched-budget multimodal evaluation. It also accelerates decoding by up to 1.5× with KV-cache quantization and 3.4× without it. These results show that global cross-token redundancy supports accurate, memory-efficient, and computation-efficient long-context inference. The source code is available at https://github.com/CURRENTF/ResidualKV.
Figures & tables
Fig. 1: Three-way trade-off in KV cache optimization. Existing method families typically balance only two of information fidelity, memory reduction, and computational efficiency, whereas ResidualKV jointly achieves all three.
Fig. 2: Two empirical properties underlying ResidualKV. KV states exhibit long-range inter-token similarity , with highly similar references often located beyond local neighborhoods. Subtracting these references removes shared components and yields a smooth residual distribution concentrated near zero, facilitating compression and low-bit quantization.
Fig. 3: Empirical analysis of cross-token KV redundancy and residualization. (a) The histogram of maximum historical cosine similarity for each token. (b) The histogram of log distance to the most similar historical token for each token. (c) Cumulative variance of retrieved-reference residuals, original KV states, and random residuals. (d) The histogram of L2 norm distributions of original and residual KV states. (e) and (f) Value distributions before and after reference subtraction, respectively; Residualization removes dominant shared components and produces low-magnitude, near-zero token-specific residuals.
Fig. 4: Overview of ResidualKV. This framework encompasses both inference and training workflows across four key components: (A) Residual Compression (§ IV-A ), which retrieves reference KV cache via flexible scheduling to generate highly quantizable compressed residuals; (B) Residual Reconstruction (§ IV-B ), which employs a lightweight linear decompressor during the memory-bound decoding phase to reconstruct original KV cache; (C) Joint Sparse Inference (§ IV-C ), which seamlessly integrates with sparse attention methods to perform selective KV reconstruction; and (D) Hybrid Objective Training (§ IV-D ), a low-cost training scheme (costing only $8) using a joint objective.
Fig. 5: Reference count and RULER-Variable Tracking (VT) performance under fixed- and dynamic-stride scheduling. The left and right axes report the number of reference tokens and variable-tracking scores, respectively, under the same compressor and sparse-attention configuration.
Fig. 6: Latency benefit of the heavy-compressor, lightweight-decompressor (heavy-light) design. With comparable parameter counts (7.37M for the symmetric baseline and 7.34M for ours), the heavier one-time compressor increases prefill cost but reduces per-step reconstruction latency from 0.93 to 0.47 ms and total latency from 480 to 243 ms over a 4,096-token prefill and 512 decode steps.
Fig. 7: Joint sparse inference with OmniKV and ResidualKV. An OmniKV filter layer selects historical token indices using full attention. ResidualKV then fetches and reconstructs only the corresponding residual codes through lightweight linear decompression and reference addition.
Model
2dk
dc
dff
Lfull
LLaVA-OV-7B
1,024
384
3,072
0,1,4,7,14
Qwen3-VL-8B-Ins
2,048
384
3,072
0,1,2,3,10,21,29
Llama-3.1-8B-Ins
2,048
512
3,072
0,1,2,8,18
Qwen2.5-7B-Inst
1,024
384
3,072
0,1,2,4,7,14
Qwen2.5-32B-Ins
2,048
512
4,096
0,1,2,3,4,5,17,29,40
Qwen3-4B-Ins
2,048
512
3,072
0,3,9,13,16,21,28
TABLE I: ResidualKV configuration by model. 2dk is the original per-token KV dimension, dc is the compressed residual dimension, dff is the SwiGLU intermediate width, and Lfull lists the full-attention layers used.
Model
Method
( KR , CR ) ↓
MMB ↑
MME ↑
MMMU ↑
POPE ↑
SQA ↑
V-MME ↑
MVB ↑
Average ↑
LLaVA-OV -7B
FastV [ 47 ]
( 10 , 10 )
80.3
81.7/1816.0
44.9
76.8
83.3
53.6
51.0
67.4
PACT [ 48 ]
( 10 , 10 )
84.9
84.7/2013.9
47.8
88.7
92.3
58.7
53.1
72.9
DivPrune [ 49 ]
( 10 , 10 )
80.4
79.6/1745.2
46.4
85.4
83.5
54.7
52.1
68.9
FastVID [ 50 ]
( 10 , 10 )
–
–
–
–
–
56.3
53.5
–
ResidualKV (Ours)
( 8 , 10 )
85.1
84.9/2032.3
48.6
89.2
95.2
58.6
53.7
73.6
Qwen3-VL -8B-Ins
FastV [ 47 ]
( 10 , 10 )
72.4
80.0/1641.6
47.7
74.2
78.7
47.4
48.4
64.1
TABLE II: Main results on multimodal benchmarks. We evaluate ResidualKV on LLaVA-OV-7B and Qwen3-VL-8B-Ins backbones across seven representative multimodal benchmarks.
Model
Method
( KR , CR ) ↓
Single-Doc ↑
Multi-Doc ↑
Summ. ↑
Few-Shot ↑
Synthetic ↑
Code ↑
Average ↑
Llama3.1 -8B-Ins
SnapKV [ 12 ]
( 30 , 30 )
43.6
46.3
26.4
68.4
55.4
60.2
50.1
PyramidKV [ 66 ]
( 30 , 30 )
43.7
46.6
26.2
68.8
55.6
59.7
50.1
Quest [ 15 ]
( 100 , 30 )
42.6
46.0
29.1
69.1
54.8
57.8
49.9
KIVI-2bit [ 18 ]
( 19 , 100 )
43.2
46.3
28.9
69.1
54.8
59.9
50.3
AdaKV [ 67 ]
( 30 , 30 )
43.1
46.0
26.4
69.5
55.0
58.2
49.7
OmniKV [ 16 ]
( 100 , 30 )
42.9
46.0
28.5
68.7
54.9
60.0
50.2
TABLE III: Main results on LongBench. We compare ResidualKV against state-of-the-art baselines across varying model scales. Variants marked † omit the neural compressor; all methods are matched by target ratio r and averaged over 16 datasets.
Model
Method
( KR , CR ) ↓
R.KV ↑
R.PS ↑
En.QA ↑
S+N ↑
MS ↑
Avg. ↑
Qwen2.5 7B-Ins
SnapKV
( 20 , 20 )
1.4
1.6
18.0
28.6
59.3
22.5
KIVI-2bit
( 19 , 100 )
0.0
0.0
25.6
58.9
59.3
24.4
KVzip
( 20 , 20 )
36.2
3.0
44.0
66.5
33.7
36.7
ResidualKV
( 19 , 20 )
52.6
23.4
29.9
47.6
61.5
40.7
Qwen3 4B-Ins
SnapKV
( 20 , 20 )
2.6
0.0
16.8
50.1
61.1
22.8
KIVI-2bit
( 19 , 100 )
0.2
0.2
16.5
55.4
53.0
20.7
TABLE IV: Main results on SCBench. R.KV, R.PS, En.QA, S+N, and MS denote key-value retrieval, prefix-suffix retrieval, English question answering, summarization plus needle-in-a-haystack retrieval, and many-shot in-context learning.
Model
Full
SnapKV
OmniKV
ResidualKV
Qwen3-4B-Think
80.0
60.0
76.7
80.0
TABLE V: Main results on the AIME reasoning benchmark.
Fig. 8: Batch-size-1 decoding throughput across context lengths. Comparison of full attention and ResidualKV with and without KV-cache quantization.
α
Nref @ 262K ↓
Ref. Fraction ↓
LongBench ↑
0.000
26,189
9.99%
47.4
0.001
3,354
1.28%
47.4
0.020
319
0.12%
47.3
0.050
149
0.06%
47.0
0.100
83
0.03%
47.1
TABLE VI: Dynamic stride trade-off on Qwen2.5-7B-Ins. Increasing α sharply reduces the reference set at 262K context with a modest change in the LongBench average score.
Fig. 9: Distribution of compressed KV representations. (a) Magnitudes remain smooth across tokens and channels without localized spikes. (b) Values concentrate near zero, facilitating low-bit quantization.
Method
( KR , CR ) ↓
Average ↑
ResidualKV Standalone
( 16 , 100 )
50.5
ResidualKV + OmniKV
( 13 , 30 )
50.3
ResidualKV + SnapKV
( 17.5 , 40 )
50.1
TABLE VII: Standalone and compositional use of ResidualKV on LongBench with Llama-3.1-8B-Ins. Standalone ResidualKV reconstructs all compressed KV pairs without an external sparsity mask, while ResidualKV+OmniKV and ResidualKV+SnapKV compose residual compression with dynamic and static token selection, respectively.
Fig. 10: Joint vs. separate K/V compression. Step-aligned NTP loss under matched trainable parameters, compression ratios, and training steps, smoothed over 200 steps.
Method
( KR , CR ) ↓
SD ↑
MD ↑
Sum ↑
FS ↑
Syn ↑
Code ↑
Avg. ↑
ResidualKV
( 45 , 30 )
43.2
47.1
27.9
69.2
54.4
59.8
50.3
w/o fc and fd
( 45 , 30 )
34.4
43.2
24.5
65.9
55.1
57.2
46.7
w/o Ref. Tokens T
( 47 , 30 )
39.3
45.3
22.1
62.3
50.8
55.4
45.9
TABLE VIII: Component-wise ablation study on Llama-3.1-8B-Ins. SD, MD, Sum, FS, Syn, and Code denote single-document QA, multi-document QA, summarization, few-shot learning, synthetic tasks, and code generation; KV quantization is disabled and dc=768 so that all KRs are comparable.
Fig. 11: Hyperparameter sensitivity. We vary the compressed dimension, number of references, compressor intermediate width, and reference stride while holding other settings fixed. Blue/red curves show NTP loss/relative decoding throughput, red circles denote the defaults, and red crosses indicate OOM.
Fig. 12: Ablation of the hybrid training objective. Comparison of NTP-only, MSE-only, and joint training using total, NTP, and MSE losses. Joint training preserves prediction quality while minimizing reconstruction error.
Fig. 13: Multimodal reference retrieval and residual distribution. Left and middle: target tokens T on object boundaries retrieve references from semantically and structurally similar boundaries elsewhere in the image. Right: reference subtraction transforms the long-tailed raw KV distribution into a compact, near-zero residual distribution.
Technical University of Darmstadt, Darmstadt, Germany · University of Notre Dame, Notre Dame, IN, USA · Technical University of Ilmenau, Ilmenau, Germany