Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating linear attention (\eg, Qwen3.5), prior methods designed for softmax attention struggle to generalize. Our analysis reveals that both attention- and similarity-based approaches suffer notable performance degradation, underscoring the urgent need for compression methods tailored to this regime. To this end, we propose \textbf{V-CoLA}, an efficient training-free token compression framework specifically designed for linear attention. V-CoLA introduces a novel \textit{uniqueness-aware importance criterion} for identifying critical vision tokens, coupled with an \textit{adaptive token merging strategy} that performs compression. All components are optimized at the implementation level to remain compatible with the chunk-wise parallelism of linear attention, ensuring strong practical value. Extensive experiments across multiple benchmarks demonstrate the superiority of V-CoLA: it achieves 99.5% of the original performance with only 50.0% of vision tokens, and over 88.0% with as few as 12.5%, while delivering a 1.86× to 6.15× prefill speedup.
Figures & tables
Figure 1: Comparison of token selection. FastV suffers from attention bias, missing the digits at the top of the image. DART selects pivot tokens from only sparse semantic regions, leading to insufficient awareness of some primary objects. Our method successfully identifies all main regions and provides an accurate response.
Figure 2: FastV (attention-based) vs. DART (similarity-based) vs. random baseline on Qwen3.5-9B.
Figure 3: Average proportion of attention scores allocated to vision tokens across different layers in Qwen3-VL-8B and Qwen3.5-9B.
Figure 4: Comparison of similarity-based token selection between Qwen3-VL-8B and Qwen3.5-9B.
Figure 5: Overview of our method. (left) Original hybrid architectures. (right) Forwarding with V-CoLA, equipped with token-wise importance criterion and adaptive chunk-wise token merging.
Figure 6: Layer-wise retention rate (blue) and importance correlation (red) on Qwen3.5-9B.
Method
MME
MMB
GQA
SQA
T-VQA
POPE
VizWiz
MMStar
Avg(%)
Upper Bound, 880 Tokens (100%)
Qwen3.5-9B
2398.2
85.6
61.1
92.2
83.2
89.9
69.2
49.3
100.0
Remain 440 Tokens in Average ( ↓ 50.0%)
FastV Chen et al. (2024a)
2351.6
84.8
60.4
92.1
79.5
88.8
67.6
45.3
97.5
SparseVLM Zhang et al. (2024b)
2333.1
84.7
60.1
92.2
80.1
88.7
67.5
45.6
97.5
DART Wen et al. (2025)
2220.4
83.3
59.6
91.7
69.3
88.9
67.0
48.3
95.5
Table 1: We compare our method with state-of-the-art approaches using Qwen3.5-9B. The best results are bold.
# V-Tokens
100%
50.0%
25.0%
12.5%
2K
269
145 (1.86 × )
88 (3.06 × )
61 (4.41 × )
4K
511
273 (1.87 × )
153 (3.34 × )
94 (5.44 × )
8K
1034
542 (1.91 × )
289 (3.58 × )
168 (6.15 × )
Table 2: Comparison of wall-clock runtime (ms) of the language model prefilling across various vision token numbers and different remain rates.
Component
# Tokens
Runtime (ms)
Softmax Attention
8K × 1
68.5
Gated DeltaRule
8K × 1
34.1
8K × 4
135.1
Extended Gated DeltaRule
8K × 1
34.2
8K × 4
35.1 (3.85 × )
Adaptive Token Merging
8K × 1
0.86
Table 3: Comparison of wall-clock runtime across different components.
λ
MME
SQA
POPE
MMStar
0.00 (only Rtimp )
2360.5
92.3
89.5
48.9
0.05
2388.0
92.5
89.6
49.6
0.10
2398.5
92.7
89.8
49.8
0.90
2333.1
91.7
87.3
48.0
Table 4: Ablation of parameter λ in Rt calculation.
(w1,w2)
MME
SQA
POPE
MMStar
(0,1)
2387.6
92.5
89.3
49.1
(0,4)
2398.5
92.7
89.8
49.8
(−4,0)
2376.4
92.4
87.9
49.3
(−4,4)
2395.2
92.6
89.9
49.5
Table 5: Ablation of window parameters (w1,w2) .
Method
MME
SQA
POPE
MMStar
↓ 75.0% Vision Tokens
V-CoLA
2268.7
92.4
88.9
47.2
- w/o Ada. Merging
2197.4
91.5
88.4
44.3
- w/o Early Exit
2235.1
91.9
88.2
45.6
↓ 87.5% Vision Tokens
V-CoLA
2012.0
90.9
88.3
41.4
Table 6: Ablation of main components.
Method
MME
SQA
POPE
MMStar
Qwen3.5-27B
2517.4
97.0
90.5
58.1
- DART
2173.0
91.9
82.4
44.5
- VisionZip
2405.2
96.0
88.4
53.3
- DTP
1798.2
89.2
70.2
43.1
- V-CoLA (Ours)
2414.5
96.1
88.8
53.7
InfiniteVL
1998.9
86.0
87.9
53.7
Table 7: We extend comparison with state-of-the-art approaches to include Qwen3.5-27B and InfiniteVL, with fixed ↓ 75.0% token compression.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Parallelism algorithm of token uniqueness. (left) We extend Gated DeltaNet with bypass input and batched queries, it enables injecting additional queries when computing output. (right) We use shifted keys as pseudo queries, and calculate token uniqueness on output matrix Ut .
Figure 8: Efficient implementation of adaptive token merging.
Vision Tokens
100%
50.0%
25.0%
12.5%
Creativity
61.0
61.0
60.6
60.0
Richness
66.4
66.7
66.0
65.2
Visual Perception
62.2
62.8
61.7
61.2
Logical Coherence
77.5
77.9
77.3
76.7
Answer Accuracy
70.2
70.7
69.7
68.8
Image Relationship
63.1
64.0
63.0
62.2
Appendix
Table 8: MMDU results on Qwen3.5-9B under varying average vision token retention rates. All dimensions are scored by the official protocol (higher is better).
Setting
Exit
Pre-exit
Direct
Relative
Overall
Early exit only
12
100%
70.43
77.63
73.30
Early exit only
24
100%
71.30
80.26
74.87
Original
32
100%
73.04
80.26
75.92
V-CoLA w/o exit
32
48.39%
66.96
78.95
71.73
V-CoLA w/ exit
24
65.22%
71.30
77.63
73.82
Appendix
Table 9: V*-Bench accuracy on Qwen3.5-9B. “Pre-exit” denotes the average vision token retention rate before the exit layer.
Retention
ViT
Overhead
Prefill
TTFT
Decode
E2E
Speedup
100%
243.3
–
1233.0
1476.3
778.6
2254.9
1.000×
50.0%
243.3
0.47
662.9
906.2
778.6
1684.8
1.338×
25.0%
243.3
0.48
402.9
646.2
778.6
1424.8
1.583×
12.5%
243.3
0.46
279.6
522.9
778.6
1301.5
1.733×
Appendix
Table 10: End-to-end latency (ms) of INT4 TensorRT Qwen3.5-9B on Jetson Thor with 2K visual input tokens and 32 output tokens. Overhead denotes V-CoLA’s compression cost, which is already included in prefill.