Visual token pruning, which aims to compress and prune redundant visual tokens, plays a critical role in efficient inference with large vision-language models (LVLMs). However, existing methods fail to disentangle intra-modal visual redundancy from cross-modal redundancy between vision and language. We show that visual token diversity and task-specific token relevance are two crucial yet orthogonal factors that complement each other in conveying useful information and should therefore be treated separately for more effective visual token pruning. Building upon this insight, we design TODRE, a two-stage and training-free framework that incorporates Token Diversity and task RElevance for effective token compression and efficient LVLM inference. Instead of pruning redundant tokens, we introduce a greedy max-sum diversification algorithm that selects and retains a subset of diverse and representative visual tokens after the vision encoder. On top of that, ToDRE leverages an ``information migration'' mechanism to eliminate task-irrelevant visual tokens within certain decoder layers of the large language model (LLM), further improving token pruning and LVLM inference. Extensive experiments show that ToDRE prunes 90% of visual tokens after the vision encoder as well as all visual tokens in certain LLM decoder layers, leading to a 2.6x speed-up in total inference time while maintaining 95.0% model performance plus excellent model compatibility. The code is available at: this https URL.
Figures & tables
Figure 1 : (a–b) Existing methods fail to disentangle intra-modal visual redundancy from cross-modal vision-language redundancy, leading to erroneous pruning [ 11 , 61 ] ; as shown in (b), attention-guided pruning discards the informative white coffee-cup region. (c) ToDRE disentangles redundancy into two orthogonal factors, visual token diversity and task-specific token relevance, and prunes them in dedicated stages, preserving critical visual cues. (d) Across eight image-language benchmarks, ToDRE consistently outperforms all baselines.
Figure 2 : Text-to-visual attention (blue) and visual-to-text attention (orange) in each LLM decoder layer. We observe a clear pattern of “information migration” : cross-modal attention (both visual-to-text and text-to-visual) is high in early layers, reflecting active information exchange, but gradually diminishes in deeper layers as the model shifts toward unimodal text reasoning.
Figure 3 : Overall framework of ToDRE. Given the visual and textual inputs, the proposed Diversity-driven Token Selection first selects a pivot token from global thumbnail or video frames with [CLS] -based attention and then performs max-sum diversification to retain a diverse set of k visual tokens. The proposed Relevance-driven Token Reduction then dynamically identifies a pivot decoder layer and prunes all its visual tokens—the layer is identified if its visual-to-text and text-to-visual attention ratios both fall below a threshold τ . EvG , EvC , and EvF denote the embeddings of thumbnail, local crops, and video frames, respectively.
Method
MME
ScienceQA
GQA
POPE
MMBench-EN
MMBench-CN
VizWiz
VQAv2
Average
Upper Bound, 2880 Tokens
LLaVA-NeXT-7B [ 38 ]
1519.6
72.0
64.2
87.7
68.5
59.0
60.6
80.1
100.0%
Ratio=25%, Retain up to 720 Tokens
FastV [ 11 ]
1477.3
69.8
60.4
83.1
65.6
55.4
57.2
77.2
95.4%
SparseVLM [ 62 ]
1446.1
67.5
60.9
71.0
63.8
55.4
58.6
77.2
93.1%
FasterVLM [ 61 ]
1454.6
67.1
61.3
87.2
66.0
56.8
58.4
76.4
96.0%
Table 1 : Performance of training-free token compression methods across eight image-language benchmarks. “Average” denotes the mean performance ratio between each token compression method and the vanilla MLLM model. We evaluate all methods at retention ratios of 25% and 10%, with the best results highlighted in bold.
Method
Retain Ratio
# Token
VideoMME
Egoschema
MLVU
LongVideoBench
Average
LLaVA-NeXT-7B [ 38 ]
0%
2880
33.3
35.7
20.1
42.5
100.0%
FastV [ 11 ]
25%
720
32.3
31.2
16.5
40.3
90.4%
FasterVLM [ 61 ]
33.8
41.0
19.6
40.5
102.3%
DivPrune [ 2 ]
33.6
41.8
19.5
40.4
102.5%
ToDRE (Ours)
33.3
42.4
19.6
41.0
103.1%
FastV [ 11 ]
10%
288
30.4
30.4
11.2
38.7
80.8%
Table 2 : Performance of training-free token compression methods across four video-language benchmarks.
Qwen2.5-VL-7B-Instruct [ 5 ]
InternVL2-8B [ 51 ]
Benchmark
Original
Ret. 25%
Ret. 10%
Original
Ret. 25%
Ret. 10%
MME [ 15 ]
1687.7
1680.6
1573.9
1628.8
1566.4
1424.4
ScienceQA [ 42 ]
88.5
86.4
85.4
96.5
95.3
92.5
GQA [ 21 ]
60.9
57.3
52.9
62.8
56.9
52.1
POPE [ 32 ]
87.7
85.4
80.8
87.8
83.7
76.9
MMBench-EN [ 41 ]
82.9
79.9
72.8
81.2
77.2
71.4
Table 3 : Performance of ToDRE on Qwen2.5-VL-7B-Instruct and InternVL2-8B. Benchmarks as rows. “Ret.” = Retention Ratio. Averages are normalized to each model’s Original (=100%).
Method
FLOPs ↓ (T)
Memory ↓ (GB)
Throughput ↑ (samples/s)
Performance ↑
Upper Bound, 2880 Tokens
LLaVA-NeXT-7B [ 38 ]
31.4
15.9
1.5
100%
Ratio=10%, Retain up to 288 Tokens
FastV [ 11 ]
8.2 ( ↓ 73.9%)
14.1 ( ↓ 11.3%)
2.1 (1.4 × )
88.8%
SparseVLM [ 62 ]
6.9 ( ↓ 78.0%)
14.1 ( ↓ 11.3%)
2.5 (1.7 × )
86.3%
FasterVLM [ 61 ]
6.1 ( ↓ 80.6%)
13.6 ( ↓ 14.5%)
2.7 (1.8 × )
91.4%
Table 4 : Inference efficiency comparisons. All experiments were conducted on a single NVIDIA RTX 3090 GPU. “Memory”: peak GPU memory usage; “Throughput”: number of POPE samples processed per second; “Performance”: average score across 8 image understanding benchmarks.
Method
Total Time ↓ (Min:Sec)
MME
ScienceQA
GQA
POPE
Average
Upper Bound, 2880 Tokens
LLaVA-NeXT-7B [ 38 ]
77:04
1519.6
72.0
64.2
87.7
100.0%
Stage 2 only
70:15
1522.7
71.9
64.3
87.6
100.0%
Ratio=25%, Retain up to 720 Tokens
Stage 1 only
48:10
1503.8
70.6
63.1
87.5
98.8%
Stage 1 + Stage 2 (ToDRE)
44:18
1504.3
70.7
63.3
87.5
98.9%
Table 5 : Ablation study on two-stage token compression. We evaluated the individual and combined effects of the proposed two-stage pruning pipeline under retention ratios of 25% and 10%.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
MME
ScienceQA
GQA
POPE
Average
Original
1519.6
72.0
64.2
87.7
100%
Retention Ratio = 25%
[CLS]
1504.3
70.7
63.3
87.5
98.9%
Random
1500.0
70.6
63.2
87.1
98.6%
Center
1502.7
70.6
63.3
87.4
98.8%
Farthest
1500.3
70.7
63.2
87.4
98.7%
Appendix
Table A1 : Ablations on Pivot Token Selection Strategy. All results are based on LLaVA-NeXT-7B. “ [CLS] ” = token with highest attention to encoder [CLS] ; “Center” = token nearest to mean visual feature; “Farthest” = farthest token from mean; “Random” = random token.
Layers
Total Time ↓ (Min:Sec)
MME
ScienceQA
GQA
POPE
Average
LLaVA-NeXT-7B [ 38 ]
77:04
1519.6
72.0
64.2
87.7
100.0%
L
78:51
1519.6
71.8
64.2
87.6
99.9%
L/2∼L
39:20
1527.8
65.4
39.0
87.1
87.9%
L/2+5L/8+6L/8+7L/8
65:22
1527.8
65.4
39.1
87.2
87.9%
L/2
64:19
1527.8
65.4
39.0
87.1
87.9%
5L/8
58:14
1528.6
71.8
54.3
87.5
96.2%
Appendix
Table A2 : Ablation study on selected layers in relevance-driven visual token reduction. L denotes the total number of decoder layers in the LLM. The bolded row corresponds to the default setting used in the main paper.
Threshold τ
Total Time ↓ (Min:Sec)
MME
ScienceQA
GQA
POPE
Average
LLaVA-NeXT-7B [ 38 ]
77:04
1519.6
72.0
64.2
87.7
100.0%
0.03
79:22
1519.6
71.7
64.2
87.6
99.9%
0.05
73:24
1530.2
71.7
64.2
87.6
100.0%
0.10 (Ours)
72:35
1530.4
71.7
64.2
87.7
100.1%
0.15
72:25
1524.0
71.7
64.2
87.6
100.0%
Appendix
Table A3 : Ablation study on threshold τ in relevance-driven visual token reduction. When both attention ratios αt→v and αv→t fall below the threshold τ , all remaining visual tokens are removed from the LLM input.
Figure A1 : Output token attention toward different input token types across LLM layers during decoding. Results are averaged over 100 samples per benchmark.
Figure A2 : Qualitative comparison of free-form video-grounded QA on the Video Detail Caption benchmark [ 10 ] . Green text highlights correctly identified events and objects; red text indicates incorrect predictions; yellow text marks missing but essential information.
Figure A3 : Supplementary visualizations comparing attention-driven and ToDRE-based token compression. The visualization is based on seven benchmarks: MME [ 15 ] , SQA [ 42 ] , GQA [ 21 ] , POPE [ 32 ] , MMBench and MMBench-CN [ 41 ] , VQAv2 [ 17 ] . Best viewed when zoomed in.