Visual token pruning, which aims to compress and prune redundant visual tokens, plays a critical role in efficient inference with large vision-language models (LVLMs). However, existing methods fail to disentangle intra-modal visual redundancy from cross-modal redundancy between vision and language. We show that visual token diversity and task-specific token relevance are two crucial yet orthogonal factors that complement each other in conveying useful information and should therefore be treated separately for more effective visual token pruning. Building upon this insight, we design TODRE, a two-stage and training-free framework that incorporates Token Diversity and task RElevance for effective token compression and efficient LVLM inference. Instead of pruning redundant tokens, we introduce a greedy max-sum diversification algorithm that selects and retains a subset of diverse and representative visual tokens after the vision encoder. On top of that, ToDRE leverages an ``information migration'' mechanism to eliminate task-irrelevant visual tokens within certain decoder layers of the large language model (LLM), further improving token pruning and LVLM inference. Extensive experiments show that ToDRE prunes 90% of visual tokens after the vision encoder as well as all visual tokens in certain LLM decoder layers, leading to a 2.6x speed-up in total inference time while maintaining 95.0% model performance plus excellent model compatibility. The code is available at: this https URL.
Figures & tables
Figure 1 : (a–b) Existing methods fail to disentangle intra-modal visual redundancy from cross-modal vision-language redundancy, leading to erroneous pruning [ 11 , 61 ] ; as shown in (b), attention-guided pruning discards the informative white coffee-cup region. (c) ToDRE disentangles redundancy into two orthogonal factors, visual token diversity and task-specific token relevance, and prunes them in dedicated stages, preserving critical visual cues. (d) Across eight image-language benchmarks, ToDRE consistently outperforms all baselines.
Figure 2 : Text-to-visual attention (blue) and visual-to-text attention (orange) in each LLM decoder layer. We observe a clear pattern of “information migration” : cross-modal attention (both visual-to-text and text-to-visual) is high in early layers, reflecting active information exchange, but gradually diminishes in deeper layers as the model shifts toward unimodal text reasoning.
Figure 3 : Overall framework of ToDRE. Given the visual and textual inputs, the proposed Diversity-driven Token Selection first selects a pivot token from global thumbnail or video frames with [CLS] -based attention and then performs max-sum diversification to retain a diverse set of k visual tokens. The proposed Relevance-driven Token Reduction then dynamically identifies a pivot decoder layer and prunes all its visual tokens—the layer is identified if its visual-to-text and text-to-visual attention ratios both fall below a threshold τ . EvG , EvC , and EvF denote the embeddings of thumbnail, local crops, and video frames, respectively.
Method
MME
ScienceQA
GQA
POPE
MMBench-EN
MMBench-CN
VizWiz
VQAv2
Average
Upper Bound, 2880 Tokens
LLaVA-NeXT-7B [ 38 ]
1519.6
72.0
64.2
87.7
68.5
59.0
60.6
80.1
100.0%
Ratio=25%, Retain up to 720 Tokens
FastV [ 11 ]
1477.3
69.8
60.4
83.1
65.6
55.4
57.2
77.2
95.4%
SparseVLM [ 62 ]
1446.1
67.5
60.9
71.0
63.8
55.4
58.6
77.2
93.1%
FasterVLM [ 61 ]
1454.6
67.1
61.3
87.2
66.0
56.8
58.4
76.4
96.0%
Table 1 : Performance of training-free token compression methods across eight image-language benchmarks. “Average” denotes the mean performance ratio between each token compression method and the vanilla MLLM model. We evaluate all methods at retention ratios of 25% and 10%, with the best results highlighted in bold.
Method
Retain Ratio
# Token
VideoMME
Egoschema
MLVU
LongVideoBench
Average
LLaVA-NeXT-7B [ 38 ]
0%
2880
33.3
35.7
20.1
42.5
100.0%
FastV [ 11 ]
25%
720
32.3
31.2
16.5
40.3
90.4%
FasterVLM [ 61 ]
33.8
41.0
19.6
40.5
102.3%
DivPrune [ 2 ]
33.6
41.8
19.5
40.4
102.5%
ToDRE (Ours)
33.3
42.4
19.6
41.0
103.1%
FastV [ 11 ]
10%
288
30.4
30.4
11.2
38.7
80.8%
Table 2 : Performance of training-free token compression methods across four video-language benchmarks.
Qwen2.5-VL-7B-Instruct [ 5 ]
InternVL2-8B [ 51 ]
Benchmark
Original
Ret. 25%
Ret. 10%
Original
Ret. 25%
Ret. 10%
MME [ 15 ]
1687.7
1680.6
1573.9
1628.8
1566.4
1424.4
ScienceQA [ 42 ]
88.5
86.4
85.4
96.5
95.3
92.5
GQA [ 21 ]
60.9
57.3
52.9
62.8
56.9
52.1
POPE [ 32 ]
87.7
85.4
80.8
87.8
83.7
76.9
MMBench-EN [ 41 ]
82.9
79.9
72.8
81.2
77.2
71.4
Table 3 : Performance of ToDRE on Qwen2.5-VL-7B-Instruct and InternVL2-8B. Benchmarks as rows. “Ret.” = Retention Ratio. Averages are normalized to each model’s Original (=100%).
Method
FLOPs ↓ (T)
Memory ↓ (GB)
Throughput ↑ (samples/s)
Performance ↑
Upper Bound, 2880 Tokens
LLaVA-NeXT-7B [ 38 ]
31.4
15.9
1.5
100%
Ratio=10%, Retain up to 288 Tokens
FastV [ 11 ]
8.2 ( ↓ 73.9%)
14.1 ( ↓ 11.3%)
2.1 (1.4 × )
88.8%
SparseVLM [ 62 ]
6.9 ( ↓ 78.0%)
14.1 ( ↓ 11.3%)
2.5 (1.7 × )
86.3%
FasterVLM [ 61 ]
6.1 ( ↓ 80.6%)
13.6 ( ↓ 14.5%)
2.7 (1.8 × )
91.4%
Table 4 : Inference efficiency comparisons. All experiments were conducted on a single NVIDIA RTX 3090 GPU. “Memory”: peak GPU memory usage; “Throughput”: number of POPE samples processed per second; “Performance”: average score across 8 image understanding benchmarks.
Method
Total Time ↓ (Min:Sec)
MME
ScienceQA
GQA
POPE
Average
Upper Bound, 2880 Tokens
LLaVA-NeXT-7B [ 38 ]
77:04
1519.6
72.0
64.2
87.7
100.0%
Stage 2 only
70:15
1522.7
71.9
64.3
87.6
100.0%
Ratio=25%, Retain up to 720 Tokens
Stage 1 only
48:10
1503.8
70.6
63.1
87.5
98.8%
Stage 1 + Stage 2 (ToDRE)
44:18
1504.3
70.7
63.3
87.5
98.9%
Table 5 : Ablation study on two-stage token compression. We evaluated the individual and combined effects of the proposed two-stage pruning pipeline under retention ratios of 25% and 10%.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
MME
ScienceQA
GQA
POPE
Average
Original
1519.6
72.0
64.2
87.7
100%
Retention Ratio = 25%
[CLS]
1504.3
70.7
63.3
87.5
98.9%
Random
1500.0
70.6
63.2
87.1
98.6%
Center
1502.7
70.6
63.3
87.4
98.8%
Farthest
1500.3
70.7
63.2
87.4
98.7%
Appendix
Table A1 : Ablations on Pivot Token Selection Strategy. All results are based on LLaVA-NeXT-7B. “ [CLS] ” = token with highest attention to encoder [CLS] ; “Center” = token nearest to mean visual feature; “Farthest” = farthest token from mean; “Random” = random token.
Layers
Total Time ↓ (Min:Sec)
MME
ScienceQA
GQA
POPE
Average
LLaVA-NeXT-7B [ 38 ]
77:04
1519.6
72.0
64.2
87.7
100.0%
L
78:51
1519.6
71.8
64.2
87.6
99.9%
L/2∼L
39:20
1527.8
65.4
39.0
87.1
87.9%
L/2+5L/8+6L/8+7L/8
65:22
1527.8
65.4
39.1
87.2
87.9%
L/2
64:19
1527.8
65.4
39.0
87.1
87.9%
5L/8
58:14
1528.6
71.8
54.3
87.5
96.2%
Appendix
Table A2 : Ablation study on selected layers in relevance-driven visual token reduction. L denotes the total number of decoder layers in the LLM. The bolded row corresponds to the default setting used in the main paper.
Threshold τ
Total Time ↓ (Min:Sec)
MME
ScienceQA
GQA
POPE
Average
LLaVA-NeXT-7B [ 38 ]
77:04
1519.6
72.0
64.2
87.7
100.0%
0.03
79:22
1519.6
71.7
64.2
87.6
99.9%
0.05
73:24
1530.2
71.7
64.2
87.6
100.0%
0.10 (Ours)
72:35
1530.4
71.7
64.2
87.7
100.1%
0.15
72:25
1524.0
71.7
64.2
87.6
100.0%
Appendix
Table A3 : Ablation study on threshold τ in relevance-driven visual token reduction. When both attention ratios αt→v and αv→t fall below the threshold τ , all remaining visual tokens are removed from the LLM input.
Figure A1 : Output token attention toward different input token types across LLM layers during decoding. Results are averaged over 100 samples per benchmark.
Figure A2 : Qualitative comparison of free-form video-grounded QA on the Video Detail Caption benchmark [ 10 ] . Green text highlights correctly identified events and objects; red text indicates incorrect predictions; yellow text marks missing but essential information.
Figure A3 : Supplementary visualizations comparing attention-driven and ToDRE-based token compression. The visualization is based on seven benchmarks: MME [ 15 ] , SQA [ 42 ] , GQA [ 21 ] , POPE [ 32 ] , MMBench and MMBench-CN [ 41 ] , VQAv2 [ 17 ] . Best viewed when zoomed in.
Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens, and positional distortion caused by sequence compaction. Based on these observations, we propose \textbf{VPRune}, a training-free pre-LLM pruning framework consisting of visual-only diversity selection, similarity-guided token recycling, and position-preserving restoration. Experiments on FastVLM-1.5B across multiple vision-language benchmarks demonstrate that VPRune achieves a favorable accuracy--compression trade-off, with particularly pronounced advantages under aggressive compression. Furthermore, evaluations on edge-device show that VPRune effectively reduces end-to-end inference latency while maintaining superior task performance, demonstrating its practicality for resource-constrained LVLM deployment.
Guangchuan Lv, Dianxing Shi, Dingjie Fu
Northeastern University · Beihang University · Huazhong University of Science and Technology
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning methods follow an absolute-ranking paradigm, assigning importance scores to visual tokens and retaining a fixed top-K subset. In this work, we argue that this paradigm is fundamentally brittle: attention sinks distort token importance rankings, while image redundancy and query-dependent visual evidence make fixed token budgets unreliable across inputs. We propose OccamToken, a training-free framework that replaces absolute token ranking with register-anchored relative evidence testing. Instead of asking which tokens are globally important, OccamToken evaluates whether a visual token provides information beyond a register-based reference. Our key insight is that register tokens naturally absorb low-information attention patterns, making them a stable reference for identifying genuinely informative visual evidence. Based on this principle, OccamToken performs both image-adaptive redundancy pruning and query-adaptive relevance pruning through dynamic thresholds derived from register attention. Across LLaVA-NeXT, LLaVA-v1.5, and Qwen3-VL, OccamToken consistently improves the accuracy-efficiency trade-off without additional training. Notably, on LLaVA-NeXT, it reduces 2,880 visual tokens to approximately 40 while preserving over 93% of the original accuracy, enabling stable visual token compression even in the extreme 1.4% retention regime.
Large Vision-Language Models (LVLMs) face significant computational inefficiencies caused by the large number of visual tokens. Existing visual token pruning methods mainly focus on either retaining individually important tokens or selecting mutually diverse ones. In this work, we revisit visual token pruning from a coverage perspective and formulate it as a biased attention coverage maximization problem. The key idea is to select a compact token subset whose encoder-side outgoing attention can jointly cover the image while assigning higher coverage priority to more informative regions. From this perspective, we propose ACPruner, a training-free visual token pruning framework for efficient LVLM inference. ACPruner first estimates token importance by combining intra-modal saliency and inter-modal relevance, then derives token-wise coverage from attention patterns within the vision encoder, and finally performs greedy selection to maximize the proposed coverage objective. Extensive experiments across multiple LVLM backbones, including LLaVA-1.5-7B/13B, LLaVA-NeXT-7B/13B, Qwen2.5-VL-7B, and LLaVA-OneVision-7B, show that ACPruner consistently achieves strong performance retention while delivering substantial end-to-end inference speedups.
Xu Li, Yuxuan Liang, Yi Zheng +6
College of Computer Science and Artificial Intelligence Fudan University, Shanghai, China