Organizations: Global Innovation Exchange, Tsinghua University, Beijing, China · Department of Automation, Tsinghua University, Beijing, China · School of Computer Science, Peking University, Beijing, China
Vision-language models (VLMs) have achieved impressive performance on multimodal inference tasks, but the cost remains a significant challenge due to the large number of vision tokens processed during the prefill stage. Existing token pruning methods often rely on utilizing the static attention patterns directly, failing to exploit the dynamic internal signals within VLMs. To address the issue, we propose AdaptInfer, a novel plug-and-play framework for vision token pruning. First, we introduce a dynamic text-guided pruning mechanism that construct soft priors over text-token importance on each pruning layer, allowing more informed scoring of vision tokens at each stage. Second, we observe a highly consistent distribution of cross-modal attention shifts by architecture, which inspires us to introduce a efficient data-driven schedule which determines the pruning locations automatically. Experimental results have verified the effectiveness and generalization of the proposed method. Under the same token budget, AdaptInfer surpasses existing approaches in accuracy. For instance, AdaptInfer maintains averagely 99.4% accuracy on Qwen2-VL while 70% of the prefilling vision token overhead is reduced. The source code is available on: https://github.com/weiczh02/AdaptInfer-base.
Figures & tables
Figure 1: MIoU of key text tokens in different layers. The consistent low mIoU shows the key text token varies across layers.
Figure 2: Highly consistent distribution of attention shifts on MME and TextVQA.
Figure 3: The architecture of AdaptInfer . Text token importance is computed and guides vision token selection at every pruning layer adaptively. The right panel illustrates the implementation details of the Rank & Prune module.
Method
Avg. Tokens
MME
GQA
MMB
SQA
TextVQA
MMVet
Ratio (%)
Vanilla
576
1864
61.9
64.6
69.5
58.3
31.1
100
ToMe (ICLR 23)
128
1343
52.4
53.3
59.6
49.1
27.2
82.8 ( ↓ 17.2)
FastV (ECCV 24)
128
1490
49.6
56.1
68.6
52.5
26.3
87.9 ( ↓ 12.1)
PDrop (CVPR 25)
128
1761
57.1
61.6
68.4
56.6
28.1
94.7 ( ↓ 5.3)
SparseVLM (ICML 25)
128
1746
58.4
64.5
68.6
56.7
29.0
96.2 ( ↓ 3.8)
DART (EMNLP 25)
128
1804
57.7
62.3
69.4
55.9
29.8
96.3 ( ↓ 3.7)
Table 1: Comparison of methods under different pruning budgets on LLava-1.5-7B .
Methods
Tokens
Image-Based
Video-Based
Ratio
MME
TextVQA
GQA
MMB
POPE
TGIF
MSRVTT
MSVD
Vanilla
100%
1901
77.8
60.4
71.7
86.9
9.9
30.9
42.3
100%
SparseVLM
50%
1900
76.9
59.9
71.0
86.4
11.7
31.2
42.9
102.1%
AdaptInfer
50%
1909
76.6
60.0
72.1
86.7
11.1
30.5
43.6
101.7%
SparseVLM
30%
1867
73.6
57.5
70.1
84.6
9.3
29.5
41.8
96.4%
AdaptInfer
30%
1885
73.9
58.5
71.4
85.8
10.4
30.2
43.5
99.4%
Table 2: Comparison of Performance on Qwen2-VL-2B.
Figure 4: Performance trends of AdaptInfer on two representative benchmarks.
Method
Tokens
MME
TextVQA
Ratio (%)
Vanilla
576
1864
58.3
100
SparseVLM
48
1416
48.1
79.2
AdaptInfer
48
1494
52.8
85.4
SparseVLM
32
1098
47.7
70.4
AdaptInfer
32
1190
49.3
74.2
Table 3: Extreme pruning comparsions.
Method
Tokens
TVQA
MME
MMB
POPE
GQA
Ratio
Vanilla
576
58.3
1864
64.6
85.9
61.9
100%
VisionZip
64
55.5
1690
60.1
77.0
55.1
91.5%
VisionZip + AdaptInfer
64
56.5
1724
62.8
83.3
56.6
95.0%
Table 4: Complementarity with vision-encoder-side compression.
Method
Tokens
Metrics
MME
GQA
MMB
TextVQA
Average
Vanilla
576
FLOPs (T)
4.268
4.250
4.623
4.611
4.438
Latency (ms)
82.0
76.5
91.3
91.2
85.3
PDrop
64
FLOPs (T)
0.975
0.958
1.316
1.305
1.138 ( ↓ 74.3%)
Latency (ms)
33.0
32.0
36.4
36.6
34.5 ( ↓ 59.6%)
SparseVLM
64
FLOPs (T)
0.974
0.958
1.316
1.305
1.138 ( ↓ 74.3%)
Latency (ms)
34.4
34.7
38.1
39.5
36.7 ( ↓ 57.0%)
Table 5: Latency test of AdaptInfer.
Tokens
Text Prior
MME
TextVQA
SQA
POPE
30%
Single-token
1878
35.9
86.3
84.7
30%
Random
1990
48.6
88.6
86.4
30%
Static
1989
49.3
88.7
86.9
30%
Dynamic
2015
49.4
89.5
87.3
Table 6: Controlled ablation of different text priors on InternVL3-2B.
Tokens
K
Pruning Layers
MME
MMB
64
3
1,10,20
1684
61.7
64
2
1,20
1520
56.3
64
4
1,7,14,20
1645
60.7
64
3
0,10,20
1660
60.7
64
3
2,10,20
1589
60.1
64
3
1,9,20
1642
60.9
Table 7: Performance of different pruning locations on LLava-1.5-7B.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Layer-wise distribution of attention shifts on LLava-1.5-13B.
Tokens
MME
MMB
TextVQA
Ratio (%)
576
1826
61.2
68.5
100
128
1813
59.8
67.4
98.5
64
1735
58.1
66.2
95.5
Appendix
Table 8: Performance of AdaptInfer on LLava-1.5-13B. Our method still maintain high overall accuracy.
Figure 6: Layer-wise distribution of attention shifts on Qwen2-VL-2B on MME.
Task
Vanilla
AdaptInfer
Ratio
Existence
200
200
100%
Count
145
135
93.1%
Position
158
158
100%
Color
190
180
94.7%
Posters
125
123
98.4%
Celebrity
138
136
98.6%
Appendix
Table 9: Pruning-tolerance evaluation on MME. For AdaptInfer, 30% tokens are retained averagely.
C to C
C to E
E to C
E to E
Percentage
78.2%
1.7%
1.3%
18.8%
Appendix
Table 10: Instance-Level change before and after pruning. C: Correct; E: Error.
Figure 7: Visualization of AdaptInfer on different examples.
Figure 8: Examples where AdaptInfer outperforms SparseVLM. AdaptInfer performs better when text prompts are rich in descriptive modifiers.
Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in visual understanding and reasoning, but they also impose significant computational burdens due to long visual sequence inputs. Recent works address this issue by pruning unimportant visual tokens, achieving substantial computational reduction while maintaining model performance. The core of token pruning lies in determining token importance, with current approaches primarily relying on attention scores from vision encoders or Large Language Models (LLMs). In this paper, we analyze the effectiveness of attention mechanisms in both vision encoders and LLMs. We find that vision encoders suffer from attention sink, leading to poor focus on informative foreground regions, while in LLMs, although prior studies have identified attention bias toward token positions, text-to-vision attention demonstrates resistance to this bias and enables effective pruning guidance in middle layers. Based on these observations, we propose LearnPruner, a two-stage token pruning framework that first removes redundant vision tokens via a learnable pruning module after the vision encoder, then retains only task-relevant tokens in the LLM's middle layer. Experimental results show that our LearnPruner can preserve approximately 95% of the original performance while using only 5.5% of vision tokens, and achieve 3.2× inference acceleration, demonstrating a superior accuracy-efficiency trade-off.
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning methods follow an absolute-ranking paradigm, assigning importance scores to visual tokens and retaining a fixed top-K subset. In this work, we argue that this paradigm is fundamentally brittle: attention sinks distort token importance rankings, while image redundancy and query-dependent visual evidence make fixed token budgets unreliable across inputs. We propose OccamToken, a training-free framework that replaces absolute token ranking with register-anchored relative evidence testing. Instead of asking which tokens are globally important, OccamToken evaluates whether a visual token provides information beyond a register-based reference. Our key insight is that register tokens naturally absorb low-information attention patterns, making them a stable reference for identifying genuinely informative visual evidence. Based on this principle, OccamToken performs both image-adaptive redundancy pruning and query-adaptive relevance pruning through dynamic thresholds derived from register attention. Across LLaVA-NeXT, LLaVA-v1.5, and Qwen3-VL, OccamToken consistently improves the accuracy-efficiency trade-off without additional training. Notably, on LLaVA-NeXT, it reduces 2,880 visual tokens to approximately 40 while preserving over 93% of the original accuracy, enabling stable visual token compression even in the extreme 1.4% retention regime.
Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score that measures whether a token is useful. Existing methods typically rely on Gumbel-Softmax to approximate discrete selection during training. Such selectors make the score depend on the behavior of a relaxed pruning operator, not directly on the consequence of information loss. In this paper, we propose DiffPrune, which gives token scores a direct meaning. During training, DiffPrune keeps all tokens and weakens each token's information according to its score. If weakening a token hurts the task, the scorer is pushed to protect it; if not, the token can receive a lower score. Because the loss is differentiated through this actual information-throttling path, the scorer avoids the unstable surrogate path of relaxed token selection. DiffPrune implements this idea with an Information Throttler, which injects variance-preserving noise into visual tokens, where high-score tokens remain close to their original representations, while low-score tokens carry less original information. At inference, the throttler is removed, and hard top-K pruning is applied using the learned scores. Across ten VLM benchmarks, DiffPrune retains 96.5% of full-model accuracy while accelerating LLM prefill by 2.85x, with only 0.69 ms inference overhead. Code will be publicly available.
Landi He, Mingde Yao, Shawn Young +1
Shenzhen University of Advanced Technology · CUHK MMLab, CPII under InnoHK