Organizations: Global Innovation Exchange, Tsinghua University, Beijing, China · Department of Automation, Tsinghua University, Beijing, China · School of Computer Science, Peking University, Beijing, China
Vision-language models (VLMs) have achieved impressive performance on multimodal inference tasks, but the cost remains a significant challenge due to the large number of vision tokens processed during the prefill stage. Existing token pruning methods often rely on utilizing the static attention patterns directly, failing to exploit the dynamic internal signals within VLMs. To address the issue, we propose AdaptInfer, a novel plug-and-play framework for vision token pruning. First, we introduce a dynamic text-guided pruning mechanism that construct soft priors over text-token importance on each pruning layer, allowing more informed scoring of vision tokens at each stage. Second, we observe a highly consistent distribution of cross-modal attention shifts by architecture, which inspires us to introduce a efficient data-driven schedule which determines the pruning locations automatically. Experimental results have verified the effectiveness and generalization of the proposed method. Under the same token budget, AdaptInfer surpasses existing approaches in accuracy. For instance, AdaptInfer maintains averagely 99.4% accuracy on Qwen2-VL while 70% of the prefilling vision token overhead is reduced. The source code is available on: https://github.com/weiczh02/AdaptInfer-base.
Figures & tables
Figure 1: MIoU of key text tokens in different layers. The consistent low mIoU shows the key text token varies across layers.
Figure 2: Highly consistent distribution of attention shifts on MME and TextVQA.
Figure 3: The architecture of AdaptInfer . Text token importance is computed and guides vision token selection at every pruning layer adaptively. The right panel illustrates the implementation details of the Rank & Prune module.
Method
Avg. Tokens
MME
GQA
MMB
SQA
TextVQA
MMVet
Ratio (%)
Vanilla
576
1864
61.9
64.6
69.5
58.3
31.1
100
ToMe (ICLR 23)
128
1343
52.4
53.3
59.6
49.1
27.2
82.8 ( ↓ 17.2)
FastV (ECCV 24)
128
1490
49.6
56.1
68.6
52.5
26.3
87.9 ( ↓ 12.1)
PDrop (CVPR 25)
128
1761
57.1
61.6
68.4
56.6
28.1
94.7 ( ↓ 5.3)
SparseVLM (ICML 25)
128
1746
58.4
64.5
68.6
56.7
29.0
96.2 ( ↓ 3.8)
DART (EMNLP 25)
128
1804
57.7
62.3
69.4
55.9
29.8
96.3 ( ↓ 3.7)
Table 1: Comparison of methods under different pruning budgets on LLava-1.5-7B .
Methods
Tokens
Image-Based
Video-Based
Ratio
MME
TextVQA
GQA
MMB
POPE
TGIF
MSRVTT
MSVD
Vanilla
100%
1901
77.8
60.4
71.7
86.9
9.9
30.9
42.3
100%
SparseVLM
50%
1900
76.9
59.9
71.0
86.4
11.7
31.2
42.9
102.1%
AdaptInfer
50%
1909
76.6
60.0
72.1
86.7
11.1
30.5
43.6
101.7%
SparseVLM
30%
1867
73.6
57.5
70.1
84.6
9.3
29.5
41.8
96.4%
AdaptInfer
30%
1885
73.9
58.5
71.4
85.8
10.4
30.2
43.5
99.4%
Table 2: Comparison of Performance on Qwen2-VL-2B.
Figure 4: Performance trends of AdaptInfer on two representative benchmarks.
Method
Tokens
MME
TextVQA
Ratio (%)
Vanilla
576
1864
58.3
100
SparseVLM
48
1416
48.1
79.2
AdaptInfer
48
1494
52.8
85.4
SparseVLM
32
1098
47.7
70.4
AdaptInfer
32
1190
49.3
74.2
Table 3: Extreme pruning comparsions.
Method
Tokens
TVQA
MME
MMB
POPE
GQA
Ratio
Vanilla
576
58.3
1864
64.6
85.9
61.9
100%
VisionZip
64
55.5
1690
60.1
77.0
55.1
91.5%
VisionZip + AdaptInfer
64
56.5
1724
62.8
83.3
56.6
95.0%
Table 4: Complementarity with vision-encoder-side compression.
Method
Tokens
Metrics
MME
GQA
MMB
TextVQA
Average
Vanilla
576
FLOPs (T)
4.268
4.250
4.623
4.611
4.438
Latency (ms)
82.0
76.5
91.3
91.2
85.3
PDrop
64
FLOPs (T)
0.975
0.958
1.316
1.305
1.138 ( ↓ 74.3%)
Latency (ms)
33.0
32.0
36.4
36.6
34.5 ( ↓ 59.6%)
SparseVLM
64
FLOPs (T)
0.974
0.958
1.316
1.305
1.138 ( ↓ 74.3%)
Latency (ms)
34.4
34.7
38.1
39.5
36.7 ( ↓ 57.0%)
Table 5: Latency test of AdaptInfer.
Tokens
Text Prior
MME
TextVQA
SQA
POPE
30%
Single-token
1878
35.9
86.3
84.7
30%
Random
1990
48.6
88.6
86.4
30%
Static
1989
49.3
88.7
86.9
30%
Dynamic
2015
49.4
89.5
87.3
Table 6: Controlled ablation of different text priors on InternVL3-2B.
Tokens
K
Pruning Layers
MME
MMB
64
3
1,10,20
1684
61.7
64
2
1,20
1520
56.3
64
4
1,7,14,20
1645
60.7
64
3
0,10,20
1660
60.7
64
3
2,10,20
1589
60.1
64
3
1,9,20
1642
60.9
Table 7: Performance of different pruning locations on LLava-1.5-7B.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Layer-wise distribution of attention shifts on LLava-1.5-13B.
Tokens
MME
MMB
TextVQA
Ratio (%)
576
1826
61.2
68.5
100
128
1813
59.8
67.4
98.5
64
1735
58.1
66.2
95.5
Appendix
Table 8: Performance of AdaptInfer on LLava-1.5-13B. Our method still maintain high overall accuracy.
Figure 6: Layer-wise distribution of attention shifts on Qwen2-VL-2B on MME.
Task
Vanilla
AdaptInfer
Ratio
Existence
200
200
100%
Count
145
135
93.1%
Position
158
158
100%
Color
190
180
94.7%
Posters
125
123
98.4%
Celebrity
138
136
98.6%
Appendix
Table 9: Pruning-tolerance evaluation on MME. For AdaptInfer, 30% tokens are retained averagely.
C to C
C to E
E to C
E to E
Percentage
78.2%
1.7%
1.3%
18.8%
Appendix
Table 10: Instance-Level change before and after pruning. C: Correct; E: Error.
Figure 7: Visualization of AdaptInfer on different examples.
Figure 8: Examples where AdaptInfer outperforms SparseVLM. AdaptInfer performs better when text prompts are rich in descriptive modifiers.