Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.
Figures & tables
Figure 1: Cross-Layer Attention Evolution. Attention maps across depth (i.e., Layer 2/8/12) exhibit two behaviors: the answer-critical region (yellow circle) is weakly activated in early layers but gains attention with depth, while boxed tokens (white rectangle) remain highly attended across layers.
Figure 2: Analysis of Numerical Inertia. Visualization of the attention score and rank evolution of the Top-20 tokens from the second layer, please refer to Fig. 7 for more details.
Figure 3: Analysis of Rank Discrepancy. Comparison between the final-layer key tokens and their rank distribution in early layers (2,4,8,12).
Figure 4: Overview of the DIPrune Framework. (a) The pipeline illustrates the multimodal encoding process (left) and the pruning operation at specific layers (right). (b-d) The DIPrune Mechanism: A dual importance scoring design that captures (c) Inter-layer Dynamic Semantic Evolution and evaluates (d) Intra-layer Static Feature Saliency. The two scores are jointly used to select the optimal top- K tokens.
Methods
GQA
MMB
MMB CN
MME
POPE
SQA
VQA v2
VQA Text
Average
Upper Bound, 576 Tokens
61.9
64.7
58.1
1862
85.9
69.5
78.4
58.2
100%
LLaVA-1.5 7B Retain 192 Tokens ( ↓66.7% )
FastV (ECCV’24)
52.7
61.2
57.0
1612
64.8
67.3
67.1
52.5
89.4%
PDrop (CVPR’25)
57.1
63.2
56.8
1766
82.3
68.8
75.1
56.1
96.4%
VisionZip (CVPR’25)
59.3
64.5
57.3
1767
86.4
68.9
76.8
57.3
98.6%
SparseVLM (ICML’25)
57.6
62.5
53.7
1721
83.6
69.1
75.6
56.1
96.0%
Table 1: Performance Comparison on LLaVA-v1.5-7B across Diverse Benchmarks. Results are reported under different pruning ratios. Gray shading denotes our method, and bold indicates the best results. Gray-text rows ( † ) apply baselines at the same pruning layer as DIPrune (L6/L10); DIPrune(Ours) prunes at L10 by default.
Methods
MMB
MME
POPE
SQA
Avg.
Upper Bound
82.8
2304
86.1
84.7
100%
Token Pruning Rate = 66.7%
FastV
75.7
2072
82.2
78.5
92.6%
HoloV
78.3
2093
85.0
79.8
94.6%
DIPrune (Ours)
81.4
2193
87.5
80.8
97.6%
Token Pruning Rate = 77.8%
Table 3: Performance Comparison on Qwen2.5-VL-7B.
Figure 5: Ablation studies on LLaVA-NeXT-7B.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Window Ratio
0.3
0.5
0.7
0.8
0.9
1.0
TextVQA
56.11
56.83
57.11
57.21
57.13
57.19
GQA
57.74
57.74
58.01
58.88
58.68
58.72
Appendix
Table 5: Performance under different query-window ratios on LLaVA-v1.5-7B.
Methods
GQA
MMB
MMB CN
MME
POPE
SQA
VQA Text
Average
Upper Bound, 576 Tokens
63.3
68.9
62.3
1818
85.9
72.8
61.3
100%
LLaVA-1.5 13B Retain 192 Tokens ( ↓66.7% )
FastV (ECCV’24)
59.1
54.0
51.2
1641
82.3
56.4
51.6
86.0%
VisionZip (CVPR’25)
59.1
66.9
-
1754
85.1
73.5
59.5
97.3%
SparseVLM (ICML’25)
58.7
67.4
61.0
1768
82.2
73.1
55.4
96.0%
DART (EMNLP’25)
62.1
68.2
61.4
1855
84.0
73.6
60.2
99.3%
Appendix
Table 6: Performance Comparison on LLaVA-v1.5-13B across Diverse Benchmarks. Results are reported under different pruning ratios. Gray shading denotes our method, and bold indicates the best results.
Methods
MSVD-QA
MSRVTT-QA
Average
Acc.
Score
Acc.
Score
Acc.
Score
Upper Bound (Video-LLaVA)
70.2
3.9
57.3
3.5
63.8
3.7
Retain 50% Tokens ( ↓50% )
FastV (ECCV’24)
71.0
3.9
55.0
3.5
63.0
3.7
FasterVLM (arXiv’24)
70.5
3.9
56.2
3.5
63.4
3.7
DART (EMNLP’25)
71.0
4.0
56.7
3.6
63.8
3.8
Appendix
Table 7: Video QA Evaluations on Video-LLaVA-7B under 50% Token Retention.
Method
RefCOCO
RefCOCO+
RefCOCOg
Average
val
testA
testB
val
testA
testB
val
test
Upper Bound (Qwen2.5-VL-7B)
89.45
92.56
85.16
83.50
89.02
79.15
86.76
87.24
100%
Qwen2.5-VL-7B Retain 25% Tokens ( ↓75% )
FastV (ECCV’24)
43.57
46.81
40.86
39.47
43.78
36.02
43.04
42.69
48.5%
PyramidDrop (CVPR’25)
46.46
53.83
37.23
42.29
47.76
32.81
45.32
44.91
50.4%
VScan (TMLR)
74.32
79.05
68.22
67.22
73.72
58.95
69.42
69.43
80.7%
Appendix
Table 8: Performance Comparisons on Qwen2.5-VL-7B across Three Referring Expression Comprehension Benchmarks. Results are reported under the aggressive pruning ratio (retaining 25% tokens). Gray shading denotes our method, and bold indicates the best results.
Model
Method
66.7%
77.8%
88.9%
Qwen3-VL-8B (Dense)
HoloV
2007 / 78.6 / 83.5
1956 / 76.8 / 80.5
1798 / 67.7 / 74.6
DIPrune (Ours)
2152 / 81.3 / 88.3
2029 / 78.6 / 87.8
1918 / 75.8 / 85.9
Qwen3-VL-30B-A3B (MoE)
HoloV
2076 / 80.4 / 84.5
1941 / 76.3 / 81.2
1724 / 67.1 / 71.7
DIPrune (Ours)
2278 / 85.1 / 89.1
2089 / 83.8 / 86.9
1877 / 78.8 / 75.2
Appendix
Table 9: Performance comparison on Qwen3-VL series under different pruning ratios (MME / MMBench / POPE). Gray shading denotes our method, and bold indicates the best results.
Ratio
Method
LLaVA-1.5-7B
Qwen2.5-VL-7B
Cap. ↑
CHAIR ↓
Conv. ↑
OCR ↑
Chart ↑
Cap. ↑
CHAIR ↓
Conv. ↑
OCR ↑
Chart ↑
66.7%
FastV
.496
13.45
4.27
300
17.4
.543
8.68
4.31
426
63.4
HoloV
.489
13.55
4.29
299
15.4
.599
10.13
6.32
481
57.6
DIPrune
.501
12.91
4.32
307
17.5
.613
8.66
6.40
492
71.8
77.8%
FastV
.485
13.20
4.07
287
16.5
.520
8.23
4.24
332
57.8
HoloV
.475
14.00
4.16
291
17.0
.581
10.12
6.21
360
44.7
Appendix
Table 10: Results on generative and fine-grained benchmarks. Cap.: DetailCaps (CAPTURE); CHAIR: CHAIR i ; Conv.: ConvBench; OCR: OCRBench; Chart: ChartQA (%).
Length
#Samples
Full
FastV
HoloV
DIPrune
3–5
2,071
51.1
41.7
46.2
48.2
6–11
8,425
62.9
44.9
56.0
60.1
12–25
2,082
68.8
55.3
60.3
64.7
All
12,578
61.9
46.1
55.1
58.9
Appendix
Table 11: Stratified analysis on LLaVA-1.5-7B (88.9% pruning). (a) GQA accuracy (%) by query length; (b) CHAIR i ( ↓ ) by the number of ground-truth instances. Bold: best among pruning methods.
Table 15
Metrics
1.0 : 0.0
0.8 : 0.2
0.6 : 0.4
0.5 : 0.5
0.4 : 0.6
0.2 : 0.8
0.0 : 1.0
MME
1943
2046
2193
2134
2097
2019
1984
MMB
79.3
80.7
81.4
81.2
80.9
80.6
78.8
POPE
84.4
85.9
87.5
86.4
86.2
85.6
83.7
Appendix
Table 15: Extended Sensitivity Analysis on Qwen2.5-VL-7B. We evaluate the performance trajectory across a spectrum of fusion ratios α:β . The default setting ( α=0.6,β=0.4 ) consistently achieves high performance. Underlined ratios correspond to using only Salign or Sevol .
Method
432 Tokens
288 Tokens
32 Tokens
16 Tokens
VisionZip
58.08 / 61.57 / 1836
57.92 / 60.32 / 1811
53.03 / 52.15 / 1580
49.76 / 46.95 / 1350
HoloV
52.24 / 55.83 / 1605
54.63 / 57.83 / 1620
53.55 / 52.37 / 1596
50.40 / 48.60 / 1430
DIPrune (Ours)
58.32 / 62.00 / 1862
58.25 / 61.93 / 1828
55.96 / 56.20 / 1756
53.84 / 54.60 / 1712
Appendix
Table 16: Performance comparison across extended token budgets on LLaVA-v1.5-7B (TextVQA / GQA / MME). Bold indicates the best results.
Method
LLaVA-1.5-7B
LLaVA-NeXT-7B
Qwen2.5-VL-7B
TTFT ↓
ITT ↓
Acc ↑
TTFT ↓
ITT ↓
Acc ↑
TTFT ↓
ITT ↓
Acc ↑
Baseline
73.23
16.21
100%
102.10
17.44
100%
84.28
16.96
100%
FastV
32.43
14.95
74.4%
43.82
15.43
87.2%
53.78
15.46
87.6%
SparseVLM
33.43
14.98
86.3%
45.25
15.46
90.6%
54.35
15.44
–
HoloV
31.83
14.94
93.7%
43.47
15.44
95.8%
53.52
15.47
90.5%
DIPrune
32.50
14.96
97.7%
43.98
15.40
98.2%
54.08
15.46
93.3%
Appendix
Table 17: Efficiency across architectures (88.9% pruning ratio). TTFT (ms) and ITT (ms/token) are measured on POPE. Overhead : DIPrune’s pruning cost, split into proxy scoring and selection, and its ratio to TTFT. –: not reported.
Method
Mem (MiB)
Time (m:s)
Lat. (ms)
Acc (%)
SparseVLM
19,458
11:56
56.9
57.6
PDrop
15,616
11:56
56.9
57.1
DIPrune
15,174
11:57
57.0
61.2
Appendix
Table 18: End-to-end efficiency on LLaVA-v1.5-7B (66.7% pruning ratio). Mem: peak GPU memory; Time: total wall-clock time; Lat.: average wall-clock time per sample.
Figure 6: Visualization of visual tokens retained by LLaVA-v1.5-7B on VQAv2 samples. We present comparison results under three progressive pruning ratios (66.7%, 77.8%, and 88.9%). The red bounding boxes in the original images highlight the task-critical visual regions directly relevant to the text query. In the pruning results, the red and white dashed boxes indicate the retention status of these key regions by the DIPrune and FastV strategies, respectively.
Figure 7: Analysis of Numerical Inertia. Visualization of the attention score and rank evolution of the Top-50 tokens from the second layer. The heatmap shows that high-ranking tokens (green regions) in shallow layers tend to maintain their dominance throughout the network.
While visual token pruning is essential for efficient Multimodal Large Language Models (MLLMs), existing training-free methods suffer from a critical limitation: they rely on static, instantaneous heuristics to perform irreversible filtering. This approach ignores the hierarchical nature of MLLMs, where token importance often evolves dynamically rather than remaining fixed across layers. Consequently, tokens essential for deep-layer reasoning are often prematurely discarded by shallow-layer estimates. To address this, we propose Trend-aware Pruning, a novel framework that elevates pruning from a local snapshot decision to a temporal trajectory modeling problem. Instead of relying on isolated scores, our method captures the momentum of attention flow. This enables a dynamic rectification mechanism that selectively reactivates "late-blooming" tokens, those initially undervalued but exhibiting rising semantic importance, thereby preventing the loss of critical visual cues. Extensive experiments demonstrate that our approach achieves a superior efficiency-performance trade-off across diverse multimodal tasks. Notably, it reduces visual tokens by over 77.8%, retaining only approximately 23 tokens in the final layer while maintaining competitive performance, offering a robust and reversible solution for high-efficiency multimodal inference.
Jie Ma, Zhike Qiu, Jie Gao +4
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University · School of Information Engineering, Xiamen Ocean Vocational College · Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, Sino-Russian Research Center for Digital Economy
Visual token pruning reduces the inference overhead of multimodal large language models (MLLMs) by retaining only a subset of visual tokens. Existing methods usually select tokens based on importance or redundancy. However, we observe that these criteria produce stable spatial biases across inputs and do not always outperform simple Uniform Grid sampling, highlighting the value of broad spatial coverage. Motivated by this, we propose S2Prune, a training-free pruning method that preserves spatial coverage while adapting token density to local image structure. We first divide the image into regions and assign at least one token to each region to preserve coverage. The remaining token budget is then distributed according to Laplacian variation, giving more tokens to regions with richer structure. We then use Early Representation Change (ERC), computed from the first decoder block, to select representative tokens within each region. We evaluate S2Prune across diverse settings and two MLLM architectures. On Qwen2.5-VL-7B-Instruct, it achieves the highest average accuracy among the evaluated training-free pruning methods. With only 32 of the original 576 visual tokens, it still retains 79.3% of the full-model performance. Code is available at https://github.com/yuanyuanjia71-spec/S2Prune.
Yuanyuan Jia, Shunpu Tang, Qianqian Yang
College of Information Science and Electronic Engineering, Zhejiang University Hangzhou, China
Recent multimodal large language models (MLLMs), such as Qwen2.5-VL and InternVL3, generate large numbers of vision tokens for high-resolution inputs, leading to substantial computational cost. Existing vision token pruning methods either depend on cross-modal attention and cannot prune before the prefill stage, or rely on diversity estimation with high computational overhead. We observe that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities. Based on this observation, we propose SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informative vision tokens. SepPrune reuses the LLM's built-in projection parameters and requires no architectural changes. Experiments on Qwen2.5-VL-7B show that SepPrune achieves state-of-the-art performance, retaining 96.3% of the original accuracy while removing 80.2% of vision tokens.
Yuchen Wang, Qihui Zhu, Yang Liu +2
MoE Key Lab of BIPC, NEL-BITA, University of Science and Technology of China, Hefei, China · ChangXin Memory Technologies, Hefei, China · APKL of BIIP, IAI, Hefei Comprehensive National Science Center, Hefei, China