SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models
Authors: Tianxiang Chen, Zhentao Tan, Zi Ye, Yue Wu, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Tao Gong, +4 more
Organizations: End-User Intelligent Computing BU, Alibaba Cloud · Alibaba Group · Department of Computer Science, Maynooth University · School of Cyber Science and Technology, University of Science and Technology of China · Fudan University
Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using blockwise importance, the finer-grained inter-layer representation shifts and the distribution differences within the layers themselves have not been fully explored. In this work, we comprehensively investigate this dual-level inefficiency. We posit that intermediate layer tokens from vision encoders should be considered for effective visual token pruning, as semantic focus shifts across layers, with middle-layer tokens capturing more detailed object-centric information that deeper layers may abstract away. Furthermore, we reveal the differential contributions of Attention and FFNs across distinct LLM decoder layers. Building upon these discoveries, we propose \textbf{SPIDER}, a training-free framework that integrates multi-layer \underline{\textbf{S}}emantic visual token \underline{\textbf{P}}run\underline{\textbf{I}}ng with an a\underline{\textbf{D}}aptive sub-lay\underline{\textbf{ER}} skipping mechanism. Experimental evaluations demonstrate that SPIDER consistently maintains strong performance across various MLLM architectures and reduction ratios. For instance, on LLaVA-NeXT-7B, SPIDER reduces FLOPs by 79% while maintaining 96% of the baseline performance.
Figures & tables
Fig. 1 : Visualizations to show the importance of mid-layer features for fine-grained token pruning. We highlight the top 25 % most attentive tokens (in purple) from various layers of the CLIP vision encoder, together with their t-SNE embeddings. Mid-layer tokens exhibit a higher key object coverage ratio ( R∗ ), indicating a stronger focus on crucial object-centric details compared to deeper layers, which tend to capture broader global semantics. Averaged over 100 sampled evaluation images, R∗ peaks at 29.65% at layer 12, compared to 11.28% at layer 1 and 15.83% at layer 24.
Fig. 2 : Visualization of sub-layer skipping, where we visualize either (a) skipping attention sub-layer module or (c) FFN sub-layer module. Sub-layer skipping is achieved via replacing the LLM decoder block with corresponding sparse blocks, so that visual tokens are blocked in attention or FFN modules.
Fig. 3 : The sub-layer contribution scores ( SLC ) of LLaVA-1.5-7B and LLaVA-NeXT-7B. Lower SLC values indicate a weaker influence of the corresponding attention or FFN sub-layer on the specified tokens. Skipping transformations of visual tokens in such low-impact sub-layers yields minimal divergence from the original model’s output distribution, and the SLC distribution is very different across various benchmarks.
Fig. 4 : Illustration of SPIDER . We begin by token pruning using the semantic information from the middle and last layers. The retained tokens are fed to the LLM for adaptive sub-layer skipping.
Method
TFLOPs
Ratio
VQAv2
GQA
MMStar
MME
MMB
POPE
MMVet
TextVQA
DocVQA
Acc. (%)
LLaVA-1.5-7B (Upper Bound, All 576 Visual Tokens)
Vanilla
8.5
100%
76.5
61.9
33.7
1510.7
64.1
85.9
31.1
58.2
21.5
100%
Approximately 55 % TFLOPs
FastV ( K=2,R=50% )
4.9
58%
73.5
60.2
32.4
1475.6
64.3
84.0
29.8
57.2
17.3
95.54%
VTW ( K=16 )
4.7
55%
66.3
55.1
32.8
1497.0
64.0
82.8
19.2
55.3
16.2
88.93%
ShortV ( l=19 )
4.7
55%
75.7
60.9
33.3
1503.1
64.8
86.2
27.9
55.1
17.9
96.08%
TABLE I : Comparison of training-free MLLM efficiency methods. FLOPs Ratio denotes the proportion of FLOPs retained relative to the vanilla model. Best results are in bold .
Method
GQA
TextVQA
POPE
Acc. (%)
Upper Bound, All 2880 Tokens (100%)
LLaVA-NeXT-7B
62.9
59.6
86.3
100.0%
Retain 320 Tokens ( ↓ 88.9%)
FastV
55.9
55.7
71.7
88.47%
SparseVLM
56.5
52.4
73.5
87.64%
VisionZip
58.1
57.6
75.0
91.97%
TABLE II : Comparisons of our MSV-Prune, with other SOTA training-free token pruning methods. Best results are in bold .
Method
R
MMStar
TextVQA
MME
Acc. (%)
Upper Bound, All 576 Tokens (100%)
LLaVA-1.5-7B
100%
33.7
58.2
1510.7
100%
ShortV
81%
33.8
57.3
1503.3
99.42%
Skip All Attn
82%
34.1
51.1
1300.6
91.69%
Skip All FFN
38%
28.9
40.9
875.9
71.34%
Skip Partial FFN
81%
32.6
57.4
1483.1
97.84%
TABLE III : Comparisons of our ASL-Skip with other training-free layer skipping/pruning methods. No visual token is pruned. Most of these methods are evaluated at around 80 % of the original TFLOPs for fair comparisons. Best results are in bold . “R” denotes TFLOPs ratio.
Fig. 5 : Instance-level visualization of boundary shift.
Fig. 6 : KL divergence comparison with magnitude pruning across LLM layers.
Fig. 7 : Visualization of retained tokens by MSV-Prune and unskipped tokens by ASL-Skip; anchors in purple ( r=0.7 ), complementary tokens in yellow.
Fig. 8 : SPIDER efficiency on LLaVA-1.5-7B at varying token reduction ratios; key image regions relevant to queries marked by red dashed boxes.
Method
MMB
MMB CN
POPE
SQA IMG
VizWiz
Acc. (%)
Vanilla, 100% Tokens
Qwen2.5-VL-3B-Instruct
77.3
73.0
87.0
80.4
68.3
100 %
Approximately 35 % TFLOPs
FastV
74.4
70.6
85.0
79.3
66.9
97.4%
VisionZip
74.9
69.8
85.4
80.1
67.1
97.7%
HiPrune
75.8
71.3
86.0
80.0
67.5
98.6%
TABLE IV : Results on Qwen2.5-VL-3B-Instruct. Best results are in bold .
Method
TextVQA
ChartQA
DocVQA
GQA
Acc. (%)
Vanilla, 100% Tokens
Qwen3-VL-8B-Instruct
82.98
83.16
95.75
61.88
100%
Approximately 55% TFLOPs
PruneSID
73.46
63.56
90.51
60.08
89.15%
IVC-Prune
82.35
79.60
95.31
61.33
98.40%
iLLaVA
76.11
64.68
84.89
61.15
89.25%
TABLE V : Results on Qwen3-VL-8B-Instruct. Best results are in bold .
Method
POPE
TextVQA
MMB
Acc. (%)
Upper Bound, All 576 Tokens (100%)
LLaVA-1.5-7B
85.9
58.2
64.1
100%
Approximately 30 % TFLOPs
+ PruMerge+ (Train)
84.0
57.1
64.9
99.0%
+ SPIDER (Train-Free)
84.8
57.4
63.9
99.0%
+ SPIDER (Train)
85.0
57.6
64.4
99.5%
TABLE VI : Comparisons of training-aware modes. The TFLOPs are kept at a similar level around 30 % for fair comparison, where 1/4 of visual tokens are preserved for PruMerge+. Best results are in bold .
Method
GQA
TextVQA
POPE
Acc. (%)
Upper Bound, All 2880 Tokens (100%)
LLaVA-NeXT-7B
62.9
59.6
86.3
100.0%
Retain 1920 Tokens
(a) Semantic Cluster
w/o Semantic Cluster
61.3
59.3
87.2
99.33%
Cluster Num =2
61.9
59.4
87.9
99.98%
TABLE VII : Ablation of our token pruning method, MSV-Prune, on LLaVA-NeXT-7B. Best results are in bold . ‘*’ denotes our default setting.
Method
MMStar
TextVQA
MME
Acc. (%)
Upper Bound, All 576 Tokens (100%)
LLaVA-1.5-7B
33.7
58.2
1510.7
100%
(a) Components of Sfuse
w/o Ssa
33.5
57.6
1497.8
99.17%
w/o SLC
32.7
58.0
1483.6
98.30%
All Equipped*
33.7
58.1
1504.9
99.81%
TABLE VIII : Ablation of our sub-layer skipping method, ASL-Skip. No visual token is pruned. Best results are in bold .
Anchor Ratio
GQA
TextVQA
POPE
Acc. (%)
LLaVA-1.5-7B
61.9
58.2
85.9
100.0%
0
59.3
54.7
83.7
95.74%
0.3
60.2
57.0
86.2
98.51%
0.5
60.6
57.5
86.4
99.09%
0.7*
60.9
57.9
86.6
99.56%
1.0
60.8
56.9
86.1
98.74%
TABLE IX : Ablation of anchor token ratio in our token pruning method, MSV-Prune. Best results are in bold . ’*’ denotes our default setting.
Fig. 9 : Heatmap of token skipping proportion (a) and POPE accuracy (b) for different Tskip and w1/w2 settings on LLaVA-1.5-7B. No visual token is pruned.
Fig. 10 : Distribution of High-attention Tokens Across Vision Encoder Layers.
TABLE X : Replaced layers for different MLLM series and parameter scales.
Fig. 11 : Sensitivity of MSV-Prune to the choice of intermediate vision encoder layer on (a) LLaVA-1.5-7B (25–30 % TFLOPs) and (b) Qwen3-VL-8B-Instruct ( ∼ 55% TFLOPs). Performance remains stable across a wide range of layer choices on both architectures, confirming that the midpoint layer ⌊L/2⌋ is a robust and architecture-agnostic default.
Fig. 12 : Failure cases: red boxes highlight relevant areas; purple masks show skipped patch tokens during decoding.
Model
Retention
Attn
Attn
FFN
Dense-SLC
Pruned-SLC
Δ
Top-8
Top-16
Top-5
Acc. (%)
Acc. (%)
(%)
LLaVA-1.5-7B
385/576
7/8
14/16
5/5
99.61
99.62
+0.01
LLaVA-NeXT-7B
1920/2880
7/8
13/16
4/5
99.84
99.82
-0.02
TABLE XI : Stability of the offline SLC policy under the pruned token distribution. We compare the layer-wise SLC ranking computed on dense visual tokens with that recomputed after MSV-Prune. We further report the average accuracy on GQA, MMVet, and POPE when ASL-Skip uses either the original dense-token SLC policy or the recomputed pruned-token SLC policy.
School of Software Technology, Zhejiang University, Ningbo, China · Zhejiang Key Laboratory of Digital-Intelligence Service Technology, China · The University of Queensland, St Lucia, QLD, Australia +2
MoE Key Lab of BIPC, NEL-BITA, University of Science and Technology of China, Hefei, China · ChangXin Memory Technologies, Hefei, China · APKL of BIIP, IAI, Hefei Comprehensive National Science Center, Hefei, China