SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models
Authors: Tianxiang Chen, Zhentao Tan, Zi Ye, Yue Wu, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Tao Gong, +4 more
Organizations: End-User Intelligent Computing BU, Alibaba Cloud · Alibaba Group · Department of Computer Science, Maynooth University · School of Cyber Science and Technology, University of Science and Technology of China · Fudan University
Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using blockwise importance, the finer-grained inter-layer representation shifts and the distribution differences within the layers themselves have not been fully explored. In this work, we comprehensively investigate this dual-level inefficiency. We posit that intermediate layer tokens from vision encoders should be considered for effective visual token pruning, as semantic focus shifts across layers, with middle-layer tokens capturing more detailed object-centric information that deeper layers may abstract away. Furthermore, we reveal the differential contributions of Attention and FFNs across distinct LLM decoder layers. Building upon these discoveries, we propose \textbf{SPIDER}, a training-free framework that integrates multi-layer \underline{\textbf{S}}emantic visual token \underline{\textbf{P}}run\underline{\textbf{I}}ng with an a\underline{\textbf{D}}aptive sub-lay\underline{\textbf{ER}} skipping mechanism. Experimental evaluations demonstrate that SPIDER consistently maintains strong performance across various MLLM architectures and reduction ratios. For instance, on LLaVA-NeXT-7B, SPIDER reduces FLOPs by 79% while maintaining 96% of the baseline performance.
Figures & tables
Fig. 1 : Visualizations to show the importance of mid-layer features for fine-grained token pruning. We highlight the top 25 % most attentive tokens (in purple) from various layers of the CLIP vision encoder, together with their t-SNE embeddings. Mid-layer tokens exhibit a higher key object coverage ratio ( R∗ ), indicating a stronger focus on crucial object-centric details compared to deeper layers, which tend to capture broader global semantics. Averaged over 100 sampled evaluation images, R∗ peaks at 29.65% at layer 12, compared to 11.28% at layer 1 and 15.83% at layer 24.
Fig. 2 : Visualization of sub-layer skipping, where we visualize either (a) skipping attention sub-layer module or (c) FFN sub-layer module. Sub-layer skipping is achieved via replacing the LLM decoder block with corresponding sparse blocks, so that visual tokens are blocked in attention or FFN modules.
Fig. 3 : The sub-layer contribution scores ( SLC ) of LLaVA-1.5-7B and LLaVA-NeXT-7B. Lower SLC values indicate a weaker influence of the corresponding attention or FFN sub-layer on the specified tokens. Skipping transformations of visual tokens in such low-impact sub-layers yields minimal divergence from the original model’s output distribution, and the SLC distribution is very different across various benchmarks.
Fig. 4 : Illustration of SPIDER . We begin by token pruning using the semantic information from the middle and last layers. The retained tokens are fed to the LLM for adaptive sub-layer skipping.
Method
TFLOPs
Ratio
VQAv2
GQA
MMStar
MME
MMB
POPE
MMVet
TextVQA
DocVQA
Acc. (%)
LLaVA-1.5-7B (Upper Bound, All 576 Visual Tokens)
Vanilla
8.5
100%
76.5
61.9
33.7
1510.7
64.1
85.9
31.1
58.2
21.5
100%
Approximately 55 % TFLOPs
FastV ( K=2,R=50% )
4.9
58%
73.5
60.2
32.4
1475.6
64.3
84.0
29.8
57.2
17.3
95.54%
VTW ( K=16 )
4.7
55%
66.3
55.1
32.8
1497.0
64.0
82.8
19.2
55.3
16.2
88.93%
ShortV ( l=19 )
4.7
55%
75.7
60.9
33.3
1503.1
64.8
86.2
27.9
55.1
17.9
96.08%
TABLE I : Comparison of training-free MLLM efficiency methods. FLOPs Ratio denotes the proportion of FLOPs retained relative to the vanilla model. Best results are in bold .
Method
GQA
TextVQA
POPE
Acc. (%)
Upper Bound, All 2880 Tokens (100%)
LLaVA-NeXT-7B
62.9
59.6
86.3
100.0%
Retain 320 Tokens ( ↓ 88.9%)
FastV
55.9
55.7
71.7
88.47%
SparseVLM
56.5
52.4
73.5
87.64%
VisionZip
58.1
57.6
75.0
91.97%
TABLE II : Comparisons of our MSV-Prune, with other SOTA training-free token pruning methods. Best results are in bold .
Method
R
MMStar
TextVQA
MME
Acc. (%)
Upper Bound, All 576 Tokens (100%)
LLaVA-1.5-7B
100%
33.7
58.2
1510.7
100%
ShortV
81%
33.8
57.3
1503.3
99.42%
Skip All Attn
82%
34.1
51.1
1300.6
91.69%
Skip All FFN
38%
28.9
40.9
875.9
71.34%
Skip Partial FFN
81%
32.6
57.4
1483.1
97.84%
TABLE III : Comparisons of our ASL-Skip with other training-free layer skipping/pruning methods. No visual token is pruned. Most of these methods are evaluated at around 80 % of the original TFLOPs for fair comparisons. Best results are in bold . “R” denotes TFLOPs ratio.
Fig. 5 : Instance-level visualization of boundary shift.
Fig. 6 : KL divergence comparison with magnitude pruning across LLM layers.
Fig. 7 : Visualization of retained tokens by MSV-Prune and unskipped tokens by ASL-Skip; anchors in purple ( r=0.7 ), complementary tokens in yellow.
Fig. 8 : SPIDER efficiency on LLaVA-1.5-7B at varying token reduction ratios; key image regions relevant to queries marked by red dashed boxes.
Method
MMB
MMB CN
POPE
SQA IMG
VizWiz
Acc. (%)
Vanilla, 100% Tokens
Qwen2.5-VL-3B-Instruct
77.3
73.0
87.0
80.4
68.3
100 %
Approximately 35 % TFLOPs
FastV
74.4
70.6
85.0
79.3
66.9
97.4%
VisionZip
74.9
69.8
85.4
80.1
67.1
97.7%
HiPrune
75.8
71.3
86.0
80.0
67.5
98.6%
TABLE IV : Results on Qwen2.5-VL-3B-Instruct. Best results are in bold .
Method
TextVQA
ChartQA
DocVQA
GQA
Acc. (%)
Vanilla, 100% Tokens
Qwen3-VL-8B-Instruct
82.98
83.16
95.75
61.88
100%
Approximately 55% TFLOPs
PruneSID
73.46
63.56
90.51
60.08
89.15%
IVC-Prune
82.35
79.60
95.31
61.33
98.40%
iLLaVA
76.11
64.68
84.89
61.15
89.25%
TABLE V : Results on Qwen3-VL-8B-Instruct. Best results are in bold .
Method
POPE
TextVQA
MMB
Acc. (%)
Upper Bound, All 576 Tokens (100%)
LLaVA-1.5-7B
85.9
58.2
64.1
100%
Approximately 30 % TFLOPs
+ PruMerge+ (Train)
84.0
57.1
64.9
99.0%
+ SPIDER (Train-Free)
84.8
57.4
63.9
99.0%
+ SPIDER (Train)
85.0
57.6
64.4
99.5%
TABLE VI : Comparisons of training-aware modes. The TFLOPs are kept at a similar level around 30 % for fair comparison, where 1/4 of visual tokens are preserved for PruMerge+. Best results are in bold .
Method
GQA
TextVQA
POPE
Acc. (%)
Upper Bound, All 2880 Tokens (100%)
LLaVA-NeXT-7B
62.9
59.6
86.3
100.0%
Retain 1920 Tokens
(a) Semantic Cluster
w/o Semantic Cluster
61.3
59.3
87.2
99.33%
Cluster Num =2
61.9
59.4
87.9
99.98%
TABLE VII : Ablation of our token pruning method, MSV-Prune, on LLaVA-NeXT-7B. Best results are in bold . ‘*’ denotes our default setting.
Method
MMStar
TextVQA
MME
Acc. (%)
Upper Bound, All 576 Tokens (100%)
LLaVA-1.5-7B
33.7
58.2
1510.7
100%
(a) Components of Sfuse
w/o Ssa
33.5
57.6
1497.8
99.17%
w/o SLC
32.7
58.0
1483.6
98.30%
All Equipped*
33.7
58.1
1504.9
99.81%
TABLE VIII : Ablation of our sub-layer skipping method, ASL-Skip. No visual token is pruned. Best results are in bold .
Anchor Ratio
GQA
TextVQA
POPE
Acc. (%)
LLaVA-1.5-7B
61.9
58.2
85.9
100.0%
0
59.3
54.7
83.7
95.74%
0.3
60.2
57.0
86.2
98.51%
0.5
60.6
57.5
86.4
99.09%
0.7*
60.9
57.9
86.6
99.56%
1.0
60.8
56.9
86.1
98.74%
TABLE IX : Ablation of anchor token ratio in our token pruning method, MSV-Prune. Best results are in bold . ’*’ denotes our default setting.
Fig. 9 : Heatmap of token skipping proportion (a) and POPE accuracy (b) for different Tskip and w1/w2 settings on LLaVA-1.5-7B. No visual token is pruned.
Fig. 10 : Distribution of High-attention Tokens Across Vision Encoder Layers.
TABLE X : Replaced layers for different MLLM series and parameter scales.
Fig. 11 : Sensitivity of MSV-Prune to the choice of intermediate vision encoder layer on (a) LLaVA-1.5-7B (25–30 % TFLOPs) and (b) Qwen3-VL-8B-Instruct ( ∼ 55% TFLOPs). Performance remains stable across a wide range of layer choices on both architectures, confirming that the midpoint layer ⌊L/2⌋ is a robust and architecture-agnostic default.
Fig. 12 : Failure cases: red boxes highlight relevant areas; purple masks show skipped patch tokens during decoding.
Model
Retention
Attn
Attn
FFN
Dense-SLC
Pruned-SLC
Δ
Top-8
Top-16
Top-5
Acc. (%)
Acc. (%)
(%)
LLaVA-1.5-7B
385/576
7/8
14/16
5/5
99.61
99.62
+0.01
LLaVA-NeXT-7B
1920/2880
7/8
13/16
4/5
99.84
99.82
-0.02
TABLE XI : Stability of the offline SLC policy under the pruned token distribution. We compare the layer-wise SLC ranking computed on dense visual tokens with that recomputed after MSV-Prune. We further report the average accuracy on GQA, MMVet, and POPE when ASL-Skip uses either the original dense-token SLC policy or the recomputed pruned-token SLC policy.
Reducing visual token redundancy is critical for accelerating Multimodal Large Language Models (MLLMs) without degrading cross-modal reasoning performance. Existing token pruning methods typically rely on single-layer signals, such as attention scores or token similarities, which overlook the cross-layer transformation of visual representations and may exhibit positional bias in multimodal token sequences. To address this limitation, we propose a training-free token pruning framework based on Cross-Layer Spectral Evolution (CLSE). Instead of measuring token importance from single-layer feature magnitudes, CLSE quantifies how token representations evolve across Transformer layers in the frequency domain. This evolution reflects the transition from high-frequency structural details to low-frequency semantic abstractions. We observe that tokens with stronger spectral redistribution across layers are more likely to be semantically active and should therefore be preserved. By modeling cross-layer token dynamics, CLSE provides a stable importance criterion that mitigates positional bias. Extensive experiments on both image and video benchmarks demonstrate that CLSE achieves a superior trade-off between efficiency and accuracy under aggressive token reduction. Across multiple MLLMs, CLSE reduces FLOPs, KV cache memory, and latency while maintaining competitive or improved performance.
Bin Chen, Yuxiang Cai, Yadan Luo +3
School of Software Technology, Zhejiang University, Ningbo, China · Zhejiang Key Laboratory of Digital-Intelligence Service Technology, China · The University of Queensland, St Lucia, QLD, Australia +2
Visual token pruning reduces the inference overhead of multimodal large language models (MLLMs) by retaining only a subset of visual tokens. Existing methods usually select tokens based on importance or redundancy. However, we observe that these criteria produce stable spatial biases across inputs and do not always outperform simple Uniform Grid sampling, highlighting the value of broad spatial coverage. Motivated by this, we propose S2Prune, a training-free pruning method that preserves spatial coverage while adapting token density to local image structure. We first divide the image into regions and assign at least one token to each region to preserve coverage. The remaining token budget is then distributed according to Laplacian variation, giving more tokens to regions with richer structure. We then use Early Representation Change (ERC), computed from the first decoder block, to select representative tokens within each region. We evaluate S2Prune across diverse settings and two MLLM architectures. On Qwen2.5-VL-7B-Instruct, it achieves the highest average accuracy among the evaluated training-free pruning methods. With only 32 of the original 576 visual tokens, it still retains 79.3% of the full-model performance. Code is available at https://github.com/yuanyuanjia71-spec/S2Prune.
Yuanyuan Jia, Shunpu Tang, Qianqian Yang
College of Information Science and Electronic Engineering, Zhejiang University Hangzhou, China
Recent multimodal large language models (MLLMs), such as Qwen2.5-VL and InternVL3, generate large numbers of vision tokens for high-resolution inputs, leading to substantial computational cost. Existing vision token pruning methods either depend on cross-modal attention and cannot prune before the prefill stage, or rely on diversity estimation with high computational overhead. We observe that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities. Based on this observation, we propose SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informative vision tokens. SepPrune reuses the LLM's built-in projection parameters and requires no architectural changes. Experiments on Qwen2.5-VL-7B show that SepPrune achieves state-of-the-art performance, retaining 96.3% of the original accuracy while removing 80.2% of vision tokens.
Yuchen Wang, Qihui Zhu, Yang Liu +2
MoE Key Lab of BIPC, NEL-BITA, University of Science and Technology of China, Hefei, China · ChangXin Memory Technologies, Hefei, China · APKL of BIIP, IAI, Hefei Comprehensive National Science Center, Hefei, China