cs.CVSep 28, 2026

SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models

Authors: Tianxiang Chen, Zhentao Tan, Zi Ye, Yue Wu, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Tao Gong, +4 more

Organizations: End-User Intelligent Computing BU, Alibaba Cloud · Alibaba Group · Department of Computer Science, Maynooth University · School of Cyber Science and Technology, University of Science and Technology of China · Fudan University

Abstract

Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using blockwise importance, the finer-grained inter-layer representation shifts and the distribution differences within the layers themselves have not been fully explored. In this work, we comprehensively investigate this dual-level inefficiency. We posit that intermediate layer tokens from vision encoders should be considered for effective visual token pruning, as semantic focus shifts across layers, with middle-layer tokens capturing more detailed object-centric information that deeper layers may abstract away. Furthermore, we reveal the differential contributions of Attention and FFNs across distinct LLM decoder layers. Building upon these discoveries, we propose \textbf{SPIDER}, a training-free framework that integrates multi-layer \underline{\textbf{S}}emantic visual token \underline{\textbf{P}}run\underline{\textbf{I}}ng with an a\underline{\textbf{D}}aptive sub-lay\underline{\textbf{ER}} skipping mechanism. Experimental evaluations demonstrate that SPIDER consistently maintains strong performance across various MLLM architectures and reduction ratios. For instance, on LLaVA-NeXT-7B, SPIDER reduces FLOPs by 79%79\% while maintaining 96%\% of the baseline performance.

Figures & tables

Explore similar work

CardsList
  1. Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

    Jun 23, 2026Bin Chen, Yuxiang Cai, Yadan Luo +3Visual Token PruningCross-Modal Attention

  2. S2^2Prune: Spatially Structured Visual Token Pruning for Multimodal Large Language Models

    Sep 1, 2026Yuanyuan Jia, Shunpu Tang, Qianqian YangVisual Token PruningMultimodal Large Language Models

  3. SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models

    Jul 28, 2026Yuchen Wang, Qihui Zhu, Yang Liu +2Visual Token PruningLong Visual-Token Sequences