cs.CVOct 6, 2026

DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

Authors: Shuo Yang, Changbai Li, Linlin Yang, Huobin Tan, Rongyu Chen, Tongfei Chen, Tian Wang, Sheng Xu, +1 more

Organizations: Beihang University · Communication University of China · National University of Singapore

Abstract

Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models

    Jul 30, 2026Jie Ma, Zhike Qiu, Jie Gao +4Visual Token PruningMultimodal Large Language Models

  2. S2^2Prune: Spatially Structured Visual Token Pruning for Multimodal Large Language Models

    Sep 1, 2026Yuanyuan Jia, Shunpu Tang, Qianqian YangVisual Token PruningMultimodal Large Language Models

  3. SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models

    Jul 28, 2026Yuchen Wang, Qihui Zhu, Yang Liu +2Visual Token PruningLong Visual-Token Sequences