cs.CVMay 24, 2025

ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance

Authors: Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao, Shijian Lu

Organizations: CCDS, NTU, Singapore · CCST, ZJUT, China · Terminus AI Lab, UCAS, China

Abstract

Visual token pruning, which aims to compress and prune redundant visual tokens, plays a critical role in efficient inference with large vision-language models (LVLMs). However, existing methods fail to disentangle intra-modal visual redundancy from cross-modal redundancy between vision and language. We show that visual token diversity and task-specific token relevance are two crucial yet orthogonal factors that complement each other in conveying useful information and should therefore be treated separately for more effective visual token pruning. Building upon this insight, we design TODRE, a two-stage and training-free framework that incorporates Token Diversity and task RElevance for effective token compression and efficient LVLM inference. Instead of pruning redundant tokens, we introduce a greedy max-sum diversification algorithm that selects and retains a subset of diverse and representative visual tokens after the vision encoder. On top of that, ToDRE leverages an ``information migration'' mechanism to eliminate task-irrelevant visual tokens within certain decoder layers of the large language model (LLM), further improving token pruning and LVLM inference. Extensive experiments show that ToDRE prunes 90% of visual tokens after the vision encoder as well as all visual tokens in certain LLM decoder layers, leading to a 2.6x speed-up in total inference time while maintaining 95.0% model performance plus excellent model compatibility. The code is available at: this https URL.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VPRune: Efficient Training-free Pre-LLM Visual Token Pruning

    Sep 21, 2026Guangchuan Lv, Dianxing Shi, Dingjie FuVisual Token PruningRecent Vision-Language Models

  2. ACPruner: Visual Token Pruning as Biased Attention Coverage Maximization in LVLMs

    Sep 28, 2026Xu Li, Yuxuan Liang, Yi Zheng +6Visual Token PruningSaliency