cs.CVSep 28, 2026

When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model

Authors: Minchan Kang, Kyeonghye Park, Seoyoung Cho, Daeshik Kim, Yucheol Cho

Organizations: Korea Advanced Institute of Science and Technology (KAIST) · Hanbat National University

Abstract

Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning

    Sep 29, 2026Jing Wang, Zhiping Wu, Dongdong Ren +3Visual Token PruningRecent Vision-Language Models

  2. LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

    Apr 27, 2026Rinyoichi Takezoe, Yaqian Li, Zihao Bo +3Visual Token Pruning

  3. DiffPrune: differentiable information throttling for token pruning in vision-language models

    Aug 3, 2026Landi He, Mingde Yao, Shawn Young +1Visual Token PruningDifferentiable Physics