cs.CVAug 8, 2025

AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference via Dynamical Text Guidance

Authors: Weichen Zhang, Zhui Zhu, Ningbo Li, Shilong Tao, Hongzi Zhu, Jingao Xu, Kebin Liu, Yunhao Liu

Organizations: Global Innovation Exchange, Tsinghua University, Beijing, China · Department of Automation, Tsinghua University, Beijing, China · School of Computer Science, Peking University, Beijing, China

Abstract

Vision-language models (VLMs) have achieved impressive performance on multimodal inference tasks, but the cost remains a significant challenge due to the large number of vision tokens processed during the prefill stage. Existing token pruning methods often rely on utilizing the static attention patterns directly, failing to exploit the dynamic internal signals within VLMs. To address the issue, we propose AdaptInfer, a novel plug-and-play framework for vision token pruning. First, we introduce a dynamic text-guided pruning mechanism that construct soft priors over text-token importance on each pruning layer, allowing more informed scoring of vision tokens at each stage. Second, we observe a highly consistent distribution of cross-modal attention shifts by architecture, which inspires us to introduce a efficient data-driven schedule which determines the pruning locations automatically. Experimental results have verified the effectiveness and generalization of the proposed method. Under the same token budget, AdaptInfer surpasses existing approaches in accuracy. For instance, AdaptInfer maintains averagely 99.4% accuracy on Qwen2-VL while 70% of the prefilling vision token overhead is reduced. The source code is available on: https://github.com/weiczh02/AdaptInfer-base.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

    Apr 27, 2026Rinyoichi Takezoe, Yaqian Li, Zihao Bo +3Visual Token Pruning

  2. DiffPrune: differentiable information throttling for token pruning in vision-language models

    Aug 3, 2026Landi He, Mingde Yao, Shawn Young +1Visual Token PruningDifferentiable Physics