Adaptive Visual Token Reduction for Accelerated Image Understanding
Authors: Seyoung Jeong, Jong Pil Yun, Sang Jun Lee
Organizations: Jeonbuk National University, Jeonju, Republic of Korea · Korea Institute of Industrial Technology (KITECH), Incheon, Republic of Korea · Chung-Ang University, Seoul, Republic of Korea
Large Vision-Language Models achieve strong VQA performance, but processing high-resolution, information-rich images requires substantial computation, motivating visual token reduction. However, existing methods often prune individual tokens or rely on fixed-size cropping, limiting their ability to preserve spatially structured information such as horizontally or vertically elongated text. To address this limitation, we propose ReFIT, an instruction-guided visual token reduction framework for efficient LVLM inference. ReFIT consists of Relevance-Guided Window Reshaping (RWR) and Instruction-Guided Token Refinement (ITR), where RWR captures instruction-relevant regions by adapting to their spatial characteristics, while ITR further removes unnecessary visual tokens. Experiments on four VQA benchmarks demonstrate that ReFIT improves answer accuracy while reducing computational cost, and qualitative results demonstrate its effectiveness in localizing relevant regions and removing unnecessary visual information.
Figures & tables
Figure 1: Overall architecture of ReFIT. Given an input image and instruction, RWR adaptively reshapes the window configuration to localize instruction-relevant regions, which are then cropped and re-encoded for fine-grained visual representation. ITR further removes unnecessary visual tokens within the cropped regions using an instruction-guided relevance map, and the remaining tokens are forwarded to the LLM for answer generation.
InfoVQA
SPDocVQA
MPDocVQA
GQA
LLM
Method
ANLS↑
FLOPs (T)↓
ANLS↑
FLOPs (T) ↓
ANLS↑
FLOPs (T)↓
Acc↑
FLOPs (T)↓
Vanilla [ 10 ]
0.2552
38.98
0.6628
51.68
0.3758
50.01
0.7598
30.39
Random sampling
0.2387
24.62
0.4888
27.70
0.2988
26.66
0.7484
19.26
ToMe [ 3 ]
0.1975
26.67
0.3215
39.41
0.2135
38.32
0.7293
19.10
FastV [ 4 ]
0.2306
26.22
0.6099
28.10
0.3523
29.31
0.7478
19.23
Pdrop [ 18 ]
0.2335
26.00
0.5507
26.93
0.3637
30.82
0.7436
20.33
Table 1: Quantitative comparison with state-of-the-art methods on VQA benchmarks: InfoVQA, SPDocVQA, MPDocVQA, and GQA
Figure 2: Qualitative comparison of instruction-relevant region cropping and answer generation between existing and proposed methods. Red boxes indicate the ground-truth regions, while green boxes indicate the regions cropped by each model. The proposed method better localizes horizontally extended text and further removes unnecessary visual tokens within the cropped regions.
RWR
ITR
InfoVQA
SPDocVQA
MPDocVQA
GQA
ANLS ↑
ANLS ↑
ANLS ↑
Acc ↑
0.3024
0.6472
0.3866
0.7608
✓
0.3058
0.6529
0.3948
0.7642
✓
0.3077
0.6552
0.3907
0.7612
✓
✓
0.3189
0.6638
0.4125
0.7669
Table 2: Ablation study on the effectiveness of RWR and ITR.
Region selection
InfoVQA
SPDocVQA
MPDocVQA
GQA
ANLS ↑
Acc ↑
ANLS ↑
Acc ↑
ANLS ↑
Acc ↑
Acc ↑
Single best window
0.3077
0.2535
0.6514
0.5431
0.3994
0.3019
0.7532
Random multi-window
0.3179
0.2667
0.6643
0.5524
0.4077
0.3071
0.7526
Fixed Top- K IoU
0.3186
0.2596
0.6552
0.5487
0.4060
0.3044
0.7569
RWR (Ours)
0.3189
0.2645
0.6638
0.5562
0.4125
0.3112
0.7669
Table 3: Ablation study on region selection strategies.
Token selection
InfoVQA
SPDocVQA
MPDocVQA
GQA
ANLS ↑
Acc ↑
ANLS ↑
Acc ↑
ANLS ↑
Acc ↑
Acc ↑
Random
0.3180
0.2638
0.6634
0.5554
0.4119
0.3104
0.7567
Spatially Connected
0.3185
0.2642
0.6638
0.5562
0.4120
0.3106
0.7568
ITR (Ours)
0.3189
0.2645
0.6638
0.5562
0.4125
0.3112
0.7669
Table 4: Ablation study on token selection methods.
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning methods follow an absolute-ranking paradigm, assigning importance scores to visual tokens and retaining a fixed top-K subset. In this work, we argue that this paradigm is fundamentally brittle: attention sinks distort token importance rankings, while image redundancy and query-dependent visual evidence make fixed token budgets unreliable across inputs. We propose OccamToken, a training-free framework that replaces absolute token ranking with register-anchored relative evidence testing. Instead of asking which tokens are globally important, OccamToken evaluates whether a visual token provides information beyond a register-based reference. Our key insight is that register tokens naturally absorb low-information attention patterns, making them a stable reference for identifying genuinely informative visual evidence. Based on this principle, OccamToken performs both image-adaptive redundancy pruning and query-adaptive relevance pruning through dynamic thresholds derived from register attention. Across LLaVA-NeXT, LLaVA-v1.5, and Qwen3-VL, OccamToken consistently improves the accuracy-efficiency trade-off without additional training. Notably, on LLaVA-NeXT, it reduces 2,880 visual tokens to approximately 40 while preserving over 93% of the original accuracy, enabling stable visual token compression even in the extreme 1.4% retention regime.
Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed to the language model. Existing token reduction methods rely on attention-based scores or pairwise similarity, without an explicit semantic representation of each token. We introduce TORINO (TOken Reduction via Interpretable coNcept Overlap), a plug-and-play framework for adaptive visual token reduction in VLMs that requires no fine-tuning of the underlying model. TORINO leverages Sparse Autoencoders (SAEs) to project visual tokens into an interpretable latent space where token relationships can be analyzed through shared concept activations. Specifically, we define concept overlap as the degree of agreement between active SAE latents and use it to group tokens that share semantic content. Reduction within each group is then performed by either pruning or merging, providing a unified framework that preserves semantically important visual information while removing redundancy. Unlike fixed-budget approaches, TORINO dynamically adapts the reduction rate to input complexity, allowing different images to retain different numbers of tokens. Experiments across multiple vision-language benchmarks show that TORINO achieves favorable efficiency-accuracy trade-offs, reducing the number of visual tokens with minimal performance loss.
Recent advancements in Vision-Language Models (VLMs) enable large language models (LLMs) to process high-resolution images, significantly improving real-world multimodal understanding. However, this capability introduces a large number of vision tokens, resulting in substantial computational overhead. To mitigate this issue, various vision token pruning methods have been proposed. Nevertheless, existing approaches predominantly rely on learned semantic features within the model to capture visual redundancy. Moreover, they lack adaptive mechanisms to adjust pruning strategies according to the complexity of the input image. In this paper, we propose ERASE, a two-stage vision token pruning framework that identifies and retains salient tokens through pruning strategies adaptive to image complexity. Experiment results demonstrate that ERASE significantly reduces vision tokens while preserving accuracy. For Qwen2.5-VL-7B, at a token pruning ratio of 85%, ERASE retains 89.46% of the original model accuracy, whereas the best prior method retains only 78.1%. Our code is available at https://github.com/Tuna-Luna/ERASE.
Yuna Lee, Kyoungho Min, Yulhwa Kim
Department of Electrical and Computer Engineering, Sungkyunkwan University, Republic of Korea · Department of Semiconductor Systems Engineering, Sungkyunkwan University, Republic of Korea