Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as reasoning unfolds, making it difficult for a fixed compressed context to retain all the details needed across stages. To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve. Deformable Aggregation of Region-wise Tokens (DART) learns content-adaptive groups and aggregation capacities, constructing compact Coarse representations linked to recoverable original Fine tokens. Temporal Routing for Adaptive Contextual Evidence (TRACE) integrates decoding history to anticipate upcoming evidence needs and select, retain, or replace active Fine-token groups. Selected Fine tokens augment the persistent Coarse context in the frozen backbone, enabling stage-specific evidence access without continuously attending to all visual tokens. On Qwen3-VL-4B, ViMoD outperforms all evaluated baselines on all eight reasoning benchmarks at a 20% target visual token budget, improving the mean normalized score by 39.0% over the strongest evaluated one-shot baseline. These gains are achieved with only 0.0546% additional trainable parameters relative to the frozen backbone.
Figures & tables
Figure 1: Visual contribution changes across different reasoning stages. The model first forms an overview of the triangle diagram, determines the scale factor from 5 and 12, and reads the known base and hypotenuse, 35 and 37. Heatmaps show gradient-based visual token contribution at selected decoding stages; warmer colors indicate higher contribution relative to other tokens in the same stage. The right panel combines the corresponding regions into U=E1∪E2∪E3∪E4 . These regions occupy 11.8% of the image on average per stage, whereas their union occupies 36.5%, counting overlapping areas once.
Figure 2: Visual demand and evidence access. (a) On MathVerse at 50% token retention, the fraction of baseline-correct examples remaining correct after compression, grouped by visual-demand variation; error bars show SEM. (b) On MMMU-Pro, the same Qwen3.8 -annotated regions are accessed stage by stage (Dynamic access) or as their union (Fixed oracle), alongside VisionZip and FastV. Accuracy is normalized to Full; bands show standard deviation.
Figure 3: From per-token savings to response-level cost. FastV on MathVerse; horizontal axes show percentage changes relative to the uncompressed baseline. (a) Early next-token entropy reflects predictive uncertainty. (b) Neighboring-window embedding similarity measures the similarity of successive content. (c) Output length counts generated tokens. (d) Estimated decoder FLOPs compare amortized cost per generated token (circles) with total response cost (diamonds). At several retention ratios, longer responses outweigh per-token savings, increasing total decoder computation.
Figure 4: Overview of ViMoD. DART constructs persistent Coarse context and a C2F mapping to the original Fine-token groups. TRACE proposes evidence groups and retains or replaces the active set. Joint-KV exposes the selected Fine evidence alongside Coarse context and text history.
Figure 5: DART adapts compression granularity and group shape to image content. (a) Blank regions generally undergo stronger compression, while text-bearing regions retain finer granularity. Lower allocation-map values indicate fewer Fine tokens per Coarse token. (b) Deformable grouping follows the diagram’s geometric structure rather than fixed grid boundaries: C0 pools blank cells, C2 covers much of the hypotenuse and part of the base, and C1 captures the middle portion of the vertical leg. The cell containing the right-angle marker forms a separate group, C3 . Each group defines a Coarse token and its C2F entry; parentheses indicate the number of Fine members.
Mathematics
Logic
Multidisciplinary
Method
We-Math ‡
DynaMath ‡
MathVerse ‡
MathVista
MathVision
LogicVista
VisualPuzzles ‡
MMMU-Pro ‡
Avg. (%)
Full (100%)
45.62 (100.0%)
61.88 (100.0%)
58.60 (100.0%)
72.80 (100.0%)
47.40 (100.0%)
37.36 (100.0%)
19.18 (100.0%)
30.92 (100.0%)
100.00
Retain 20% Visual Tokens
FastV (ECCV24)
16.95 (37.2%)
32.08 (51.8%)
37.31 (63.7%)
43.40 (59.6%)
35.56 (75.0%)
22.82 (61.1%)
11.47 (59.8%)
10.46 (33.8%)
55.25
VisionZip (CVPR25)
17.43 (38.2%)
35.33 (57.1%)
37.79 (64.5%)
46.30 (63.6%)
33.91 (71.5%)
19.69 (52.7%)
8.99 (46.9%)
9.88 (32.0%)
53.31
Table 1: Mathematical, logical, and multidisciplinary reasoning. Parenthesized percentages are relative to Full; Avg. equally weights eight normalized scores. Bold and underlining mark the best and second-best scores within each target budget; ‡ denotes explicit CoT prompting. ViMoD budgets specify average Coarse-plus-Fine occupancy. † denotes our adaptation of DSTP’s attention-change access rule to the same DART memory as ViMoD.
Scene Understanding
General
Documents and Charts
Method
GQA
POPE
V ∗
MME-P
MME-C
MMBench
DocVQA
OCR EN
OCR CN
ChartQA
Avg. (%)
Full (100%)
62.11 (100.0%)
89.89 (100.0%)
68.59 (100.0%)
1693.06
610.71
83.51 (100.0%)
91.68 (100.0%)
44.89 (100.0%)
41.40 (100.0%)
80.64 (100.0%)
100.00
Retain 20% Visual Tokens
FastV (ECCV24)
50.00 (80.5%)
77.24 (85.9%)
56.54 (82.4%)
1346.63
312.50
74.14 (88.8%)
36.14 (39.4%)
29.74 (66.3%)
19.87 (48.0%)
30.84 (38.2%)
66.03
VisionZip (CVPR25)
56.41 (90.8%)
84.37 (93.9%)
59.16 (86.3%)
1423.19
373.57
74.91 (89.7%)
47.99 (52.3%)
28.12 (62.6%)
18.51 (44.7%)
48.44 (60.1%)
72.56
Table 2: General visual understanding and document–chart QA. Avg. equally weights ten Full-normalized metric columns. OCR EN/CN denote OCRBench v2 English/Chinese; MME-P/C report official perception/cognition scores. The DSTP adaptation ( † ), parenthesized percentages, and highlighting follow Tab. 1 .
Table 8
Figure 6: Accuracy–efficiency trade-offs on MMMU-Pro. Higher decoding throughput need not reduce total time. Lines connect eight target budgets (20–90%) in increasing order; stars denote uncompressed Full, with its total time marked by the vertical guide in (a). Total time covers the complete run; decode throughput is generated tokens divided by decoding time; TTFT is time to first token. Visual occupancy averages visual-token usage over decoding forward passes relative to Full, counting Coarse and active Fine tokens for ViMoD.
DART
TRACE
0.078%
0.616%
Table 5: Network runtime (% of total), averaged over eight budgets.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Training-data composition by task type. The two DART compressors share one dataset, as do the two TRACE SFT stages; the RL pool is a subset of TRACE SFT and is not counted separately.
Setting
DART
TRACE SFT: Set
TRACE SFT: Joint
TRACE RL
Dataset size
57,420
80,000
80,000
4,800
Training duration
1,200 steps
3 epochs
3 epochs
50 steps
Global batch size
192
64
64
96
Learning rate
10−5
10−4
10−4
3×10−5
Updated components
DART
TRACE except gate
Layer fusion, adapter, SSM, gate
Entire TRACE
Appendix
Table 6: Training schedule and updated components. Dataset sizes count DART response trajectories, SFT examples, and RL questions. The two SFT stages share one 80,000-example set. DART is trained independently at each ratio; all TRACE stages use 9:1 compression.
Setting
Value
Input hidden dimension
2,560
Backbone hidden-state indices
27,31,35 (backbone indexing)
SSM state dimension
256
Region key/query dimension
128
Initial timescales
1, 2, 4, 8 routing steps
Routing interval
16 generated tokens
Appendix
Table 7: TRACE architecture and routing configuration.
Setting
Offline SFT
Online RL
Inputs
Image and question
Image, question, and student prefix
Prediction target
Reasoning steps, evidence boxes for each upcoming step, and final answer
Evidence boxes for approximately the next 16 student tokens
Output format
XML with JSON evidence fields
Schema-constrained JSON
Appendix
Table 8: Offline and online evidence annotation. Both use Qwen3.8-27B ( temperature=0 , optional reasoning disabled) and image-relative box coordinates in [0,1000] . Complete templates appear in Appendix E .
With the rapid advancement of large multimodal models (LMMs), inference-time overhead has become a key bottleneck for real-world deployment. Existing methods typically prune visual tokens at prefill, assuming the required visual evidence remains static during reasoning. However, we empirically show that visual evidence is strongly step-dependent: only a sparse subset of visual tokens is critical at each decoding step, and the critical set evolves across reasoning. Furthermore, we identify a coupled bottleneck where redundant visual context can steer the model toward query-irrelevant regions, lengthening the reasoning trace. Guided by these insights, we propose VisionPulse, a step-wise visual token pruning framework during reasoning. VisionPulse computes a lightweight visual attention mass to estimate the step-wise retention budget by exploiting its strong positive correlation with LMMs' effective visual token usage and retain only the most critical tokens under this budget. By enforcing visual sparsity during reasoning, VisionPulse filters redundant visual context while preserving relevant visual evidence, shortening reasoning traces naturally. Extensive experiments show that VisionPulse only retains 5% of visual tokens per step with reasoning traces shortened by 11.2%, while keeping accuracy almost unchanged.
Hengbo Xu, Shengjie Jin, Yanbiao Ma +1
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China.
Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose δ-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, δ-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.
Jingdi lei, Junxian Li, Di Zhang +3
Nanyang Technological University · Shanghai Jiao Tong University · Fudan University +2
Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these representations can introduce errors and increasingly interfere with textual reasoning as reasoning progresses. We propose Visual-to-Text Chain-of-Thought Distillation (V2T), a framework that enables LVLMs to internalize visual reasoning without generating intermediate visual representations at inference time. V2T first trains a teacher LVLM using interleaved visual and textual chains of thought, and then uses knowledge distillation to train a student LVLM using the teacher's logits and cross-entropy supervision from ground-truth textual reasoning. When reasoning images can be mapped to the original image, V2T can additionally distill the teacher's attention to corresponding regions, while ground-truth bounding boxes can further guide a subsequent reinforcement learning stage. Experiments across multiple multimodal reasoning benchmarks show that V2T consistently outperforms the teacher and existing baselines, improving average accuracy by 14.3% on a held-out set and 2.7% on the broader visual evaluation suite. Moreover, lightweight SFT and substantially reduced RL make V2T up to 42x faster to train than state-of-the-art baselines.
Wenhan Yang, Nilay Naharas, Ali Payani +1
University of California, Los Angeles · Cisco Research