Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs). Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding. However, visual evidence needs can shift as reasoning unfolds, making it difficult for a fixed compressed context to retain all the details needed across stages. To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve. Deformable Aggregation of Region-wise Tokens (DART) learns content-adaptive groups and aggregation capacities, constructing compact Coarse representations linked to recoverable original Fine tokens. Temporal Routing for Adaptive Contextual Evidence (TRACE) integrates decoding history to anticipate upcoming evidence needs and select, retain, or replace active Fine-token groups. Selected Fine tokens augment the persistent Coarse context in the frozen backbone, enabling stage-specific evidence access without continuously attending to all visual tokens. On Qwen3-VL-4B, ViMoD outperforms all evaluated baselines on all eight reasoning benchmarks at a 20% target visual token budget, improving the mean normalized score by 39.0% over the strongest evaluated one-shot baseline. These gains are achieved with only 0.0546% additional trainable parameters relative to the frozen backbone.
Figures & tables
Figure 1: Visual contribution changes across different reasoning stages. The model first forms an overview of the triangle diagram, determines the scale factor from 5 and 12, and reads the known base and hypotenuse, 35 and 37. Heatmaps show gradient-based visual token contribution at selected decoding stages; warmer colors indicate higher contribution relative to other tokens in the same stage. The right panel combines the corresponding regions into U=E1∪E2∪E3∪E4 . These regions occupy 11.8% of the image on average per stage, whereas their union occupies 36.5%, counting overlapping areas once.
Figure 2: Visual demand and evidence access. (a) On MathVerse at 50% token retention, the fraction of baseline-correct examples remaining correct after compression, grouped by visual-demand variation; error bars show SEM. (b) On MMMU-Pro, the same Qwen3.8 -annotated regions are accessed stage by stage (Dynamic access) or as their union (Fixed oracle), alongside VisionZip and FastV. Accuracy is normalized to Full; bands show standard deviation.
Figure 3: From per-token savings to response-level cost. FastV on MathVerse; horizontal axes show percentage changes relative to the uncompressed baseline. (a) Early next-token entropy reflects predictive uncertainty. (b) Neighboring-window embedding similarity measures the similarity of successive content. (c) Output length counts generated tokens. (d) Estimated decoder FLOPs compare amortized cost per generated token (circles) with total response cost (diamonds). At several retention ratios, longer responses outweigh per-token savings, increasing total decoder computation.
Figure 4: Overview of ViMoD. DART constructs persistent Coarse context and a C2F mapping to the original Fine-token groups. TRACE proposes evidence groups and retains or replaces the active set. Joint-KV exposes the selected Fine evidence alongside Coarse context and text history.
Figure 5: DART adapts compression granularity and group shape to image content. (a) Blank regions generally undergo stronger compression, while text-bearing regions retain finer granularity. Lower allocation-map values indicate fewer Fine tokens per Coarse token. (b) Deformable grouping follows the diagram’s geometric structure rather than fixed grid boundaries: C0 pools blank cells, C2 covers much of the hypotenuse and part of the base, and C1 captures the middle portion of the vertical leg. The cell containing the right-angle marker forms a separate group, C3 . Each group defines a Coarse token and its C2F entry; parentheses indicate the number of Fine members.
Mathematics
Logic
Multidisciplinary
Method
We-Math ‡
DynaMath ‡
MathVerse ‡
MathVista
MathVision
LogicVista
VisualPuzzles ‡
MMMU-Pro ‡
Avg. (%)
Full (100%)
45.62 (100.0%)
61.88 (100.0%)
58.60 (100.0%)
72.80 (100.0%)
47.40 (100.0%)
37.36 (100.0%)
19.18 (100.0%)
30.92 (100.0%)
100.00
Retain 20% Visual Tokens
FastV (ECCV24)
16.95 (37.2%)
32.08 (51.8%)
37.31 (63.7%)
43.40 (59.6%)
35.56 (75.0%)
22.82 (61.1%)
11.47 (59.8%)
10.46 (33.8%)
55.25
VisionZip (CVPR25)
17.43 (38.2%)
35.33 (57.1%)
37.79 (64.5%)
46.30 (63.6%)
33.91 (71.5%)
19.69 (52.7%)
8.99 (46.9%)
9.88 (32.0%)
53.31
Table 1: Mathematical, logical, and multidisciplinary reasoning. Parenthesized percentages are relative to Full; Avg. equally weights eight normalized scores. Bold and underlining mark the best and second-best scores within each target budget; ‡ denotes explicit CoT prompting. ViMoD budgets specify average Coarse-plus-Fine occupancy. † denotes our adaptation of DSTP’s attention-change access rule to the same DART memory as ViMoD.
Scene Understanding
General
Documents and Charts
Method
GQA
POPE
V ∗
MME-P
MME-C
MMBench
DocVQA
OCR EN
OCR CN
ChartQA
Avg. (%)
Full (100%)
62.11 (100.0%)
89.89 (100.0%)
68.59 (100.0%)
1693.06
610.71
83.51 (100.0%)
91.68 (100.0%)
44.89 (100.0%)
41.40 (100.0%)
80.64 (100.0%)
100.00
Retain 20% Visual Tokens
FastV (ECCV24)
50.00 (80.5%)
77.24 (85.9%)
56.54 (82.4%)
1346.63
312.50
74.14 (88.8%)
36.14 (39.4%)
29.74 (66.3%)
19.87 (48.0%)
30.84 (38.2%)
66.03
VisionZip (CVPR25)
56.41 (90.8%)
84.37 (93.9%)
59.16 (86.3%)
1423.19
373.57
74.91 (89.7%)
47.99 (52.3%)
28.12 (62.6%)
18.51 (44.7%)
48.44 (60.1%)
72.56
Table 2: General visual understanding and document–chart QA. Avg. equally weights ten Full-normalized metric columns. OCR EN/CN denote OCRBench v2 English/Chinese; MME-P/C report official perception/cognition scores. The DSTP adaptation ( † ), parenthesized percentages, and highlighting follow Tab. 1 .
Table 8
Figure 6: Accuracy–efficiency trade-offs on MMMU-Pro. Higher decoding throughput need not reduce total time. Lines connect eight target budgets (20–90%) in increasing order; stars denote uncompressed Full, with its total time marked by the vertical guide in (a). Total time covers the complete run; decode throughput is generated tokens divided by decoding time; TTFT is time to first token. Visual occupancy averages visual-token usage over decoding forward passes relative to Full, counting Coarse and active Fine tokens for ViMoD.
DART
TRACE
0.078%
0.616%
Table 5: Network runtime (% of total), averaged over eight budgets.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Training-data composition by task type. The two DART compressors share one dataset, as do the two TRACE SFT stages; the RL pool is a subset of TRACE SFT and is not counted separately.
Setting
DART
TRACE SFT: Set
TRACE SFT: Joint
TRACE RL
Dataset size
57,420
80,000
80,000
4,800
Training duration
1,200 steps
3 epochs
3 epochs
50 steps
Global batch size
192
64
64
96
Learning rate
10−5
10−4
10−4
3×10−5
Updated components
DART
TRACE except gate
Layer fusion, adapter, SSM, gate
Entire TRACE
Appendix
Table 6: Training schedule and updated components. Dataset sizes count DART response trajectories, SFT examples, and RL questions. The two SFT stages share one 80,000-example set. DART is trained independently at each ratio; all TRACE stages use 9:1 compression.
Setting
Value
Input hidden dimension
2,560
Backbone hidden-state indices
27,31,35 (backbone indexing)
SSM state dimension
256
Region key/query dimension
128
Initial timescales
1, 2, 4, 8 routing steps
Routing interval
16 generated tokens
Appendix
Table 7: TRACE architecture and routing configuration.
Setting
Offline SFT
Online RL
Inputs
Image and question
Image, question, and student prefix
Prediction target
Reasoning steps, evidence boxes for each upcoming step, and final answer
Evidence boxes for approximately the next 16 student tokens
Output format
XML with JSON evidence fields
Schema-constrained JSON
Appendix
Table 8: Offline and online evidence annotation. Both use Qwen3.8-27B ( temperature=0 , optional reasoning disabled) and image-relative box coordinates in [0,1000] . Complete templates appear in Appendix E .