EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception
Authors: Yaoxin Niu, Zhangquan Chen, Yang Zhang, Xiang An, Zhumei Wang, Chih-Ting Liao, Hongkun Cao, Ruqi Huang
Organizations: Tsinghua University · Peng Cheng Laboratory · The Hong Kong University of Science and Technology · LMMs-Lab · Beijing Institute of Technology · University of New South Wales
Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.
Figures & tables
Figure 2: EviViT pipeline. Human interactions are distilled into a question-conditioned evidence target. PTEA predicts this distribution inside the host ViT and yields one or two non-redundant decisive regions plus one context region. These regions are re-read from the original pixels, fused with coordinate-matched global tokens, and processed by the remaining frozen ViT blocks and LLM.
Figure 3: Strong-model accuracy rankings. VisualProbe VP-All (515 questions) and ZoomBench (845 questions), sorted within each panel. Open-checkpoint bars use our local evaluation. The SeProD VP-All bar is derived from its author-reported Easy/Medium/Hard scores under the original avg@32 protocol, so it is contextual rather than a matched comparison.
Figure 4: Budget curves and paired gains. Accuracy versus tokens on Qwen3-VL-4B/8B.
TTFT (s)
E2E (s)
Budget
Path
Vis. tok. (K)
Median
P90
Median
P90
Peak (GiB)
Δ Acc. (pp)
Qwen3-VL-4B
1K
Global
1.00
0.206
0.218
0.557
0.619
8.62
–
EviViT
0.90
0.384
0.403
0.738
0.804
8.66
+6.05
4K
Global
3.83
0.559
0.643
0.921
1.038
9.48
–
EviViT
3.62
0.621
0.712
0.999
1.091
9.47
+4.76
Table 3: Whole-host-isolated efficiency. Each budget pairs Global and EviViT. Tokens and accuracy use the complete sweep; latency and peak memory use 64 isolated requests per setting.
MMBench
MMStar
Visual-CoT
Host
Base
+EviViT
Δ
Base
+EviViT
Δ
Base
+EviViT
Δ
Qwen3-VL-4B
87.36
88.43
+1.06
61.73
62.80
+1.07
77.53
75.52
-2.01
Qwen3-VL-8B
88.59
88.94
+0.35
65.07
65.27
+0.20
77.92
78.35
+0.43
Qwen3.5-4B
86.37
87.76
+1.39
63.20
64.40
+1.20
77.18
78.17
+1.00
Qwen3.5-9B
87.13
86.76
-0.37
64.33
65.20
+0.87
78.98
80.34
+1.36
Table 4: General-capability compatibility on four foundation hosts.
Supervision
VP-All
VP-Hard
V ∗
Final-box
56.12
52.83
89.01
Full-trace
57.28
56.60
90.05
Δ
+1.17
+3.77
+1.05
Table 5: Full-trace versus final-box supervision at epoch three (%).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: PTEA training and checkpoint selection. Five-fold trace-only training of the Qwen3-VL-4B Block-16 selector. Thin curves show individual folds and thick curves their mean. The held-out evidence score is shown with a one-standard-deviation band across folds; dashed lines mark the selected third epoch.
Variant
VP-E
VP-M
VP-H
V ∗
HR-4K
HR-8K
PB-1img
Zoom
Avg.
Δ
Qwen3-VL-4B
Global B16
60.28
40.30
41.51
84.29
76.25
73.38
30.77
45.80
56.57
± 0.00
+ Reread
64.54
47.01
39.62
86.91
79.50
77.12
32.05
53.25
60.00
+3.43
+ Boxes
63.83
51.49
45.28
86.39
79.25
76.12
31.60
51.60
60.70
+4.13
EviViT (full)
63.83
50.37
48.11
87.96
79.62
77.75
31.83
53.37
61.61
+5.04
Qwen3-VL-8B
Appendix
Table 6: Where the gains enter. Reread uses fixed-budget source-pixel views; Boxes adds continuous crop geometry; EviViT is the full model with input-dependent capacity allocation. Accuracy is in percent. Avg. equally weights the eight benchmarks, and Δ is relative to Global B16 (global-only, 16×10242 -pixel cap) at the same scale. Bold marks each column’s best score within a scale.
Regional locations
VP-All
V ∗
Pooled
Δ
Deployed locations
57.09
92.15
66.57
–
Random translation
47.77
82.20
57.08
-9.49
Cross-image shuffled locations
47.77
81.68
56.94
-9.63
Appendix
Table 8: Location control with native multi-image input. Only crop positions change; view count, size, and token budget are fixed. Accuracy (%); Δ versus deployed locations.
Host
VP tier
Base
+EviViT
Δ
Rescue/loss
Retain
p
Qwen3-VL-4B
Easy
60.28
63.83
+3.55
16/11
87.06%
0.44
Medium
40.30
50.37
+10.07
41/14
87.04%
3.6e-04
Hard
41.51
48.11
+6.60
17/10
77.27%
0.25
Qwen3-VL-8B
Easy
67.38
69.50
+2.13
14/11
88.42%
0.69
Medium
42.16
53.36
+11.19
45/15
86.73%
1.3e-04
Hard
44.34
51.89
+7.55
19/11
76.60%
0.2
Appendix
Table 9: Paired gains come from rescues rather than answer churn. Rescue/loss compares identical questions. Retain is the fraction of originally correct answers that remain correct; p is the two-sided exact McNemar p -value. Qwen3-VL and Qwen3.5 results come from completed answer-only paired runs using the same executor.
Model
Trainable
VCoT
Weak-5
Protect-5
VP-E
VP-M
VP-H
Zoom
Avg.
Base
Frozen
78.50
89.53
70.86
67.38
42.16
44.34
43.91
55.26
EviViT
Frozen
78.31
87.94
71.82
69.50
53.36
51.89
50.65
60.74
Base+SFT
LLM LoRA
79.20
89.55
72.18
70.21
44.03
46.23
42.96
56.53
(+1.27)
EviViT+SFT
LLM LoRA
79.18
88.63
73.10
70.92
53.73
53.77
50.30
61.58
(+0.84)
Appendix
Table 10: Matched recovery SFT on Qwen3-VL-8B. Matched 10K data, order, LoRA, and epoch-3 endpoint; EviViT stays frozen. Avg. equally weights VCoT, VP-E/M/H, and Zoom. Green gains compare each SFT row with its frozen path.
Figure 6: Matched SFT training dynamics. Qwen3-VL-8B Base+SFT and EviViT+SFT follow the same 10K-example sequence for three epochs. Answer loss (left) and answer-token accuracy (right) are shown as 400-example moving means; only language LoRA is updated.
MMBench
MMStar
Visual-CoT
Host
Base
+EviViT
Δ
Base
+EviViT
Δ
Base
+EviViT
Δ
Qwen3-VL-4B
87.36
88.43
+1.06
61.73
62.80
+1.07
77.53
75.52
-2.01
Qwen3-VL-8B
88.59
88.94
+0.35
65.07
65.27
+0.20
77.92
78.35
+0.43
Qwen3.5-4B
86.37
87.76
+1.39
63.20
64.40
+1.20
77.18
78.17
+1.00
Qwen3.5-9B
87.13
86.76
-0.37
64.33
65.20
+0.87
78.98
80.34
+1.36
ZwZ-4B
87.43
85.59
-1.85
62.20
59.20
-3.00
77.77
74.96
-2.81
Appendix
Table 11: Broad-task compatibility across foundation and post-trained hosts. Accuracy uses matched inputs, deterministic decoding, and the Qwen3-VL-8B semantic judge.
Interface
Regional pixel source
VP-All
V ∗
Pooled
Internal, learned bridge
Original image
57.48
91.10
66.57
Internal, identity bridge
Original image
56.89
90.05
65.86
Native multi-image
Original image
57.09
92.15
66.57
Internal, identity bridge
Resized global view
57.28
91.62
66.57
Appendix
Table 13: Interface portability with frozen evidence views. Qwen3-VL-8B uses the same regions and per-view token counts. Pooled accuracy weights all 706 questions equally.
Region layout
Mean
Min
All90
Area
Deployed EviViT regions
73.69
68.28
64.52
12.39
Joint random translation
18.47
15.95
12.38
12.39
Joint centering
25.83
22.55
15.59
12.39
Independent random translation
14.47
11.73
9.00
12.20
Appendix
Table 14: Coverage of official targets by deployed local crops. All 186 V ∗ examples and 240 targets are retained. Mean averages target coverage within each sample, then across samples; Min averages each sample’s least-covered target. All90 requires every target to reach 90% coverage. Area is crop-union/image area. All values are percentages.
We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transformers compress each fixed-size patch into a single token, although fine-grained distinctions often depend on localized variations within only a few patches. SubViT addresses this mismatch by representing discriminative patches with multiple subtokens while retaining the original token sequence for global context, thereby allocating additional capacity where it is most needed. Since attention heads encode complementary semantics and extracting attention maps at inference requires an extra backbone forward, we adopt a two-stage training strategy. Stage 1 fine-tunes the ViT using subdivision regions sampled from random attention heads, exposing the model to diverse subdivision patterns. Stage 2 identifies informative attention maps through feature-degradation distances and distills them into a lightweight single-map router, which directly predicts deterministic token-importance scores without a separate attention forward. We evaluate SubViT on Generalized Category Discovery (GCD), a challenging task requiring both fine-grained discrimination and generalization to unlabeled novel categories. Across CUB, FGVC-Aircraft, and Stanford-Cars, SubViT improves the average novel-category accuracy of DINOv2 from 81.3% to 84.7%, with only 0.50 ms additional latency and 3.4% more FLOPs, while reducing latency by 73.8% relative to Retina Patch. Code: \href{https://github.com/jiezhu23/SubViT_ACCV26}{SubViT}.
Jie Zhu, Ivy Zhang, Minchul Kim +1
Michigan State University · Cranbrook Kingswood School · University of North Carolina at Chapel Hill
Multimodal large language models (MLLMs) fail at fine-grained visual questions less because they cannot reason than because they never see the evidence: high-resolution images are downsampled before encoding, so the model answers from linguistic priors. The standard remedies are expensive: annotated answers (SFT), hand-engineered verifiers (RLVR), or a large external teacher (on-policy distillation). We ask whether the visual evidence itself can supply the signal for free. We formalize the contrastive evidence gap, the per-token log-likelihood ratio that a model assigns to its own output when conditioned on a question-relevant region versus an irrelevant one, and study it across Qwen2.5-VL-7B, Qwen3-VL-8B, and Qwen3-VL-30B-A3B on V*Bench. Our main positive result is training-free: selecting the candidate crop under which the model's answer distribution is most peaked, using a single-view, label-free criterion, discovers the answer-bearing region with no bounding boxes, training, or labels. It localizes the target 4.4 to 5.1 times better than chance and raises fine-grained accuracy from 70 percent to 85 percent at inference. We further show that the gap is complementary to the model's own confidence. Combining them predicts correctness better than either alone, with AUC up to 0.99, and flags confidently wrong answers, with AUC ranging from 0.97 to 1.00 within the high-confidence subset. All effects concentrate on perception-bottleneck questions and vanish on a global-context control. Finally, we report an honest negative result: converting the same signal into a training method, gated self-distillation (SEG-Distill), does not outperform the base model at pilot scale across three gate designs, while more aggressive gating degrades accuracy. The signal is real, but converting it into training gains remains an open problem.
Santi Ram Tiwari, Nihal Naik, Devbrat Pandey +1
KGraph AI Solutions Pvt. Ltd. Bangalore, India – 560016
Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottleneck for efficient inference. Abundant recent methods address this bottleneck by scoring token importance and pruning low-scoring tokens in a single pass. However, one-shot scoring is insufficient because a token's prompt-relevant usefulness depends on the evidence already retained. Motivated by this insight, we introduce DIVE (Dynamic Iterative Visual Evidence Construction), a training-free framework that recasts visual-token pruning as dynamic evidence construction. DIVE repeatedly selects the remaining token with the highest residual-conditioned score, updates the visual and prompt residuals to discount the evidence already explained, and re-evaluates the remaining tokens. This select-update-re-evaluate process builds a retained set of complementary, prompt-relevant evidence. Experiments across eight image-understanding benchmarks show that DIVE consistently preserves performance across token budgets. With an 88.9% reduction in visual tokens, DIVE retains 98.2% of the uncompressed model's average performance. Code is available at https://github.com/Zhong-Chenchen/DIVE.git.
Chen Zhong, Xiao An, Zijie Wang +3
State Key Laboratory of Information Engineering in Surveying Mapping and Remote Sensing, Wuhan University, China · Electronic Information School, Wuhan University, China