EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception
Authors: Yaoxin Niu, Zhangquan Chen, Yang Zhang, Xiang An, Zhumei Wang, Chih-Ting Liao, Hongkun Cao, Ruqi Huang
Organizations: Tsinghua University · Peng Cheng Laboratory · The Hong Kong University of Science and Technology · LMMs-Lab · Beijing Institute of Technology · University of New South Wales
Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.
Figures & tables
Figure 2: EviViT pipeline. Human interactions are distilled into a question-conditioned evidence target. PTEA predicts this distribution inside the host ViT and yields one or two non-redundant decisive regions plus one context region. These regions are re-read from the original pixels, fused with coordinate-matched global tokens, and processed by the remaining frozen ViT blocks and LLM.
Figure 3: Strong-model accuracy rankings. VisualProbe VP-All (515 questions) and ZoomBench (845 questions), sorted within each panel. Open-checkpoint bars use our local evaluation. The SeProD VP-All bar is derived from its author-reported Easy/Medium/Hard scores under the original avg@32 protocol, so it is contextual rather than a matched comparison.
Figure 4: Budget curves and paired gains. Accuracy versus tokens on Qwen3-VL-4B/8B.
TTFT (s)
E2E (s)
Budget
Path
Vis. tok. (K)
Median
P90
Median
P90
Peak (GiB)
Δ Acc. (pp)
Qwen3-VL-4B
1K
Global
1.00
0.206
0.218
0.557
0.619
8.62
–
EviViT
0.90
0.384
0.403
0.738
0.804
8.66
+6.05
4K
Global
3.83
0.559
0.643
0.921
1.038
9.48
–
EviViT
3.62
0.621
0.712
0.999
1.091
9.47
+4.76
Table 3: Whole-host-isolated efficiency. Each budget pairs Global and EviViT. Tokens and accuracy use the complete sweep; latency and peak memory use 64 isolated requests per setting.
MMBench
MMStar
Visual-CoT
Host
Base
+EviViT
Δ
Base
+EviViT
Δ
Base
+EviViT
Δ
Qwen3-VL-4B
87.36
88.43
+1.06
61.73
62.80
+1.07
77.53
75.52
-2.01
Qwen3-VL-8B
88.59
88.94
+0.35
65.07
65.27
+0.20
77.92
78.35
+0.43
Qwen3.5-4B
86.37
87.76
+1.39
63.20
64.40
+1.20
77.18
78.17
+1.00
Qwen3.5-9B
87.13
86.76
-0.37
64.33
65.20
+0.87
78.98
80.34
+1.36
Table 4: General-capability compatibility on four foundation hosts.
Supervision
VP-All
VP-Hard
V ∗
Final-box
56.12
52.83
89.01
Full-trace
57.28
56.60
90.05
Δ
+1.17
+3.77
+1.05
Table 5: Full-trace versus final-box supervision at epoch three (%).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: PTEA training and checkpoint selection. Five-fold trace-only training of the Qwen3-VL-4B Block-16 selector. Thin curves show individual folds and thick curves their mean. The held-out evidence score is shown with a one-standard-deviation band across folds; dashed lines mark the selected third epoch.
Variant
VP-E
VP-M
VP-H
V ∗
HR-4K
HR-8K
PB-1img
Zoom
Avg.
Δ
Qwen3-VL-4B
Global B16
60.28
40.30
41.51
84.29
76.25
73.38
30.77
45.80
56.57
± 0.00
+ Reread
64.54
47.01
39.62
86.91
79.50
77.12
32.05
53.25
60.00
+3.43
+ Boxes
63.83
51.49
45.28
86.39
79.25
76.12
31.60
51.60
60.70
+4.13
EviViT (full)
63.83
50.37
48.11
87.96
79.62
77.75
31.83
53.37
61.61
+5.04
Qwen3-VL-8B
Appendix
Table 6: Where the gains enter. Reread uses fixed-budget source-pixel views; Boxes adds continuous crop geometry; EviViT is the full model with input-dependent capacity allocation. Accuracy is in percent. Avg. equally weights the eight benchmarks, and Δ is relative to Global B16 (global-only, 16×10242 -pixel cap) at the same scale. Bold marks each column’s best score within a scale.
Regional locations
VP-All
V ∗
Pooled
Δ
Deployed locations
57.09
92.15
66.57
–
Random translation
47.77
82.20
57.08
-9.49
Cross-image shuffled locations
47.77
81.68
56.94
-9.63
Appendix
Table 8: Location control with native multi-image input. Only crop positions change; view count, size, and token budget are fixed. Accuracy (%); Δ versus deployed locations.
Host
VP tier
Base
+EviViT
Δ
Rescue/loss
Retain
p
Qwen3-VL-4B
Easy
60.28
63.83
+3.55
16/11
87.06%
0.44
Medium
40.30
50.37
+10.07
41/14
87.04%
3.6e-04
Hard
41.51
48.11
+6.60
17/10
77.27%
0.25
Qwen3-VL-8B
Easy
67.38
69.50
+2.13
14/11
88.42%
0.69
Medium
42.16
53.36
+11.19
45/15
86.73%
1.3e-04
Hard
44.34
51.89
+7.55
19/11
76.60%
0.2
Appendix
Table 9: Paired gains come from rescues rather than answer churn. Rescue/loss compares identical questions. Retain is the fraction of originally correct answers that remain correct; p is the two-sided exact McNemar p -value. Qwen3-VL and Qwen3.5 results come from completed answer-only paired runs using the same executor.
Model
Trainable
VCoT
Weak-5
Protect-5
VP-E
VP-M
VP-H
Zoom
Avg.
Base
Frozen
78.50
89.53
70.86
67.38
42.16
44.34
43.91
55.26
EviViT
Frozen
78.31
87.94
71.82
69.50
53.36
51.89
50.65
60.74
Base+SFT
LLM LoRA
79.20
89.55
72.18
70.21
44.03
46.23
42.96
56.53
(+1.27)
EviViT+SFT
LLM LoRA
79.18
88.63
73.10
70.92
53.73
53.77
50.30
61.58
(+0.84)
Appendix
Table 10: Matched recovery SFT on Qwen3-VL-8B. Matched 10K data, order, LoRA, and epoch-3 endpoint; EviViT stays frozen. Avg. equally weights VCoT, VP-E/M/H, and Zoom. Green gains compare each SFT row with its frozen path.
Figure 6: Matched SFT training dynamics. Qwen3-VL-8B Base+SFT and EviViT+SFT follow the same 10K-example sequence for three epochs. Answer loss (left) and answer-token accuracy (right) are shown as 400-example moving means; only language LoRA is updated.
MMBench
MMStar
Visual-CoT
Host
Base
+EviViT
Δ
Base
+EviViT
Δ
Base
+EviViT
Δ
Qwen3-VL-4B
87.36
88.43
+1.06
61.73
62.80
+1.07
77.53
75.52
-2.01
Qwen3-VL-8B
88.59
88.94
+0.35
65.07
65.27
+0.20
77.92
78.35
+0.43
Qwen3.5-4B
86.37
87.76
+1.39
63.20
64.40
+1.20
77.18
78.17
+1.00
Qwen3.5-9B
87.13
86.76
-0.37
64.33
65.20
+0.87
78.98
80.34
+1.36
ZwZ-4B
87.43
85.59
-1.85
62.20
59.20
-3.00
77.77
74.96
-2.81
Appendix
Table 11: Broad-task compatibility across foundation and post-trained hosts. Accuracy uses matched inputs, deterministic decoding, and the Qwen3-VL-8B semantic judge.
Interface
Regional pixel source
VP-All
V ∗
Pooled
Internal, learned bridge
Original image
57.48
91.10
66.57
Internal, identity bridge
Original image
56.89
90.05
65.86
Native multi-image
Original image
57.09
92.15
66.57
Internal, identity bridge
Resized global view
57.28
91.62
66.57
Appendix
Table 13: Interface portability with frozen evidence views. Qwen3-VL-8B uses the same regions and per-view token counts. Pooled accuracy weights all 706 questions equally.
Region layout
Mean
Min
All90
Area
Deployed EviViT regions
73.69
68.28
64.52
12.39
Joint random translation
18.47
15.95
12.38
12.39
Joint centering
25.83
22.55
15.59
12.39
Independent random translation
14.47
11.73
9.00
12.20
Appendix
Table 14: Coverage of official targets by deployed local crops. All 186 V ∗ examples and 240 targets are retained. Mean averages target coverage within each sample, then across samples; Min averages each sample’s least-covered target. All90 requires every target to reach 90% coverage. Area is crop-union/image area. All values are percentages.
State Key Laboratory of Information Engineering in Surveying Mapping and Remote Sensing, Wuhan University, China · Electronic Information School, Wuhan University, China