Fine-grained visual understanding requires models to recognize small details within complex images. Multimodal on-policy self-distillation (OPSD) addresses this challenge by using a teacher conditioned on evidence-centered crops to supervise a student conditioned on original images along student-generated trajectories. Ideally, teacher corrections, the distributional changes from the student toward the privileged teacher, should be driven by task-relevant visual evidence. However, the designs that make the teacher effective also introduce other interference. Using a lagged or frozen teacher improves training stability but introduces a model-state gap from the evolving student, while cropping enhances task-relevant evidence but also loses the visual context. These two sources of interference make the teacher corrections not purely rely on the visual evidence. We introduce Evidence-Aligned multimodal on-policy self-Distillation (EAD), which retains the crop-conditioned teacher as the target but constructs a separate evidence reference for weighting the corrections. To exclude the effect of lagged model-state from this reference, EAD measures prediction changes using the current student. To avoid crop-induced context changes, EAD masks the evidence region in the original image while preserving the other visual context. The change from the student's masked-image prediction to its original-image prediction provides a controlled reference for the direction in which the visual evidence shifts the student's prediction. EAD weights each teacher correction by its cosine alignment with the reference, i.e., retaining aligned corrections and downweighting the rest. Retaining only 6% of the supervision mass of dense OPSD, EAD consistently outperforms previous state-of-the-art methods.
Figures & tables
Figure 1: Motivation and overview of EAD. The privileged teacher correction in multimodal OPSD contains both the desired evidence-driven signal and interference arising from model-state gap and crop-induced context changes (top left). EAD constructs a separate student-side evidence reference from the current student’s prediction shift induced by evidence restoration, which keeps the model state fixed and preserves the surrounding visual context (top right). EAD then weights teacher corrections by the alignment with this evidence reference, concentrating supervision on visually supported corrections (bottom).
Figure 2: Diagnosing interference in teacher corrections. (a) Total Variation (TV) distance between the normalized token-wise JS distributions of Dtteacher and Dtcrop over 512 images, showing the effect of model-state gap across training. (b) TV distance between Dtcrop and Dtmask under the current student, showing the difference between crop-induced changes and evidence removal. (c) The corresponding token-wise JS distributions for one example. Dtmask places greater emphasis on tokens associated with the task-relevant visual evidence. Dtcrop shows large divergence on less vision-relevant tokens.
Figure 3: Overview of EAD. The crop-conditioned EMA teacher provides the privileged target pθˉ,tC , while the current student produces pθ,tR and pθ,tR− on the original and masked images, respectively. The Original–Mask prediction change pθ,tR−pθ,tR− forms the evidence reference. EAD compares the teacher correction and evidence reference after a shared JS-motivated normalization, and uses their clipped cosine similarity to weight the distillation loss.
Model
Params
V*Bench
ZoomBench
HR-Bench
MME-RealWorld
Avg.
4K
8K
EN
CN
Closed-Source Models (Single Forward Pass)
GPT-5.2 ( OpenAI, 2025 )
–
79.06
50.89
81.12
78.38
72.60
68.80
71.81
GPT-5.4 ( OpenAI, 2026 )
–
76.96
52.66
84.00
77.88
74.20
70.93
72.77
Gemini-3.1-Pro ( Google, 2026a )
–
87.96
61.18
89.63
86.88
76.53
73.31
79.25
Gemini-3.5-Flash ( Google, 2026b )
–
89.01
61.42
89.12
86.62
75.31
73.97
79.24
Table 1: Fine-grained visual understanding. We report results on six fine-grained visual benchmarks under the single-forward-pass evaluation setting. Results for closed-source and open-weight models are taken from Yuan et al. (2026) , while all Qwen3.5-4B/9B-based methods are obtained from our evaluation. All post-training methods use the same backbone and training data, are trained for one epoch, and are evaluated using the final checkpoint. Avg. is the unweighted mean over the six benchmarks. Best results within each Qwen3.5 scale are shown in bold.
Model State
V*Bench
ZoomBench
HR-Bench
MME-RealWorld
Avg.
4K
8K
EN
CN
Cross-state
89.53
59.05
85.12
81.38
71.05
69.75
75.98
Teacher-side
89.01
59.76
85.25
83.62
73.18
71.35
77.03
Student-side (EAD)
91.62
60.83
86.38
82.88
72.39
71.10
77.53
Table 2: Effect of model state on the evidence reference. All variants use the same Original–Mask image pair. Cross-state uses different model states for the two predictions, whereas Teacher-side and Student-side (EAD) use the EMA teacher and current student alone, respectively. Avg. denotes the unweighted mean over the six benchmarks. Best results in each column are shown in bold.
Visual Contrast
V*Bench
ZoomBench
HR-Bench
MME-RealWorld
Avg.
4K
8K
EN
CN
Original–Crop
88.48
60.00
83.00
80.25
70.61
68.62
75.16
Original–Mask (EAD)
91.62
60.83
86.38
82.88
72.39
71.10
77.53
Crop–Downsampled Crop (VAD)
91.62
58.82
86.50
83.00
69.33
68.73
76.33
Table 3: Effect of visual contrast on the evidence reference. The first two variants use the current student and differ only in the visual inputs, while the VAD-style variant uses the EMA teacher on the cropped image and its downsampled counterpart. All variants use the same EAD objective. Avg. denotes the unweighted mean over the benchmarks. Best results in each column are shown in bold.
Variant
V*Bench
ZoomBench
HRBench-4K
HRBench-8K
MME-EN
MME-CN
Avg.
No JS Scaling
89.01
59.29
85.12
82.38
72.37
71.29
76.58
Shifted-Cosine
90.58
61.18
83.50
81.25
72.22
70.05
76.46
EAD
91.62
60.83
86.38
82.88
72.39
71.10
77.53
Table 4: Effect of local JS scaling and clipped cosine similarity. No JS Scaling removes the local normalization before computing directional alignment, while Shifted-Cosine retains negatively aligned teacher corrections.
Figure 4: Sparsity and concentration of EAD supervision. (a) Fraction of the original distillation-loss mass retained by EAD throughout training. (b) Concentration of the retained loss across response tokens, compared with the original dense distillation loss on the same responses. (c) On-policy rollout accuracy of EAD and Vision-OPD during training. Results are shown for Qwen3.5-4B and Qwen3.5-9B.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
MMVP
CV-Bench
MMStar
POPE
Avg.
Qwen3.5-4B Post-Training Strategies
Vanilla
79.00
86.97
74.00
89.33
82.33
GRPO ( Shao et al., 2024 )
78.00
85.78
69.33
87.87
80.25
Vision-OPD ( Yuan et al., 2026 )
76.67
85.48
70.80
88.99
80.49
OPD-V ( Bi et al., 2026 )
77.00
84.38
71.40
89.43
80.55
VAD ( Zhang et al., 2026 )
79.67
87.30
71.80
88.79
81.89
Appendix
Table 5: Preservation of broader visual capabilities. We evaluate the same 4B and 9B post-training checkpoints from Table 1 on four held-out benchmarks without additional training. Avg. denotes the unweighted mean over the four benchmarks. Best results within each model scale are shown in bold.
Benchmark
Vision-OPD 4B
V*Bench
90.05
ZoomBench
59.29
HRBench-4K
82.25
HRBench-8K
79.25
MME-RealWorld-EN
70.71
MME-RealWorld-CN
69.43
Appendix
Table 6: Released Vision-OPD 4B checkpoint under our evaluation pipeline. All results are obtained using the same revised evaluator applied to the Qwen3.5-based models in this work.
Statistic
4B
9B
Cumulative retained loss mass
5.87
6.15
Top-10% loss mass: EAD
95.98
96.81
Top-10% loss mass: dense
61.10
62.27
Training rollout accuracy: EAD
59.53
64.31
Training rollout accuracy: Vision-OPD
55.06
57.49
Appendix
Table 7: Summary of supervision retention, concentration, and rollout accuracy. All values are percentages. Top-10% loss mass measures how much of each response’s total loss comes from its 10% highest-loss tokens, averaged across responses. Rollout accuracy is computed over all responses from training steps 1–65.
Evidence Reference
MMVP
CV-Bench
MMStar
POPE
Holdout Avg.
Fine Avg.
All-10
Cross-state masking
77.00
85.49
73.53
86.66
80.67
75.98
77.86
Crop-induced change
78.33
86.61
70.87
89.11
81.23
75.16
77.59
Teacher-side masking
78.67
87.59
72.67
89.24
82.04
77.03
79.03
Teacher crop degradation
76.33
84.15
71.87
83.30
78.91
76.33
77.37
Current student (EAD)
80.33
86.95
73.07
87.70
82.01
77.53
79.33
Appendix
Table 8: Held-out generalization of evidence-reference variants. The same checkpoints evaluated in Tables 2 and 3 are evaluated on MMVP, CV-Bench, MMStar, and POPE. Holdout Avg. is the unweighted mean over the four held-out benchmarks, Fine Avg. is the mean over the six fine-grained benchmarks, and All-10 averages all ten benchmarks.
Benchmark
Vanilla Cosine
EAD
V*Bench
89.01
91.62
ZoomBench
59.29
60.83
HRBench-4K
85.12
86.38
HRBench-8K
82.38
82.88
MME-RealWorld-EN
72.37
72.39
MME-RealWorld-CN
71.29
71.10
Appendix
Table 9: Effect of local JS scaling. Vanilla cosine directly compares the unscaled probability differences, whereas EAD applies the shared local normalization 1/max(pθ,tR,ϵ) before measuring directional alignment. Both variants otherwise use the same evidence reference and training procedure.
Benchmark
Shifted-Cosine
EAD
V*Bench
90.58
91.62
ZoomBench
61.18
60.83
HRBench-4K
83.50
86.38
HRBench-8K
81.25
82.88
MME-RealWorld-EN
72.22
72.39
MME-RealWorld-CN
70.05
71.10
Appendix
Table 10: Effect of excluding negatively aligned corrections. Shifted-Cosine retains negatively aligned corrections by mapping cosine similarity from [−1,1] to [0,1] . All other settings are unchanged.
Figure 5: Additional token-level case studies. Each example shows the original, cropped, and masked images together with the normalized token-wise JS divergence of Dtteacher , Dtcrop , and Dtmask . Across different fine-grained visual tasks, Dtmask tends to emphasize tokens more closely associated with the task-relevant visual evidence, while the other comparisons can assign large divergence to less directly related tokens.
Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Xiaohongshu Inc.