Fine-grained visual understanding requires models to recognize small details within complex images. Multimodal on-policy self-distillation (OPSD) addresses this challenge by using a teacher conditioned on evidence-centered crops to supervise a student conditioned on original images along student-generated trajectories. Ideally, teacher corrections, the distributional changes from the student toward the privileged teacher, should be driven by task-relevant visual evidence. However, the designs that make the teacher effective also introduce other interference. Using a lagged or frozen teacher improves training stability but introduces a model-state gap from the evolving student, while cropping enhances task-relevant evidence but also loses the visual context. These two sources of interference make the teacher corrections not purely rely on the visual evidence. We introduce Evidence-Aligned multimodal on-policy self-Distillation (EAD), which retains the crop-conditioned teacher as the target but constructs a separate evidence reference for weighting the corrections. To exclude the effect of lagged model-state from this reference, EAD measures prediction changes using the current student. To avoid crop-induced context changes, EAD masks the evidence region in the original image while preserving the other visual context. The change from the student's masked-image prediction to its original-image prediction provides a controlled reference for the direction in which the visual evidence shifts the student's prediction. EAD weights each teacher correction by its cosine alignment with the reference, i.e., retaining aligned corrections and downweighting the rest. Retaining only 6% of the supervision mass of dense OPSD, EAD consistently outperforms previous state-of-the-art methods.
Figures & tables
Figure 1: Motivation and overview of EAD. The privileged teacher correction in multimodal OPSD contains both the desired evidence-driven signal and interference arising from model-state gap and crop-induced context changes (top left). EAD constructs a separate student-side evidence reference from the current student’s prediction shift induced by evidence restoration, which keeps the model state fixed and preserves the surrounding visual context (top right). EAD then weights teacher corrections by the alignment with this evidence reference, concentrating supervision on visually supported corrections (bottom).
Figure 2: Diagnosing interference in teacher corrections. (a) Total Variation (TV) distance between the normalized token-wise JS distributions of Dtteacher and Dtcrop over 512 images, showing the effect of model-state gap across training. (b) TV distance between Dtcrop and Dtmask under the current student, showing the difference between crop-induced changes and evidence removal. (c) The corresponding token-wise JS distributions for one example. Dtmask places greater emphasis on tokens associated with the task-relevant visual evidence. Dtcrop shows large divergence on less vision-relevant tokens.
Figure 3: Overview of EAD. The crop-conditioned EMA teacher provides the privileged target pθˉ,tC , while the current student produces pθ,tR and pθ,tR− on the original and masked images, respectively. The Original–Mask prediction change pθ,tR−pθ,tR− forms the evidence reference. EAD compares the teacher correction and evidence reference after a shared JS-motivated normalization, and uses their clipped cosine similarity to weight the distillation loss.
Model
Params
V*Bench
ZoomBench
HR-Bench
MME-RealWorld
Avg.
4K
8K
EN
CN
Closed-Source Models (Single Forward Pass)
GPT-5.2 ( OpenAI, 2025 )
–
79.06
50.89
81.12
78.38
72.60
68.80
71.81
GPT-5.4 ( OpenAI, 2026 )
–
76.96
52.66
84.00
77.88
74.20
70.93
72.77
Gemini-3.1-Pro ( Google, 2026a )
–
87.96
61.18
89.63
86.88
76.53
73.31
79.25
Gemini-3.5-Flash ( Google, 2026b )
–
89.01
61.42
89.12
86.62
75.31
73.97
79.24
Table 1: Fine-grained visual understanding. We report results on six fine-grained visual benchmarks under the single-forward-pass evaluation setting. Results for closed-source and open-weight models are taken from Yuan et al. (2026) , while all Qwen3.5-4B/9B-based methods are obtained from our evaluation. All post-training methods use the same backbone and training data, are trained for one epoch, and are evaluated using the final checkpoint. Avg. is the unweighted mean over the six benchmarks. Best results within each Qwen3.5 scale are shown in bold.
Model State
V*Bench
ZoomBench
HR-Bench
MME-RealWorld
Avg.
4K
8K
EN
CN
Cross-state
89.53
59.05
85.12
81.38
71.05
69.75
75.98
Teacher-side
89.01
59.76
85.25
83.62
73.18
71.35
77.03
Student-side (EAD)
91.62
60.83
86.38
82.88
72.39
71.10
77.53
Table 2: Effect of model state on the evidence reference. All variants use the same Original–Mask image pair. Cross-state uses different model states for the two predictions, whereas Teacher-side and Student-side (EAD) use the EMA teacher and current student alone, respectively. Avg. denotes the unweighted mean over the six benchmarks. Best results in each column are shown in bold.
Visual Contrast
V*Bench
ZoomBench
HR-Bench
MME-RealWorld
Avg.
4K
8K
EN
CN
Original–Crop
88.48
60.00
83.00
80.25
70.61
68.62
75.16
Original–Mask (EAD)
91.62
60.83
86.38
82.88
72.39
71.10
77.53
Crop–Downsampled Crop (VAD)
91.62
58.82
86.50
83.00
69.33
68.73
76.33
Table 3: Effect of visual contrast on the evidence reference. The first two variants use the current student and differ only in the visual inputs, while the VAD-style variant uses the EMA teacher on the cropped image and its downsampled counterpart. All variants use the same EAD objective. Avg. denotes the unweighted mean over the benchmarks. Best results in each column are shown in bold.
Variant
V*Bench
ZoomBench
HRBench-4K
HRBench-8K
MME-EN
MME-CN
Avg.
No JS Scaling
89.01
59.29
85.12
82.38
72.37
71.29
76.58
Shifted-Cosine
90.58
61.18
83.50
81.25
72.22
70.05
76.46
EAD
91.62
60.83
86.38
82.88
72.39
71.10
77.53
Table 4: Effect of local JS scaling and clipped cosine similarity. No JS Scaling removes the local normalization before computing directional alignment, while Shifted-Cosine retains negatively aligned teacher corrections.
Figure 4: Sparsity and concentration of EAD supervision. (a) Fraction of the original distillation-loss mass retained by EAD throughout training. (b) Concentration of the retained loss across response tokens, compared with the original dense distillation loss on the same responses. (c) On-policy rollout accuracy of EAD and Vision-OPD during training. Results are shown for Qwen3.5-4B and Qwen3.5-9B.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
MMVP
CV-Bench
MMStar
POPE
Avg.
Qwen3.5-4B Post-Training Strategies
Vanilla
79.00
86.97
74.00
89.33
82.33
GRPO ( Shao et al., 2024 )
78.00
85.78
69.33
87.87
80.25
Vision-OPD ( Yuan et al., 2026 )
76.67
85.48
70.80
88.99
80.49
OPD-V ( Bi et al., 2026 )
77.00
84.38
71.40
89.43
80.55
VAD ( Zhang et al., 2026 )
79.67
87.30
71.80
88.79
81.89
Appendix
Table 5: Preservation of broader visual capabilities. We evaluate the same 4B and 9B post-training checkpoints from Table 1 on four held-out benchmarks without additional training. Avg. denotes the unweighted mean over the four benchmarks. Best results within each model scale are shown in bold.
Benchmark
Vision-OPD 4B
V*Bench
90.05
ZoomBench
59.29
HRBench-4K
82.25
HRBench-8K
79.25
MME-RealWorld-EN
70.71
MME-RealWorld-CN
69.43
Appendix
Table 6: Released Vision-OPD 4B checkpoint under our evaluation pipeline. All results are obtained using the same revised evaluator applied to the Qwen3.5-based models in this work.
Statistic
4B
9B
Cumulative retained loss mass
5.87
6.15
Top-10% loss mass: EAD
95.98
96.81
Top-10% loss mass: dense
61.10
62.27
Training rollout accuracy: EAD
59.53
64.31
Training rollout accuracy: Vision-OPD
55.06
57.49
Appendix
Table 7: Summary of supervision retention, concentration, and rollout accuracy. All values are percentages. Top-10% loss mass measures how much of each response’s total loss comes from its 10% highest-loss tokens, averaged across responses. Rollout accuracy is computed over all responses from training steps 1–65.
Evidence Reference
MMVP
CV-Bench
MMStar
POPE
Holdout Avg.
Fine Avg.
All-10
Cross-state masking
77.00
85.49
73.53
86.66
80.67
75.98
77.86
Crop-induced change
78.33
86.61
70.87
89.11
81.23
75.16
77.59
Teacher-side masking
78.67
87.59
72.67
89.24
82.04
77.03
79.03
Teacher crop degradation
76.33
84.15
71.87
83.30
78.91
76.33
77.37
Current student (EAD)
80.33
86.95
73.07
87.70
82.01
77.53
79.33
Appendix
Table 8: Held-out generalization of evidence-reference variants. The same checkpoints evaluated in Tables 2 and 3 are evaluated on MMVP, CV-Bench, MMStar, and POPE. Holdout Avg. is the unweighted mean over the four held-out benchmarks, Fine Avg. is the mean over the six fine-grained benchmarks, and All-10 averages all ten benchmarks.
Benchmark
Vanilla Cosine
EAD
V*Bench
89.01
91.62
ZoomBench
59.29
60.83
HRBench-4K
85.12
86.38
HRBench-8K
82.38
82.88
MME-RealWorld-EN
72.37
72.39
MME-RealWorld-CN
71.29
71.10
Appendix
Table 9: Effect of local JS scaling. Vanilla cosine directly compares the unscaled probability differences, whereas EAD applies the shared local normalization 1/max(pθ,tR,ϵ) before measuring directional alignment. Both variants otherwise use the same evidence reference and training procedure.
Benchmark
Shifted-Cosine
EAD
V*Bench
90.58
91.62
ZoomBench
61.18
60.83
HRBench-4K
83.50
86.38
HRBench-8K
81.25
82.88
MME-RealWorld-EN
72.22
72.39
MME-RealWorld-CN
70.05
71.10
Appendix
Table 10: Effect of excluding negatively aligned corrections. Shifted-Cosine retains negatively aligned corrections by mapping cosine similarity from [−1,1] to [0,1] . All other settings are unchanged.
Figure 5: Additional token-level case studies. Each example shows the original, cropped, and masked images together with the normalized token-wise JS divergence of Dtteacher , Dtcrop , and Dtmask . Across different fine-grained visual tasks, Dtmask tends to emphasize tokens more closely associated with the task-relevant visual evidence, while the other comparisons can assign large divergence to less directly related tokens.
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.
Kangning Zhang, Yixing Li, Shuai Shao +9
Shanghai Jiao Tong University · Xiaohongshu Inc. · The Chinese University of Hong Kong +1
Multimodal Large Language Models (MLLMs) still struggle with fine-grained visual understanding, where answers often depend on small but decisive evidence in the full image. We observe a regional-to-global perception gap: the same MLLM answers fine-grained questions more accurately when conditioned on evidence-centered crops than on the corresponding full images, suggesting that many failures stem from difficulty to focus on relevant evidence rather than insufficient local recognition ability. Motivated by this observation, we propose Vision-OPD (Vision On-Policy Distillation), a regional-to-global self-distillation framework that transfers the model's own privileged regional perception to its full-image policy. Vision-OPD instantiates two conditional policies from the same MLLM: a crop-conditioned teacher and a full-image-conditioned student. The student generates on-policy rollouts, and Vision-OPD minimizes token-level divergence between the teacher and student next-token distributions along these rollouts. This enables the model to internalize the benefit of visual zooming without external teacher models, ground-truth labels, reward verifiers, or inference-time tool use. Experiments on multiple fine-grained visual understanding benchmarks show that Vision-OPD models achieve competitive or superior performance against much larger open-source, closed-source, and "Thinking-with-Images" agentic models. The code is available at https://github.com/VisionOPD/Vision-OPD
Qianhao Yuan, Jie Lou, Xing Yu +4
Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Xiaohongshu Inc.
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories while conditioning on the same student-generated prefix. When a student misinterprets an image early in a response, this accumulating erroneous rationale eventually pulls the teacher away from its visual evidence. The teacher and student converge on the same hallucination, causing standard cross-model supervision to collapse precisely where correction is most needed. We find that the teacher's visual corrective preference is not lost under this misleading agreement. Comparing the predictions of the identical teacher given the real image and a visual null reveals that the privileged evidence still pushes the model toward the correct interpretation. We introduce OPD-Aha, which reconstructs the distillation target directly from this isolated visual preference rather than relying on the fragile teacher-student discrepancy. This reconstructed target aggressively suppresses continuations that contradict the image. Trained with this objective, students learn to naturally interrupt their own flawed reasoning with reflection tokens such as wait and actually. After reflection, subsequent generation relies less on the accumulated erroneous text and more on the visual evidence. Correcting these trajectories mid-generation fundamentally alters the reasoning process, yielding broad and consistent improvements across diverse fine-grained perception and complex multimodal reasoning benchmarks. Our code and models are available at https://github.com/Echochef/OPD-Aha.
Chenhao Qiu, Dawei Li, Yechao Zhang +2
Arizona State University · University of Virginia · Stevens Institute of Technology