A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. However, existing methods primarily supervise the student's output, leaving visual understanding implicit. Our analysis reveals that a student can match the teacher's answer without relying on the same visual evidence, raising the question: how can we ensure the student responds to the visual information that actually determines the answer? To this end, we propose \textbf{Cross-World On-Policy Distillation (CW-OPD)}, which explicitly supervises the student's response to changes in visual evidence. For each example, CW-OPD constructs two visual worlds that share the question and scene context but differ in answer-critical evidence, yielding different answers. We perform on-policy distillation in both worlds and distill the teacher's cross-world belief transition, encouraging the student to match not only \emph{what} the teacher predicts but also \emph{why} its prediction changes with the evidence. A gradient analysis shows that this term is invariant to errors shared by both worlds and supplies a corrective signal invisible to endpoint matching alone. In this way, CW-OPD makes reliance on the relevant visual evidence an explicit distillation target rather than an implicit consequence of output matching. To diagnose whether a model truly grounds its answers in visual evidence, we introduce CWBench, which measures cross-world consistency via Cross-World Pair Accuracy (CWPA). Experiments on Qwen3.5-4B show that CW-OPD outperforms the strongest baseline by \textbf{1.2} points on average, and the 4B student exceeds DeepSeek-V4.1 (552B) by \textbf{22.4} CWPA points on CWBench. Code is released in https://github.com/baokou-fw2/CWAD.
Figures & tables
Figure 1: Motivating observations on the training dataset. (a) Teacher and student exhibit different cross-modal attention despite producing the same answer. (b) As the target region is progressively blurred, the student’s sensitivity to the option “E. I can’t see it” barely differs from the base model. (c) Under semantically edited worlds, the student fails to distinguish between them, whereas the teacher significantly increases the logit of the correct answer a∗ in the edited world (Welch’s t=47.6 , p<10−4 , Cohen’s d=2.13 ), confirming the significance of this effect. Student: Qwen3.5-4B. Results average the Qwen3.5-9B teacher (OPD) and the 4B EMA teacher (OPSD).
Figure 2: (a) Dual-world construction: a VLM planner writes a counterfactual edit instruction, an image editor produces the edited image, and a VLM verifier accepts the pair only if the answer flips while the question and scene are preserved. (b) CW-OPD: the student rolls out once in World A and all passes reuse its prefix τ<t ; bidirectional OPD anchors both worlds ( LWorldA+LWorldB ), and the transition loss Ltrans distills how the teacher’s belief changes with the evidence.
Model
V*Bench
HR-Bench 4K
RealWorldQA
MMVP
HallusionBench
SEED-Bench
OK-VQA
Avg.
CWPA ↑
CWFR ↓
Qwen3.5-0.8B
Base
63.87
47.09
58.56
67.21
62.36
70.44
51.50
60.15
37.71
50.00
SFT
67.21
48.53
58.42
70.46
61.97
71.22
55.80
61.94
37.90
49.72
GRPO
64.60
44.50
58.67
70.39
63.83
70.91
53.20
60.87
37.65
50.11
OPSD
71.20
69.00
66.27
67.00
64.84
74.32
62.15
67.83
41.94
48.33
OPD
70.68
70.25
65.10
70.67
63.95
70.69
61.40
67.53
39.19
48.53
Table 1: Main results. All methods are evaluated at the original image resolution. Avg. is the unweighted mean over the seven standard benchmarks. Bold denotes the best result. On CWBench we report CWPA (Eq. 9 ) and CWFR (Eq. 10 ).
Figure 3: Statistics, validation, and examples of the proposed CW-Bench. (a) Edit type distribution of the answer-critical visual interventions in the train (n=2,365) and test (n=594) splits. (b) Question type distribution of the paired queries in both splits. (c) The filtering pipeline for constructing dual-world pairs: from 5,079 candidate pairs, 4,396 pairs pass the image-edit validation (86.6%), and 2959 pairs remain after VLM filtering (67.3% pass rate, 58.3% overall yield). (d) Examples of dual-world pairs: for each example, the original and edited images are shown with the shared question; only the answer-critical evidence differs, flipping the answer between the two worlds.
Model
Size
A Acc
B Acc
CWPA ↑
CWPA-Q ↑
CWFR ↓
Kimi-K3
2.8T
70.03
72.90
50.00
62.14
42.93
DeepSeek-V4.1
552B
54.38
79.63
43.43
64.32
47.15
Qwen3.8
27B
72.05
79.29
57.58
69.93
36.18
Qwen3.5(Base)
4B
61.95
79.46
47.64
59.86
46.13
OPSD
4B
74.75
79.97
60.61
66.45
33.50
CW-OPD
4B
78.11
82.83
65.82
73.60
29.30
Table 2: Cross-world consistency of frontier models and our distilled students on CWBench. We report CWPA (Eq. 9 ) and CWFR (Eq. 10 ); higher CWPA and CWPA-Q and lower CWFR indicate more consistent evidence-grounded answering.
Method
LOPDA
LOPDB
Ltrans
A Acc.
B Acc.
CWPA ↑
CWFR ↓
Avg.
base
–
–
–
61.95
79.46
47.64
46.13
77.29
OPD
✓
–
–
60.92
82.16
49.53
44.02
78.14
OPSD
✓
–
–
74.75
79.97
60.61
33.50
79.81
DW-OPD
✓
✓
–
72.69
81.33
59.12
35.78
79.43
CW-OPD
✓
✓
✓
78.11
82.83
65.82
29.30
80.97
Table 3: Ablation of CW-OPD supervision components. DW-OPD applies bidirectional per-world OPD to both worlds but omits the transition loss( LOPDA+LOPDB ), isolating the contribution of cross-world transition supervision. We report the CWBench metrics CWPA (Eq. 9 ) and CWFR (Eq. 10 ), together with Avg., the unweighted mean over the seven standard VQA benchmarks.
Figure 4: Ablation studies on (a) the design of the cross-world transition loss and (b) the visual evidence supervision strategy, evaluated on the 4B student. Each panel reports the Avg. score over seven standard VQA benchmarks, CWPA and CWFR on CWBench, and the Answer-Transition Error (ATE): the average absolute gap between the student and the teacher in the log-likelihood change of the gold answer token between the two worlds, which is the same quantity diagnosed in Fig. 1 (c). (a) Cosine, magnitude, and KL variants of the transition loss. (b) BBox supervision, masked-answer consistency, and the full CW-OPD.
Figure 5: Training dynamics on the Student 4B (A) and Student 0.8B (B) models. (a, c) CWPA over training steps for SFT, OPD, OPSD, and CW-OPD. (b, d) Decomposition of the CW-OPD training objective into World A loss, World B loss, and cross-world transition loss.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Robustness to structured teacher-logit perturbations. Shared : the same vector b is added in both worlds, under which the transition target is invariant (Proposition 1 (i)); opposite : +b / −b offsets the teacher contrast dT by 2b . Shared corruption degrades DW-OPD while CW-OPD is nearly unaffected; opposite corruption reverses the ordering, confirming that the two error structures act on the two objectives in opposite order. Error bars: std over 3 seeds.
Teacher perturbation
Method
Clean
Shared β=0.5
Shared β=1.0
Opp. β=0.5
Opp. β=1.0
CWPA ↑
DW-OPD ( λ=0 )
59.12
53.0 ± 1.2
47.0 ± 1.4
57.5 ± 1.1
56.0 ± 1.2
CW-OPD ( λ=1 )
65.82
64.6 ± 0.9
64.0 ± 1.0
60.0 ± 0.8
55.2 ± 1.3
Avg. over seven benchmarks ↑
DW-OPD ( λ=0 )
79.43
77.8 ± 0.3
76.1 ± 0.3
79.0 ± 0.2
78.6 ± 0.1
Appendix
Table 4: CWPA and Avg. under structured teacher perturbations (4B student, Qwen3.5-9B teacher; values averaged over 3 seeds). Clean conditions reproduce Table 3 .
Method
V*Bench
HR-Bench 4K
RealWorldQA
MMVP
HalluB
SEED-B
OK-VQA
Avg.
CWPA ↑
CWFR ↓
Base
81.15
83.00
74.38
77.67
70.42
79.25
75.17
77.29
47.64
46.13
OPD
82.70
83.25
76.84
79.17
68.17
80.31
76.56
78.14
49.53
44.02
Vision-OPD
84.10
83.90
77.20
80.00
69.90
80.55
76.80
78.92
50.33
43.47
VAD
84.54
83.67
78.32
81.13
68.43
80.79
78.93
79.40
54.32
42.57
VA-OPD
85.69
84.32
76.29
78.67
72.34
79.96
78.33
79.37
54.27
40.96
FP-OPD
83.67
83.50
77.80
79.33
69.50
80.60
77.30
78.81
49.05
44.36
Appendix
Table 5: Results with a standard larger teacher (Qwen3.5-9B → Qwen3.5-4B). CW-OPD (OPD) applies cross-world transition supervision on top of standard OPD. CWPA and CWFR are defined in Eqs. 9 and 10 .
Hyperparameter
Value
Optimization & Training
Optimizer
AdamW
Learning Rate
2e-6
Weight Decay
1e-2
LR Schedule
Constant
Epochs
1
Appendix
Table 6: Key hyperparameters used for Qwen3.5-4B.
Hyperparameter
Value
Optimization & Training
Optimizer
AdamW
Learning Rate
2e-6
Weight Decay
1e-2
LR Schedule
Constant
Epochs
1
Appendix
Table 7: Key hyperparameters used for Qwen3.5-0.8B.
λ
OPD + CW-OPD
OPSD + CW-OPD
base
77.29
77.29
0.0
77.81
77.73
0.25
79.88
80.46
0.5
80.07
80.32
1.0
80.37
80.97
2.0
79.64
79.39
Appendix
Table 8: Ablation of the transition weight λ on the 4B model. λ=0 corresponds to bidirectional OPD without transition supervision. Each entry is the average over 3 seeds.
Category
Benchmark
Task
Scale
Metric
Visual Perception & Grounding
V*Bench ( Wu and Xie, 2024 )
MCQ
191 Qs
Accuracy
HR-Bench 4K ( Wang et al., 2024 )
MCQ
4K images
Accuracy
RealWorldQA
MCQ
765 Qs
Accuracy
MMVP ( Tong et al., 2024 )
MCQ
300 Qs
Accuracy
Hallucination Diagnosis
POPE ( Li et al., 2023b )
Binary QA
3 subsets
Acc / F1
HallusionBench ( Guan et al., 2024 )
Paired QA
1,129 Qs
Pair Acc
Appendix
Table 9: Overview of the eight standard benchmarks, organized by evaluation focus. “Task” denotes the primary question format; “Scale” reports the approximate number of evaluation samples or questions; “Metric” is the primary evaluation metric.
On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-policy distillation can improve a student's output quality while failing to strengthen its reliance on visual input: on vision-critical tokens, the student's predictions remain largely unchanged whether or not fine-grained visual detail is present, even though the teacher's predictions depend heavily on it.To make this difference observable, we introduce visual advantage (VA), the token-level log-probability difference when the teacher scores a student-generated rollout with versus without access to fine-grained visual detail. VA is concentrated in a small minority of tokens, and these high-VA tokens are the ones that actually carry the visual supervision signal. This motivates a distillation objective that treats them differently from language scaffolding, so their contribution is not diluted by the abundant surrounding language tokens.We propose Visual-Advantage On-Policy Distillation (VA-OPD), which uses VA at two granularities: rollout-level reweighting by trajectory-averaged VA, and token-level KL averaged within high-VA and low-VA groups separately. We train on two math datasets (Geometry3K and ViRL39K) and evaluate on eight benchmarks covering both mathematical reasoning and visual understanding, across three teacher sizes (4B, 8B, and 32B) on the Qwen3-VL family. VA-OPD improves over standard on-policy distillation on every benchmark, with the gain growing monotonically along both the teacher-size and data-scale axes, suggesting that these factors compound consistently.
Ruiqi Liu, Xiaolei Lv, Gengsheng Li +8
Institute of Automation, CAS · School of Advanced Interdisciplinary Sciences, UCAS · 3Hello +1
On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefixes along those trajectories. This aligns the distillation states with the student's own generation distribution. However, it still assumes that the complete teacher distribution is an appropriate target across student capacities. In vision--language reasoning, teacher corrections can depend on visual distinctions that a compact student cannot represent. Our target-scaling study shows that, as the target approaches the complete teacher distribution, the student realizes less of the prescribed shift and obtains worse downstream performance. We therefore propose \emph{Fisher-Projected On-Policy Distillation} (FP-OPD), which distills only locally realizable teacher corrections. FP-OPD uses continuous visual perturbations to estimate the student's local visual tangent space and projects the centered teacher--student log-probability gap onto this space under the student's Fisher metric. The resulting capacity-aware target is optimized with full-vocabulary reverse KL on student trajectories, retaining the standard OPD framework. In 8B-to-2B distillation, FP-OPD improves all seven evaluated multimodal benchmarks. It raises the average score by 2.77 points over the pretrained student and by 1.60 points over standard OPD. These results demonstrate that locally realizable teacher corrections provide a more effective target for distilling compact vision--language models.
Leyan Xue, Feng Xiong, Mingjun Ma +1
College of Intelligence and Computing Tianjin University Tianjin, China
Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annotated data, distillation from external models, reinforcement learning with human feedback, or verifiable answers. This limits their ability to improve without external supervision. To tackle this, we propose NOPD (Noisy Student On-Policy Self-Distillation), a simple yet effective self-distillation approach that improves VLMs without any external models or ground-truth answers. Our key insight is that prediction discrepancies between clean and corrupted inputs naturally induce a self-supervision signal. In NOPD, the model learns from corrupted inputs while using its own predictions under clean inputs as token-level supervision. We show the effectiveness of NOPD on five visual reasoning tasks; it can match and even outperform reinforcement learning approaches or distillation from external models. Notably, when trained with 2.1K samples from Geometry3K, NOPD improves Qwen2.5-VL-7B by 20 points on its validation set. It also shows generalization on out-of-distribution test sets and achieves 7.4 point gains on MathVista. Furthermore, we demonstrate that NOPD is a general approach to enhance VLMs, achieving improvements across three models on 12 benchmarks.
Shuai Wang, Daoan Zhang, Zhe Tang +2
The Hong Kong University of Science and Technology (Guangzhou) · University of Rochester · Zhejiang University of Technology +1