On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teacher inputs and refining teacher feedback), yet we find that limited student perception is another critical bottleneck in multimodal OPD. By providing oracle visual facts, the performance of OPD-trained students can still be substantially improved for both weak and strong teachers. To address this bottleneck, we propose S-OPD, a simple multimodal on-policy distillation framework that explicitly strengthens student perceptual learning through two objectives. Specifically, Teacher-calibrated Policy Contrast separates student policies under original and masked images with teacher-based token-level gating, strengthening the student's reliance on visual evidence during reasoning. Policy Agreement aligns student policies under original and noise-perturbed images, further improving perceptual robustness to visual noise. Notably, our method can be seamlessly plugged into existing OPD frameworks, requiring no additional data annotations, model parameters or inference operations. Extensive experiments on eight benchmarks across student scales and distillation paradigms demonstrate consistent performance improvements, with gains of up to 4.25 points on LogicVista. When combined with existing teacher-side supervision methods, our method can yield further gains. Code is available at https://github.com/Sirilaw/S-OPD.
Figures & tables
Figure 1: Beyond enhancing teacher-side supervision: Previous methods improve the teacher-side signal, while S-OPD adds two student-side objectives to enhance perception. This improves performance across model scales (2B and 4B) and distillation paradigms (OPD and OPSD).
Figure 2: Oracle visual facts evaluation on Geometry3K.
Figure 3: Performance across quartiles of estimated KL divergence between student policies under original and masked images (Q1–Q4, lowest to highest). A clear trend is observed: Higher-KL groups generally achieve higher accuracy.
Figure 4: (a) Details of S-OPD with original, masked, and noisy images. The teacher compares original and masked images to gate student policy contrast, while policy agreement aligns the student’s original and noisy policies. (b) S-OPD achieves higher performance and reduces the teacher-student log-probability gap during training on Geometry3K.
Mathematical Reasoning
Logical Reasoning
General Reasoning
Overall
Method
MathVerse
MathVista
MathVision
WeMath
LogicVista
VisualPuzzles
ZeroBench
MMMU
Avg.
Qwen3-VL-2B-Instruct
Base model
45.51
61.2
30.29
32.29
34.52
16.18
13.17
45.78
34.87
GRPO
48.43
62.3
31.62
36.57
42.73
23.16
14.67
49.67
38.64
OPD
46.41
65.2
34.35
36.67
36.24
12.16
9.28
48.44
36.09
S-OPD
47.81 (+1.40)
67.0 (+1.8)
35.93 (+1.58)
37.43 (+0.76)
37.65 (+1.41)
12.67 (+0.51)
11.98 (+2.70)
49.22 (+0.78)
37.46 (+1.37)
Table 1: Main results on eight benchmarks. We compare the distillation performance against the corresponding baseline. For OPD and S-OPD, we use either Qwen3-VL-8B-Instruct or its GRPO-trained variant as the teacher; the subscript G denotes the latter teacher setting. Avg. is computed over the eight benchmarks.
2B Student
4B Student
Method
Geometry3K
Mathematical
Logical
General
Avg.
Geometry3K
Mathematical
Logical
General
Avg.
OPD
40.93
45.98
25.83
30.40
37.48
46.76
54.38
42.89
36.76
47.06
w/ TPC
42.56
47.10
26.05
29.77
38.07
49.02
54.94
42.88
37.99
47.83
w/ PA
41.93
46.78
27.91
31.27
38.60
48.03
54.42
41.46
39.06
47.41
S-OPD
44.26
47.39
28.02
30.92
39.08
49.02
55.03
43.33
38.57
48.10
Table 2: Component analysis with 2B and 4B students on Geometry3K. Avg. denotes the arithmetic average over all benchmarks including the test split of Geometry3K.
Gating strategy
Token selection
Geometry3K
Mathematical
Logical
General
Avg.
OPD (reference)
–
40.93
45.98
25.83
30.40
37.48
No gating
All
42.56
46.37
26.55
30.72
38.06
Student-based
Adaptive
41.43
46.80
25.44
30.92
37.93
Random
Count-matched
40.61
47.47
26.57
30.95
38.39
Teacher-calibrated (Ours)
Adaptive
44.26
47.39
28.02
30.92
39.08
Table 3: Comparison of gating strategies for TPC with the remaining S-OPD configuration fixed. OPD is included as a reference. Random gating selects the same number of tokens as teacher-calibrated gating for each response. Bold values indicate the best results.
Figure 5: S-OPD gains across student scales and distillation paradigms.
Value
Geometry3K
Mathematical
Logical
General
Avg.
λTPC
0.005
44.26
47.39
28.02
30.92
39.08
0.01
44.92
47.73
25.83
30.93
38.82
0.02
43.05
47.00
26.64
30.36
38.34
λPA
0.01
44.92
46.85
26.55
30.91
38.58
Table 4: Ablation of the TPC and PA coefficients.
Value
Geometry3K
Mathematical
Logical
General
Avg.
Masking ratio ρ
0.2
43.43
47.11
24.94
32.34
38.49
0.4
42.43
47.18
26.65
31.20
38.54
0.6
44.26
47.39
28.02
30.92
39.08
0.8
44.43
47.34
26.02
31.42
38.74
Gaussian noise level σ
Table 5: Ablation of the masking ratio and Gaussian noise level.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
ViRL39K Distillation
ViRL39K GRPO
Geometry3K Distillation
Training data
ViRL39K
ViRL39K
Geometry3K train
Max. prompt length
4,096
4,096
1,024
Max. response length
4,096
4096
2,048
Prompt batch size
192
192
128
Learning rate
1×10−6
1×10−6
1×10−6
Rollouts per prompt ( n )
1
5
1
Appendix
Table 6: Key training configurations for the ViRL39K experiments and the Geometry3K ablations.
Figure 6: Masked images used for TPC at different masking ratios ρ . Randomly selected image patches are replaced with black pixels, removing increasing amounts of visual evidence as ρ increases.
Figure 7: Noisy images used for PA at different Gaussian noise levels σ . Independent zero-mean Gaussian noise is added to each RGB channel in the [0,1] pixel range, followed by clipping. Larger σ produces stronger perturbations while retaining the overall image structure.
Dataset
User prompt template
MathVerse_MINI
{question}
MathVista_MINI
{question}
MathVision
{question}
WeMath
Question: {question} Options: [Options]
LogicVista
{question}
VisualPuzzles
{question} Options: [Options] Solve the multiple-choice question and then answer with the option letter from the given choices. The last line of your response should be of the following format: ’Answer: $LETTER’ (without quotes), where LETTER is one of the options. Think step by step before answering.
Appendix
Table 7: Generation user prompts used for each evaluation dataset. Braced terms denote fields from each VLMEvalKit example, and [Options] denotes the formatted list of answer choices.
Figure 8: Oracle visual-fact intervention on Geometry3K. Top panels report accuracy with the original image alone ( w/o visual facts ) and with additional oracle visual facts derived from diagram annotations ( w/ visual facts ); bottom panels show the corresponding accuracy gap in percentage points. Compared with standard OPD, S-OPD improves accuracy without visual facts and reduces the accuracy gap from 7.8 to 2.3 points for the 2B student and from 15.1 to 12.1 points for the 4B student.
Benchmark
Trained on Geometry3K
Trained on ViRL39K
2B
4B
2B
4B
OPD
S-OPD
OPD
S-OPD
OPD
S-OPD
OPD
S-OPD
V ∗ Bench
73.30
74.35
74.87
74.35
74.87
75.92
62.67
63.30
ZoomBench
41.42
42.37
43.55
44.26
41.30
40.95
43.67
44.14
HallusionBench
55.73
56.26
64.77
64.98
52.58
53.00
74.35
74.35
Appendix
Table 8: Comparison of OPD and S-OPD across different student model sizes and training datasets. All models are evaluated on three out-of-domain visual benchmarks.
In-domain
Mathematical Reasoning
Logical Reasoning
General Reasoning
Overall
Method
Geo3K-Test
MVerse
MVista
MVision
WeMath
LogicVista
VisPuzzles
ZeroBench
MMMU-val
Avg.
Qwen3-VL-2B-Instruct
Vanilla OPD
40.93
49.28
63.3
34.96
36.38
39.83
11.82
11.68
49.11
37.48
w/ TPC
42.56
49.84
65.2
35.07
38.29
40.71
11.39
11.98
47.56
38.07
w/ PA
41.93
50.69
65.3
34.36
36.76
43.14
12.67
14.97
47.56
38.60
S-OPD
44.26
50.76
65.3
35.43
38.05
42.51
13.53
13.17
48.67
39.08
Appendix
Table 9: Detailed component analysis results on Geometry3K. The Geometry3K test split is used for in-domain performance evaluation. Bold and underlined numbers indicate the best and second-best results within each student size, respectively.
Value
Mathematical
Logical
General
Avg.
λTPC
0.005
47.07
25.09
29.28
37.13
0.01
46.20
25.05
28.71
36.54
0.02
47.04
25.16
30.60
37.46
λPA
0.01
46.41
24.37
28.57
36.44
Appendix
Table 10: Ablation of the TPC and PA coefficients on ViRL39K.
Figure 9: Token-level log-probability analysis of 2B and 4B students on ViRL39K and Geometry3K during training. We report the teacher–student token-level log-probability gap under different teacher–student configurations.
ViRL39K, 2B
Geometry3K, 4B
Metric
OPD
S-OPD
OPD
S-OPD
Tokens/response
2235.9
2233.2
802.4
800.3
Time/step (s)
189.44
243.56
92.26
107.06
Gen. (ms/token)
0.233
0.252
0.315
0.333
Overhead
–
+28.6%
–
+16.0%
Appendix
Table 11: Computational overhead of S-OPD. Both methods use four NVIDIA A100 80GB GPUs, two for the student and two for the teacher. Training time excludes initialization, checkpoint saving, and validation. Overhead is relative to OPD in time per step.
Figure 10: Qualitative examples from MathVerse, LogicVista, and MMMU-val where 2B S-OPD correctly grounds its reasoning in visual evidence, while 2B OPD produces incorrect answers.
Figure 11: Qualitative visualization of token-level policy contrast and teacher gating across three types of visual reasoning : mathematical geometry (MathVerse), abstract visual logic (LogicVista), and scientific diagram understanding (MMMU-val). In the student policy contrast , orange and blue indicate positive and negative changes in token log-probability between the original and masked visual inputs, respectively, while darker colors denote larger absolute changes. For the teacher gate strength , darker purple indicates a larger positive teacher log-probability gap between the original and masked images. For the gated policy contrast , darker yellow indicates a larger teacher-modulated policy discrepancy. The teacher and student exhibit distinct token-level emphasis across all three tasks, and the gated policy contrast selectively concentrates on key visual elements and decision-relevant reasoning steps, demonstrating the intended effect of our teacher-gating mechanism.