Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across training checkpoints and across response groups with different task rewards. Motivated by this observation, we introduce GAD-RL, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions. A frozen teacher conditions on reference transcriptions and student-generated prefixes. GAD-RL disables distillation for response groups containing an output with task reward at least 0.95 and continuously attenuates distillation strength as group-mean reward increases. It also weights forward KL by the student's probability of the teacher's Top-1 token, moderating local auxiliary updates when student support for that candidate is low. On Qwen3.5-2B, GAD-RL achieves 59.92% Micro Recall on CHAOS-Bench, surpassing GRPO and GRPO+OPD (fixed-weight) by 8.45 and 4.43 percentage points, respectively, while achieving an Overall score of 91.18 on OmniDocBench v1.6.
Figures & tables
Figure 1: Student performance improves while directional supervisory SNR declines. The same frozen teacher scores GRPO checkpoints offline on a separate analysis set of 1,000 synthetic pages. (a) Signal composition on perturbed-word-associated tokens: relative to student probabilities at the same prefixes, teacher signals are beneficial when favoring correct emitted tokens or disfavoring errors, and potentially harmful in the reverse direction; small differences are neutral (Appendix B.4 ). (b) Beneficial-to-harmful token-count ratio on a log scale; the dotted line marks equal counts. (c) Student perturbed-word Micro Recall.
Method
CHAOS-Bench
GlitchText
OmniDocBench v1.6
Micro Recall ↑
Ident ↑
Cor ↓
Overall ↑
Student: Qwen3.5-2B
Baseline
0.0402
76.48
14.94
80.06
SFT
0.3763
82.20
10.71
90.29
GRPO
0.5147
91.30
4.21
90.80
GRPO+OPD (fixed-weight)
0.5549
91.31
3.14
90.90
Table 1: Transcription faithfulness and general document parsing, grouped by student model. Micro Recall is a ratio; other scores are percentages. Bold marks the best result within each model block.
Figure 2: GRPO and GAD-RL training dynamics. (a) Mean reward. (b) Fraction of groups with at least one of eight responses achieving perturbed-word Recall of 1. For GAD-RL, (c) the fraction of groups with distillation disabled by the gate and (d) the mean attenuation factor among the remaining active groups.
Figure 3: FKL and SA-FKL training dynamics. With otherwise identical settings: (a) mean training reward (EMA 0.95); (b) CHAOS-Bench Micro Recall, with a step-0 reference of 0.0402.
Table 5
Variant
Recall ↑
GRPO
0.5147
GAD-RL (full)
0.5992
w/o SW
0.5598
w/o MGD
0.5931
w/o RAA
0.5414
Table 4: Component ablations on CHAOS-Bench at λ .
Checkpoint
All tokens
Correct tokens
Teacher errors
% of all
% of correct
Base
4,281
1,029
131
3.06
12.73
Step 100
10,325
10,159
967
9.37
9.52
Step 200
12,903
12,865
4,234
32.81
32.91
Step 300
13,391
13,366
5,731
42.80
42.88
Table 5: Teacher Top-1 errors at correct student outputs. Counts refer to perturbed-word-associated tokens. The final columns give teacher errors as percentages of all eligible tokens and of correct student tokens.
Figure 4: Teacher supervision becomes less favorable as group reward increases. GRPO+OPD groups at steps 50 and 200 are binned by mean reward Rˉ . For perturbed-word-associated tokens: (a) beneficial-to-harmful count ratio (log scale; dotted line: equal counts); (b) beneficial and harmful fractions, with the remainder neutral. The absolute probability-change margin is 0.1. n gives group counts at step 50 / 200; error bars show 95% input-group-bootstrap confidence intervals.
Method
Coefficient 1
Recall ↑
GRPO
0
0.5147
+ Linear-decay OPD
λs(u)
0.5693
+ Low-score OPD
λ
0.5775
GAD-RL
λgfκ(Rˉ)
0.5992
10λgfκ(Rˉ)
0.5882
Table 6: Distillation decay and filtering strategies. The best Recall is bold.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Example of a synthetically perturbed document page. Boxes highlight character-level substitutions in “noisy”, “model”, “robustness”, and “strongest”.
Figure 6: Teacher signals on erroneous tokens. For the same emitted token, Δp<−0.1 is beneficial suppression and Δp>0.1 is harmful reinforcement. Panel (b) uses an enlarged horizontal scale.
Method
Teacher
Recall ↑
GRPO
None
0.5147
GAD-RL
Self-teacher (2B)
0.5441
GAD-RL
SFT Qwen3.5-2B
0.5992
Appendix
Table 7: Teacher-choice ablation for GAD-RL on CHAOS-Bench.
Student
N
Beneficial
Harmful
Neutral
B/H
Recall (%)
Base
4,281
3,732
297
252
12.566
7.18
Step 100
10,325
2,270
3,443
4,612
0.659
94.16
Step 200
12,903
1,261
8,908
2,734
0.142
97.78
Step 300
13,391
1,175
10,273
1,943
0.114
98.06
Appendix
Table 8: Fixed-teacher signal counts and perturbed-word Recall on the analysis set. Signal statistics use eligible emitted tokens; Recall uses all 3,564 annotated words, including omissions.
Student
N
GT-compatible
Incompatible
Unresolved
Base
294
136 (46.26)
131 (44.56)
27 (9.18)
Step 100
3,442
2,413 (70.10)
967 (28.09)
62 (1.80)
Step 200
8,908
4,615 (51.81)
4,234 (47.53)
59 (0.66)
Step 300
10,273
4,458 (43.40)
5,731 (55.79)
84 (0.82)
Appendix
Table 9: Teacher Top-1 at correct perturbed-word-associated tokens with Δp<−0.1 . Entries are counts (percent of N ), including unresolved cases in the denominator.
Setting
Value
Student / teacher backbone
Qwen3.5-2B or Qwen3-VL-2B; matched within each pair
Initialization
Student: base checkpoint; teacher: transcription SFT from the same base
Teacher SFT data
200,000 transcription-task examples
Mixed dataset
14,400 samples (perturbed + general)
Mixture (general:perturbed)
3:2
Learning rate / duration
10−6 / 300 steps
Appendix
Table 10: Shared GAD-RL training configuration for Qwen3.5-2B and Qwen3-VL-2B. Each student uses a teacher trained from the same backbone.
Figure 7: Beneficial and harmful teacher signals. Panels show the page, enlarged perturbed-word region, GT, student response, and token probabilities. Yellow rows mark focal tokens; · denotes a space. Student and teacher score the same emitted token at the same prefix, conditioned on image and privileged text, respectively.
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.
Wenze Lin, Jiale Zhao, Xitai Jiang +5
LeapLab, Tsinghua University · Qiuzhen College, Tsinghua University · Beihang University +2
On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, \textbf{Supervision Fidelity Decay (SFD)}: as student-generated prefixes lengthen, the teacher's next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce \textbf{Lookahead Group Reward (\ours{})}. Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, \ours{} evaluates the student's top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, \ours{} improves mean@8 by \textbf{2.57} points over OPD for a 7B student, with gains increasing in longer-generation and reaching +\textbf{4.92} points on AIME-26 at 39k tokens.
Yanjiang Liu, Jie Lou, Xinyan Guan +7
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu
On-policy distillation is a powerful way to transfer reasoning ability from a strong teacher to a smaller student: the student samples trajectories from its own policy, and the teacher provides dense token-level supervision on the states the student actually visits. However, this supervision is not always reliable: a teacher can assign high likelihood to plausible but incorrect solutions, or low likelihood to correct student solutions that follow different reasoning paths. Unconditionally distilling the teacher can therefore reinforce bad modes or erase useful student behavior. To address these limitations, we introduce RG-OPD: Reward-Gated On-Policy Distillation that uses verifier feedback to decide when teacher logits should be trusted. RG-OPD bridges sparse verifier rewards and dense teacher logits, preserving token-level supervision while filtering misleading teacher signals. Across reasoning and coding benchmarks, RG-OPD produces stronger distilled students, outperforming both vanilla reverse-KL distillation and the recent TSD-KD baseline. At 1K generation length, RG-OPD improves over reverse-KL by 2.9 points and over TSD-KD by 4.9 points; in the long-generation setting, it improves over the untuned student by 8.2 points. Our code is available at https://github.com/UoC-tail/RG-OPD.
Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi +3