Organizations: National Engineering Research Center of Software Engineering, Peking University, Beijing, China · School of Computer Science, Peking University, Beijing, China · Key Laboratory of High Confidence Software Technologies, Ministry of Education, Beijing, China · Center on Frontiers of Computing Studies, Peking University, Beijing, China
Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a scalar-weighted teacher KL term to the policy-gradient objective, providing dense token-level guidance that may be unreliable at some positions. Despite the benefits of combining these signals, their interaction during optimization can destabilize joint training. To understand how this instability develops, we study the learning dynamics of hybrid reward--distillation training through a neural tangent kernel (NTK) analysis. We introduce the cross-signal NTK KDR(n), a token-level statistic that measures the alignment between reward and distillation gradients at position n. Through this analysis, we identify two failure modes: 1 Magnitude drowning, where the reward gradient exceeds the distillation gradient by orders of magnitude, so that even weak directional conflict can cause the distillation loss to rise despite its explicit inclusion in the training objective; and 2 Localized directional conflict, where the sequence-level advantage and the teacher's position-specific distribution induce opposing updates at the same token (KDR(n)<0). The severity of these effects depends on the optimization regime: the gradient-norm ratio κ=∥∇LR∥/∥∇LD∥ varies by roughly an order of magnitude across tasks, and our experiments reveal an empirical threshold beyond which naive mixing can lead to persistent training collapse. Motivated by these findings, we introduce the M3 family, which combines magnitude normalization with three strategies...
Figures & tables
Figure 1: Two failures of linear reward–distillation mixing (Qwen3-1.7B, GSM8K). (a) Distillation loss rises under OPSD, indicating magnitude drowning. (b) Positive and negative KDR(n) interleave in a correct rollout, revealing token-local conflict hidden by aggregation. (c) Baselines collapse by step 500 , while M3-Select and M3-Soft remain stable.
Figure 2: From interaction analysis to hybrid updates. Shared parameters couple dense teacher and sparse reward supervision. M3 controls residual scale and allocates teacher influence by token-level compatibility; its extragradient extension separates teacher shaping from reward correction.
Figure 3: GradNorm challenge on InternLM2.5-1.8B GSM8K: training reward (left), gradient cosine (center), and learned distillation weight (right).
Qwen3-1.7B
Qwen2.5-1.5B
InternLM2.5-1.8B
Llama-3.2-1B
Method
GSM8K
SVAMP
ARC
GSM8K
SVAMP
ARC
GSM8K
SVAMP
ARC
GSM8K
SVAMP
ARC
Reference: pure RL
GRPO
0.696 0.0
0.940 0.0
0.732 0.0
0.726 0.0
0.825 0.0
0.689 0.0
0.404 0.0
0.670 0.0
0.599 0.0
0.552 0.0
0.730 0.0
0.533 0.0
Teacher-augmented baselines
Hybrid (OPSD)
0.786 ↑ +9.0
0.932 ↓ -0.8
0.796 ↑ +6.4
0.717 ↓ -0.9
0.857 ↑ +3.2
0.693 ↑ +0.4
0.492 ↑ +8.8
0.608 ↓ -6.2
0.604 ↑ +0.5
0.563 ↑ +1.1
0.722 ↓ -0.8
0.544 ↑ +1.1
OPSD+GradNorm
0.765 ↑ +6.9
0.917 ↓ -2.3
0.752 ↑ +2.0
0.735 ↑ +0.9
0.837 ↑ +1.2
0.695 ↑ +0.6
0.459 ↑ +5.5
0.687 ↑ +1.7
0.590 ↓ -0.9
0.547 ↓ -0.5
0.717 ↓ -1.3
0.370 ↓ -16.3
Table 1: Cell reports accuracy (top) and its percentage-point change from GRPO (bottom; ↑/↓ ); † Hard masking is vacuous for single-token ARC.
Reference and adaptive baselines
Ours : boundary-gated mixing
Architecture
GRPO
Hybrid
GradNorm
M3-Select
M3-Soft
Qwen3-1.7B
↓0.003
↓0.000
↓0.000
0.809
0.866
Qwen3-0.6B
↓0.003
↓0.000
↓0.053
0.644
0.673
Llama-1B
↓0.004
↓0.003
↓0.328
0.497
0.485
Table 2: Long-horizon reward on GSM8K after 500 steps (last-50 mean). ↓ denotes collapse.
Figure 6
Region
OPSD
GRPO
Hybrid
Compatible ( {\color[rgb]{0,0,1}K_{DR}}>\epsilon , n =851)
+1.6×10−3
+2.7×10−3
+2.0×10−3
Neutral ( |{\color[rgb]{0,0,1}K_{DR}}|\leq\epsilon , n =5191)
≈0
≈0
≈0
Conflicting ( {\color[rgb]{0,0,1}K_{DR}}<-\epsilon , n =2450)
+0.3×10−4
+7.3×10−4
−2.4×10−4
Table 3: One-step sampled-token log-probability change by pre-update conflict region (30 trajectories, 8,492 tokens). Negative values indicate suppression by the update.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Hyperparameter
Value
LoRA
Rank r
64
LoRA
Scaling αLoRA
128
Distillation
Divergence
Generalized JSD
Distillation
JSD coefficient λ
21
Distillation
Token clipping τ
0.05
Reinforcement learning
Rollouts per group G
8
Appendix
Table 4: Default experimental hyperparameters.
Best non-M3 baseline
M3-Soft
M3-Soft
Architecture
GSM8K
SVAMP
ARC
GSM8K
SVAMP
ARC
≥ Baseline
Qwen3-1.7B
0.8317
0.940
0.796
0.8324
0.945
0.807
3/3
Qwen2.5-1.5B
0.735
0.860
0.695
0.752
0.905
0.736
3/3
InternLM2.5-1.8B
0.492
0.737
0.604
0.501
0.777
0.608
3/3
Llama-3.2-1B
0.568
0.730
0.545
0.568
0.750
0.556
3/3
Total: 11 strict wins and 1 exact tie
12/12
Appendix
Table 5: Accuracy of M3-Soft versus the strongest non-M3 baseline.
Architecture
Dataset
GRPO
OPSD
OPSD+GradNorm
M3-Select
M3-Soft
\bar{{\color[rgb]{1,0.5508,0}\kappa}}
Qwen3-1.7B
GSM8K
0.838
0.825
0.731
0.844
0.881
3,404
SVAMP
0.922
0.894
0.894
0.947
0.969
5,080
ARC
0.741
0.919
0.769
0.263
0.775∗
2,465
Qwen2.5-1.5B
GSM8K
0.806
0.747
0.653
0.153
0.828∗
2,467
SVAMP
0.769
0.766
0.681
0.331
0.778∗
2,234
ARC
0.659
0.650
0.644
0.319
0.700
2,184
Appendix
Table 6: Per-cell training reward (last-20 mean) and mean gradient-magnitude ratio κˉ . M3-Soft reports the selected evaluated configuration; ∗ denotes extended training, and κˉ is measured on the corresponding naive-Hybrid run.
Method
MATH
SVAMP
ARC-Challenge
Reference and adaptive baselines
GRPO
↓0.018
0.900
0.773
Hybrid (OPSD)
0.448
↓0.170
0.765
OPSD+GradNorm
↓0.015
0.973
0.790
Ours : boundary-gated mixing (M3)
M3-Select
0.367
0.943
0.282
Appendix
Table 7: Long-horizon reward across datasets on Qwen3-1.7B after 500 steps (last-50 mean). M3-Soft uses the best evaluated gate configuration per dataset; other methods use the reference configuration. Bold marks the column maximum, and ↓ denotes collapse.
Method
1–100
101–200
201–300
301–400
401–500
Full
Trend
Pure GRPO
0.840
0.801
0.809
0.802
0.075
0.665
collapse @412
Hybrid (OPSD, α=0.5 )
0.859
0.762
0.114
0.006
0.001
0.348
collapse @221
OPSD+GradNorm †
0.877
0.823
0.868
0.464
0.000
0.606
collapse @358
RLSD ( Yang et al., 2026 )
0.855
0.874
0.871
0.259
0.005
0.573
collapse @362
SDPO ( Hübotter et al., 2026 )
0.306
0.000
0.000
0.000
0.000
0.061
collapse @50
M3-Select ( αmax=0.15 )
0.846
0.818
0.853
0.846
0.810
0.835
stable
Appendix
Table 8: Phase-wise training reward over 500 steps on Qwen3-1.7B GSM8K. “Full” is the all-step mean, “Trend” gives the first collapse step, and † denotes batch size 1 .
Figure 6: Runtime token-level conflict on Qwen3-0.6B GSM8K. Hybrid’s post-collapse drop reflects a vanishing RL gradient; M3-Select remains active near {\color[rgb]{0,0,1}C_{\mathrm{NTK}}}=0.3 .
Method
β
Last-10
Last-20
αeff
GRPO
—
0.858
0.850
—
OPSD (uniform)
—
0.833
0.825
0.500
M3-Select (hard)
∞
0.855
0.844
∼0.30
M3-Soft
10
0.875
0.847
0.446
M3-Soft
5
0.813
0.838
0.439
M3-Soft
1
0.906
0.856
0.444
Appendix
Table 9: M3-Soft sharpness sweep on Qwen3-1.7B GSM8K at 100 steps ( αmax=0.5 ). αeff denotes the mean teacher weight. For hard selection, the measured conflict rate is CNTK≈0.40 , so approximately 60% of tokens retain weight αmax , giving αeff≈0.5×0.60=0.30 .
Model
Dataset
β=0.5
β=1
β=5
GRPO
Qwen3-1.7B
GSM8K
0.834
0.856
0.850
0.838
SVAMP
0.991
0.969
0.981
0.953
ARC
0.516
0.350
0.356
0.741
Llama-3.2-1B
GSM8K
0.569
0.647
0.353
↓0.000
SVAMP
0.456
0.822
0.609
0.072
ARC
0.559
0.500
0.388
0.522
Appendix
Table 10: M3-Soft sharpness sensitivity at 100 steps (last-20 mean reward). αmax=0.05 by default and 0.025 on Qwen3-SVAMP.
Mode
Conflict %
cos(gR,gD)
True opposition
{\color[rgb]{1,0.5508,0}\kappa}_{\mathrm{median}}
LˉD
On-policy OPSD
75%
+0.08
25/157 ( 16% )
388
0.098
Off-policy
75%
−0.01
90/157 ( 57% )
68
0.810
Appendix
Table 11: Response-level semantic conflict (Qwen3-0.6B, 208 trajectories): on-policy = teacher evaluates the student’s own trajectory, off-policy = ground-truth contexts; true opposition = wrong trajectories with opposing RL and distillation gradients.
Figure 7: On-policy versus off-policy semantic conflict. On-policy distillation has a higher median {\color[rgb]{1,0.5508,0}\kappa} but fewer incorrect trajectories with genuinely opposed reward and distillation gradients; off-policy supervision reverses this pattern.
Figure 8: Per-position conflict rate {\color[rgb]{0,0,1}C_{\mathrm{NTK}}} across model scales and datasets. The fraction of positions with negative cross-signal NTK remains between 0.37 and 0.46 (mean ≈0.41 , dashed line).
Figure 9: A readable token-level case study for a correct ( A>0 ) and an incorrect ( A<0 ) response to the same prompt. Darker red denotes stronger local conflict ( {\color[rgb]{0,0,1}K_{DR}(n)}<0 ), darker blue denotes stronger compatibility, and pale colors denote weak interaction. Both responses interleave the two signal types.
Figure 10: Three additional paired token-level case studies. Headers report {\color[rgb]{0,0,1}C_{\mathrm{NTK}}} over each full 200 -token response; the first 48 tokens are displayed. The incorrect rollout has the higher conflict rate in each pair ( 49%→52% , 42%→62% , and 42%→45% ), while every rollout contains both compatible and conflicting positions.
Figure 11: Aggregate conflict diagnostics: (a) mean magnitude ratio across settings; (b) parameter-level gradient cosine; and (c) token-level conflict rate versus reward for M 3 -Select.
LoRA Rank
r=8
r=16
r=32
r=64
r=128
Mean Reward
0.826
0.824
0.828
0.844
0.841
Appendix
Table 12: LoRA-rank ablation for M 3 -Soft on Qwen3-1.7B GSM8K over 100 training steps.
Failure pattern
Observable signature
Intervention
Length inflation Magnitude drowning
Large {\color[rgb]{1,0.5508,0}\kappa} ; repetitive reasoning not penalized by the final-answer verifier
On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental. Which teacher model should be used, and in the case of self-distillation, which specific context should serve as the supervisory signal? Does the optimal choice vary from one token to the next? At present, addressing these questions typically requires costly training runs whose aggregate performance metrics obscure the dynamics at the level of individual tokens. We introduce a training-free diagnostic framework that operates at the highest resolution: per token, per question, and per teacher. We derive an ideal per-node gradient defined as the parameter update that maximally increases the student's probability of success. We then develop a scalable targeted-rollout algorithm to estimate this gradient efficiently, even for long chains of intermediate thoughts. The gradient alignment score, defined as the cosine similarity between this ideal gradient and any given distillation gradient, quantifies the extent to which a particular configuration approximates the ideal signal. Across a range of self-distillation settings and external teacher models, we observe that distillation guidance exhibits substantially higher alignment with the ideal on incorrect rollouts than on correct ones, where the student already performs well and the teacher's signal tends to become noisy. Furthermore, we find that the optimal distillation context depends jointly on the student model's capacity and the target task, and that no single universally effective configuration emerges. These findings motivate the use of per-task, per-token diagnostic analyses for distillation.
Mohammadreza Armandpour, Fatih Ilhan, David Harrison +6
Sparse on-policy distillation (OPD) allocates teacher supervision to a small subset of tokens in student-generated trajectories. However, useful teacher guidance can yield a noisy update when its gradient is estimated from a sampled next token. We study this estimation problem at a fixed prefix in information geometry and propose an information-efficiency ratio (IER) based on a signal-to-noise decomposition. IER characterizes relative gradient estimation error under an optimal scalar baseline. A candidate-set approximation enables token selection based on IER and its combination with existing usefulness scores, while retaining the sampled reverse-KL training objective. On mathematical and medical reasoning tasks, adding IER improves existing selectors in multiple settings, with sparse configurations matching or exceeding full OPD without token selection at small token budgets of 0.1%--1%. These results support accounting for both usefulness and gradient-estimation reliability when allocating sparse supervision. Our code is available at https://github.com/BruceSheng1202/IER-OPD.
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.
Wenze Lin, Jiale Zhao, Xitai Jiang +5
LeapLab, Tsinghua University · Qiuzhen College, Tsinghua University · Beihang University +2