Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting distillation update improves the student. Indeed, privileged information can lead the teacher to solve tasks through shortcuts unavailable to the student, producing supervision poorly matched to the student's current behavior. Consequently, even a higher-performing teacher can provide guidance that degrades student performance. To address this, we analyze how the choice of privileged teacher affects the student's update. We derive a necessary and sufficient condition for the teacher's local distillation update to be a positive multiple of the student's reward gradient. Our analysis suggests that the teacher should not only perform well on the task, but also provide guidance suited to the student's current capabilities. This characterization motivates a practical teacher-training surrogate that combines outcome rewards with token-level Kullback-Leibler (KL) regularization toward the student. Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation. Across mathematical reasoning, coding, tool use, and terminal use, JOLT improves training efficiency and performance, with further gains from student rewards.
Figures & tables
Figure 1: Joint On-Policy Learning and Teaching (JOLT) trains a shared model as both a privileged teacher and an unprivileged student. The teacher learns from outcome rewards with KL regularization toward the student, while the student receives teacher supervision along its own generations. (Left): the joint teacher–student training process. (Right): pseudocode for a training step.
Figure 2: Training accuracy (top), teacher–student KL (middle), and final evaluation (bottom) on GSM8K, MATH, and LCBv6. JOLT improves mean evaluation accuracy over the distillation baselines with closer teacher–student agreement. Bottom-row dots show runs, and error bars show standard deviations. Evaluation protocols appear in Appendix A .
Method
MATH
LCBv6
AppWorld
GRPO
68.36±3.73
46.13±1.83
62.20±2.83
OPSD
61.20±2.10
45.34±1.50
29.49±2.48
π -Distill
42.25±1.10
38.23±1.72
59.85±2.97
JOLT
72.14±0.90
47.64±1.10
51.92±5.30
JOLT+
72.59±1.17
49.08±0.41
72.09±6.01
Table 1: Evaluation (%): MATH/LCBv6 accuracy and AppWorld development-set state-test pass rate. Mean ± standard deviation across runs. Run counts and protocols are in Appendix A .
Figure 3: Training performance on AppWorld. JOLT+ reaches GRPO’s final training performance using approximately 1.8× fewer updates, then continues improving. Combining teacher supervision with student outcome rewards outperforms both distillation baselines.
Figure 4: Performance scaling with training token budget, as measured in completion tokens (left) and processed training tokens (right). On MATH, JOLT reaches high accuracy with fewer tokens than GRPO while surpassing OPSD’s early plateau under both token counters. Appendix B details token accounting, curve fitting, and the budget comparison.
Figure 5: Teacher–student log-probability differences along responses to the same GSM8K problem under OPSD (top) and JOLT (bottom). Red indicates higher teacher log probability, while blue indicates lower teacher log probability. In this example, OPSD exhibits disagreement throughout much of its response, while JOLT shows smaller differences concentrated at fewer token positions. Appendix D extends the comparison to π -Distill and JOLT+.
Method
pass@1
pass@4
pass@16
pass@32
Base
5.18
10.84
18.35
22.81
GRPO
8.25
17.04
26.91
30.34
OPSD
2.82
6.96
13.59
16.90
π -Distill
4.99
10.50
18.30
22.31
JOLT
6.56
13.56
22.06
25.69
JOLT+
10.96
20.12
29.99
34.83
Table 2: Terminal-Bench 2.1 pass@ k (%) after TMax training, using one training run per method.
Figure 6: Empirical squared bias (left), variance (middle), and mean squared error (right) of weighted student gradients along each method’s own training trajectory. Bias and MSE use Monte Carlo reward-gradient references. GRPO’s error is dominated by variance, whereas OPSD develops increasing empirical bias. JOLT and JOLT+ retain low variance with smaller empirical bias than OPSD. Bands show pointwise 95% batch-bootstrap intervals. See Appendix E for details.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Hyperparameter
GSM8K
MATH
LCBv6
AppWorld
TMax
GRPO
Learning Rate
1×10−6
3×10−6
1×10−6
5×10−7
1×10−6
Batch Size B
8
8
32
4
8
Group Size G
8
8
8
4
32
Num Steps
200
150
150
300
50
OPSD
Learning Rate
1×10−7
2×10−6
1×10−6
1×10−7
2.5×10−7
Batch Size B
8
8
32
4
8
Appendix
Table 3: Selected training hyperparameters. B denotes the number of problems per update and G the group size per active role. On GSM8K and MATH, JOLT/JOLT+ use four teacher and four student rollouts per problem, while GRPO and OPSD use eight student rollouts. Learning rates are initial values where a schedule is used.
Figure 7: Evaluation curves on GSM8K, MATH, LCBv6, and AppWorld (top), with training-accuracy summaries for the first three tasks and final AppWorld evaluation below. Final evaluation summaries for the first three tasks appear in Figure 2 . LCBv6 evaluates training problems on complete test suites; AppWorld reports development-set state-test pass rate. Bands and error bars show standard deviations across runs; dots show individual runs.
Figure 8: Hyperparameter sensitivity on GSM8K, covering JOLT configurations (left) and learning rates across methods (right).
Figure 9: GSM8K accuracy across teacher–student coupling strengths for (a) JOLT and (b) JOLT+.
Setting
JOLT+ Runs
GRPO Runs
Probability of Improvement (%)
GSM8K
5
4
90.0
MATH
3
3
88.9
LCBv6
3
3
100.0
AppWorld
3
3
100.0
Equal-weight average
—
—
94.7
Appendix
Table 4: Empirical probability that a JOLT+ run outperforms a GRPO run at the final checkpoint, using the evaluations in Figures 2 and 7 . The average gives equal weight to each setting.
Token counter
Method
a∞ (%)
B50
k
RMSE
Completion
GRPO
70.044
4.173
1.655
0.994
OPSD
61.706
0.512
2.076
0.938
JOLT
72.123
0.440
2.413
1.126
Processed
GRPO
70.441
4.343
1.501
0.959
OPSD
61.691
1.314
2.270
0.937
JOLT
72.125
0.955
2.634
1.152
Appendix
Table 5: Fitted accuracy–token scaling parameters for Figure 4 . Budgets are in millions of tokens; RMSE is in percentage points.
Token counter
GRPO
JOLT
GRPO/JOLT
Completion
8.04
0.617
13.0×
Processed
8.40
1.30
6.5×
Appendix
Table 6: Token budgets at 63.44% MATH accuracy. GRPO uses the observed checkpoint near its learning-curve knee; JOLT uses its fitted curve. Budgets are in millions of tokens, and values are approximate.
Target
Completion tokens
Processed training tokens
accuracy
GRPO
JOLT
Ratio
GRPO
JOLT
Ratio
55%
3.48
0.368
9.5×
3.50
0.811
4.3×
60%
5.56
0.496
11.2×
5.81
1.065
5.5×
65%
9.92
0.691
14.4×
10.73
1.443
7.4×
Appendix
Table 7: Estimated token budgets and GRPO/JOLT budget ratios at multiple MATH accuracy targets. Both methods’ budgets are obtained from their fitted curves. Budgets are in millions of tokens; ratios are calculated before rounding.
Configuration
Student outcome reward
Teacher outcome coefficient α
Teacher KL coefficient β
Tuned LR
Held-out accuracy (%)
OPSD
–
–
–
2×10−6
62.95±0.78
OPSD + Student Reward
✓
–
–
2×10−6
66.52±0.73
JOLT (teacher KL only)
–
–
10−6
10−6
65.39±0.97
JOLT (teacher outcome only)
–
1
–
2×10−6
67.87±0.86
JOLT (full objective)
–
1
10−3
2×10−6
72.10±0.82
JOLT+ (full objective)
✓
1
10−3
2×10−6
73.80±0.74
Appendix
Table 8: Held-out MATH accuracy at step 75. OPSD, OPSD + Student Reward, and the two teacher-objective variants report means across five shared training seeds, with 95% confidence-interval half-widths. OPSD + Student Reward disables both teacher-objective terms in JOLT+. JOLT and JOLT+ use the step-75 evaluations in Figure 7 , each with three runs and a 95% confidence interval; their configurations are given in Table 3 . JOLT+’s interval is a Student- t interval computed from the plotted standard deviation. Dashes indicate absent objective terms.
Figure 10: Learning curves on MATH for OPSD and JOLT variants with either teacher KL regularization or teacher outcome rewards. Both teacher-training variants achieve higher accuracy than OPSD later in training. Lines and shaded regions show the mean and 95% confidence intervals across five training seeds.
Figure 11: Token-level teacher–student alignment on the same problem. Red indicates higher teacher log probability, while blue indicates lower teacher log probability. JOLT and JOLT+ show smaller cross-context log-probability differences than OPSD and π -Distill.
Figure 12: Student-gradient directions at fixed checkpoints. Distillation-only and combined estimates show lower directional dispersion (top) and higher mean cosine similarity to the Monte Carlo reward-gradient reference (bottom) than reward-only estimates. Columns follow different training methods; within each checkpoint, estimators share model weights, problems, sampled student trajectories, and the reference. Bands show pointwise 95% whole-batch bootstrap intervals.
Setting
Search Coverage
GSM8K
Five learning rates evaluated over the full training horizon, followed by repeated-seed comparisons.
MATH
We examine 42 configurations spanning learning rates, teacher–student mixing weights, teacher KL coefficients, and scheduling choices, including teacher-only controls. Selected configurations are evaluated across additional random seeds.
LCBv6
Eighteen distinct configurations spanning eight learning rates, five teacher KL coefficients, and three teacher mixing weights, followed by fresh-seed confirmation.
AppWorld
Five full-horizon learning-rate trials and additional replications of the selected configuration.
Appendix
Table 9: Search coverage for π -Distill hyperparameter tuning.
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.
Langlin Huang, Hao Liu, Mononito Goswami +5
Washington University in St. Louis · AWS AI Labs · Carnegie Mellon University +1
On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base \texttt{Qwen3-1.7B}, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.
Yi Ding, Ruqi Zhang
Department of Computer Science, Purdue University, USA
On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. However, its reliance on a stronger external teacher has driven recent work on on-policy self-distillation, where the same model serves as both teacher and student under different prompt contexts. Yet, existing self-distillation methods largely reduce learning to KL matching toward the context-augmented teacher model. This approach often suffers from training instability and can degrade reasoning performance over time. Moreover, self-distillation from the same model with prompt augmentation lacks the exploratory diversity provided by a genuine external teacher. To address these limitations, we move beyond fixed-teacher KL matching and propose \textbf{P}reference-\textbf{B}ased \textbf{S}elf-\textbf{D}istillation (\textbf{PBSD}), which revisits on-policy self-distillation through a reward-regularized perspective. Instead of directly matching the teacher distribution, we derive a reward-regularized objective whose analytic optimum is a reward-reweighted teacher distribution, yielding a target policy provably superior to the original teacher under this objective. Practically, PBSD optimizes preference gaps between teacher and student samples while maintaining on-policy student sampling. We support this framework with a statistical analysis of the induced preference-learning problem, formally establishing when on policy self-distillation is preferable to learning from an external teacher in our setting. Experiments on mathematical reasoning and tool-use benchmarks across multiple model scales demonstrate that PBSD consistently achieves the strongest average performance among comparable baselines, showing improved training stability over prior self-distillation baselines while preserving token efficiency.