Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting distillation update improves the student. Indeed, privileged information can lead the teacher to solve tasks through shortcuts unavailable to the student, producing supervision poorly matched to the student's current behavior. Consequently, even a higher-performing teacher can provide guidance that degrades student performance. To address this, we analyze how the choice of privileged teacher affects the student's update. We derive a necessary and sufficient condition for the teacher's local distillation update to be a positive multiple of the student's reward gradient. Our analysis suggests that the teacher should not only perform well on the task, but also provide guidance suited to the student's current capabilities. This characterization motivates a practical teacher-training surrogate that combines outcome rewards with token-level Kullback-Leibler (KL) regularization toward the student. Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation. Across mathematical reasoning, coding, tool use, and terminal use, JOLT improves training efficiency and performance, with further gains from student rewards.
Figures & tables
Figure 1: Joint On-Policy Learning and Teaching (JOLT) trains a shared model as both a privileged teacher and an unprivileged student. The teacher learns from outcome rewards with KL regularization toward the student, while the student receives teacher supervision along its own generations. (Left): the joint teacher–student training process. (Right): pseudocode for a training step.
Figure 2: Training accuracy (top), teacher–student KL (middle), and final evaluation (bottom) on GSM8K, MATH, and LCBv6. JOLT improves mean evaluation accuracy over the distillation baselines with closer teacher–student agreement. Bottom-row dots show runs, and error bars show standard deviations. Evaluation protocols appear in Appendix A .
Method
MATH
LCBv6
AppWorld
GRPO
68.36±3.73
46.13±1.83
62.20±2.83
OPSD
61.20±2.10
45.34±1.50
29.49±2.48
π -Distill
42.25±1.10
38.23±1.72
59.85±2.97
JOLT
72.14±0.90
47.64±1.10
51.92±5.30
JOLT+
72.59±1.17
49.08±0.41
72.09±6.01
Table 1: Evaluation (%): MATH/LCBv6 accuracy and AppWorld development-set state-test pass rate. Mean ± standard deviation across runs. Run counts and protocols are in Appendix A .
Figure 3: Training performance on AppWorld. JOLT+ reaches GRPO’s final training performance using approximately 1.8× fewer updates, then continues improving. Combining teacher supervision with student outcome rewards outperforms both distillation baselines.
Figure 4: Performance scaling with training token budget, as measured in completion tokens (left) and processed training tokens (right). On MATH, JOLT reaches high accuracy with fewer tokens than GRPO while surpassing OPSD’s early plateau under both token counters. Appendix B details token accounting, curve fitting, and the budget comparison.
Figure 5: Teacher–student log-probability differences along responses to the same GSM8K problem under OPSD (top) and JOLT (bottom). Red indicates higher teacher log probability, while blue indicates lower teacher log probability. In this example, OPSD exhibits disagreement throughout much of its response, while JOLT shows smaller differences concentrated at fewer token positions. Appendix D extends the comparison to π -Distill and JOLT+.
Method
pass@1
pass@4
pass@16
pass@32
Base
5.18
10.84
18.35
22.81
GRPO
8.25
17.04
26.91
30.34
OPSD
2.82
6.96
13.59
16.90
π -Distill
4.99
10.50
18.30
22.31
JOLT
6.56
13.56
22.06
25.69
JOLT+
10.96
20.12
29.99
34.83
Table 2: Terminal-Bench 2.1 pass@ k (%) after TMax training, using one training run per method.
Figure 6: Empirical squared bias (left), variance (middle), and mean squared error (right) of weighted student gradients along each method’s own training trajectory. Bias and MSE use Monte Carlo reward-gradient references. GRPO’s error is dominated by variance, whereas OPSD develops increasing empirical bias. JOLT and JOLT+ retain low variance with smaller empirical bias than OPSD. Bands show pointwise 95% batch-bootstrap intervals. See Appendix E for details.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Hyperparameter
GSM8K
MATH
LCBv6
AppWorld
TMax
GRPO
Learning Rate
1×10−6
3×10−6
1×10−6
5×10−7
1×10−6
Batch Size B
8
8
32
4
8
Group Size G
8
8
8
4
32
Num Steps
200
150
150
300
50
OPSD
Learning Rate
1×10−7
2×10−6
1×10−6
1×10−7
2.5×10−7
Batch Size B
8
8
32
4
8
Appendix
Table 3: Selected training hyperparameters. B denotes the number of problems per update and G the group size per active role. On GSM8K and MATH, JOLT/JOLT+ use four teacher and four student rollouts per problem, while GRPO and OPSD use eight student rollouts. Learning rates are initial values where a schedule is used.
Figure 7: Evaluation curves on GSM8K, MATH, LCBv6, and AppWorld (top), with training-accuracy summaries for the first three tasks and final AppWorld evaluation below. Final evaluation summaries for the first three tasks appear in Figure 2 . LCBv6 evaluates training problems on complete test suites; AppWorld reports development-set state-test pass rate. Bands and error bars show standard deviations across runs; dots show individual runs.
Figure 8: Hyperparameter sensitivity on GSM8K, covering JOLT configurations (left) and learning rates across methods (right).
Figure 9: GSM8K accuracy across teacher–student coupling strengths for (a) JOLT and (b) JOLT+.
Setting
JOLT+ Runs
GRPO Runs
Probability of Improvement (%)
GSM8K
5
4
90.0
MATH
3
3
88.9
LCBv6
3
3
100.0
AppWorld
3
3
100.0
Equal-weight average
—
—
94.7
Appendix
Table 4: Empirical probability that a JOLT+ run outperforms a GRPO run at the final checkpoint, using the evaluations in Figures 2 and 7 . The average gives equal weight to each setting.
Token counter
Method
a∞ (%)
B50
k
RMSE
Completion
GRPO
70.044
4.173
1.655
0.994
OPSD
61.706
0.512
2.076
0.938
JOLT
72.123
0.440
2.413
1.126
Processed
GRPO
70.441
4.343
1.501
0.959
OPSD
61.691
1.314
2.270
0.937
JOLT
72.125
0.955
2.634
1.152
Appendix
Table 5: Fitted accuracy–token scaling parameters for Figure 4 . Budgets are in millions of tokens; RMSE is in percentage points.
Token counter
GRPO
JOLT
GRPO/JOLT
Completion
8.04
0.617
13.0×
Processed
8.40
1.30
6.5×
Appendix
Table 6: Token budgets at 63.44% MATH accuracy. GRPO uses the observed checkpoint near its learning-curve knee; JOLT uses its fitted curve. Budgets are in millions of tokens, and values are approximate.
Target
Completion tokens
Processed training tokens
accuracy
GRPO
JOLT
Ratio
GRPO
JOLT
Ratio
55%
3.48
0.368
9.5×
3.50
0.811
4.3×
60%
5.56
0.496
11.2×
5.81
1.065
5.5×
65%
9.92
0.691
14.4×
10.73
1.443
7.4×
Appendix
Table 7: Estimated token budgets and GRPO/JOLT budget ratios at multiple MATH accuracy targets. Both methods’ budgets are obtained from their fitted curves. Budgets are in millions of tokens; ratios are calculated before rounding.
Configuration
Student outcome reward
Teacher outcome coefficient α
Teacher KL coefficient β
Tuned LR
Held-out accuracy (%)
OPSD
–
–
–
2×10−6
62.95±0.78
OPSD + Student Reward
✓
–
–
2×10−6
66.52±0.73
JOLT (teacher KL only)
–
–
10−6
10−6
65.39±0.97
JOLT (teacher outcome only)
–
1
–
2×10−6
67.87±0.86
JOLT (full objective)
–
1
10−3
2×10−6
72.10±0.82
JOLT+ (full objective)
✓
1
10−3
2×10−6
73.80±0.74
Appendix
Table 8: Held-out MATH accuracy at step 75. OPSD, OPSD + Student Reward, and the two teacher-objective variants report means across five shared training seeds, with 95% confidence-interval half-widths. OPSD + Student Reward disables both teacher-objective terms in JOLT+. JOLT and JOLT+ use the step-75 evaluations in Figure 7 , each with three runs and a 95% confidence interval; their configurations are given in Table 3 . JOLT+’s interval is a Student- t interval computed from the plotted standard deviation. Dashes indicate absent objective terms.
Figure 10: Learning curves on MATH for OPSD and JOLT variants with either teacher KL regularization or teacher outcome rewards. Both teacher-training variants achieve higher accuracy than OPSD later in training. Lines and shaded regions show the mean and 95% confidence intervals across five training seeds.
Figure 11: Token-level teacher–student alignment on the same problem. Red indicates higher teacher log probability, while blue indicates lower teacher log probability. JOLT and JOLT+ show smaller cross-context log-probability differences than OPSD and π -Distill.
Figure 12: Student-gradient directions at fixed checkpoints. Distillation-only and combined estimates show lower directional dispersion (top) and higher mean cosine similarity to the Monte Carlo reward-gradient reference (bottom) than reward-only estimates. Columns follow different training methods; within each checkpoint, estimators share model weights, problems, sampled student trajectories, and the reference. Bands show pointwise 95% whole-batch bootstrap intervals.
Setting
Search Coverage
GSM8K
Five learning rates evaluated over the full training horizon, followed by repeated-seed comparisons.
MATH
We examine 42 configurations spanning learning rates, teacher–student mixing weights, teacher KL coefficients, and scheduling choices, including teacher-only controls. Selected configurations are evaluated across additional random seeds.
LCBv6
Eighteen distinct configurations spanning eight learning rates, five teacher KL coefficients, and three teacher mixing weights, followed by fresh-seed confirmation.
AppWorld
Five full-horizon learning-rate trials and additional replications of the selected configuration.
Appendix
Table 9: Search coverage for π -Distill hyperparameter tuning.