A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises each complete answer's probability to a power above one and renormalizes, shifting probability toward answers the model finds most likely (sharpening). Sampling from it improves reasoning without changing parameters, but needs many scored candidates per query. We show that a model can instead be trained to produce such answers in one generation. On-policy power distillation (OPPD) runs a sequential Monte Carlo sampler in which the model being trained generates candidates and a frozen teacher's power distribution weights them; the same probabilities weight each answer in a maximum-likelihood update. Training raises single-generation accuracy by up to 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature, and one generation scores 2.4 and 3.5 points above published power sampling with 64 candidates, recovering 94 percent of the gain that 16 candidates give the untrained model. For context, against GRPO trained with verified rewards from the same checkpoint and budget, OPPD scores 3.8, 4.0 and 5.4 points higher on MATH500, GSM8K and AIME using no reference answers; the two are complementary, and OPPD applied after GRPO adds up to 9.3 points. Trained only on mathematics, OPPD raises HumanEval accuracy by up to 5.3 points. One loss coefficient moves the sharpening exponent the model absorbs between 1.19 and 2.02, against 1.14 for ordinary on-policy distillation, and it rises mostly on the model's own answers. Gains hold across model families and sizes, including a model already trained with verified rewards, where lowering the temperature gives nothing and OPPD adds 4.4 points on MATH500. Code: https://github.com/ArminAzizi98/OPPD.
Figures & tables
Figure 1: Oppd moves power sampling from inference time into training. (a) The power distribution gives each complete answer a probability proportional to the model’s probability of that answer raised to the power α . (b) One training step: (1) the trainable student generates its own candidate answers, (2) the candidates grow in blocks of tokens, and (3) a frozen teacher scores each new block. (4) The teacher probabilities give two weights. The importance weight (“SMC weight”) decides which partial answers are discarded or duplicated, and the training weight sets how much each complete answer contributes to the student update. (c) Power sampling generates and scores 16 candidates for every query, while the trained student generates one answer and obtains 94% of the sampler’s gain on MATH500 (Table 9 ).
MATH500
GSM8K
pass@1
best-of-4
pass@1
best-of-4
sampling at temperature 1
untrained
50.25
82.0
59.25
91.0
on-policy distillation (baseline)
62.25
86.5
74.88
93.5
oppd , λ=1
64.00
84.0
79.38
96.0
oppd , λ=0
73.25
88.5
86.50
95.5
Table 1: Self-distillation on Qwen2.5-Math-7B . The published rows are for the same model. Generating a single answer, the model trained with oppd scores above power sampling with 64 candidates and matches GRPO, which uses reference answers during training.
Figure 2: (a) The absorbed exponent αeff decreases monotonically as the anchor coefficient λ increases, across five settings; bars are interquartile ranges over per-context fits. (b) The same estimator on contexts written by the initialization and on the model’s own answers. The two agree for ordinary on-policy distillation, while for oppd the estimate is much higher on its own answers, as expected when the training signal comes only from those answers. Exact values in Table 10 .
model
training
untrained
trained
Δ [ 95% interval]
Qwen3-8B-Base
oppd , λ=0
76.52
81.86
+5.34 [ +1.83 , +8.84 ]
Qwen3-8B-Base
GRPO
76.52
78.51
+1.98 [ −1.68 , +5.64 ]
Qwen3-14B-Base
oppd , λ=0
80.18
83.69
+3.51 [ +0.30 , +6.71 ]
Llama-3.1-8B-Instruct
oppd , λ=1
63.41
68.14
+4.73 [ +0.76 , +8.69 ]
Qwen2.5-Math-7B
oppd , λ=0
42.53
46.49
+3.96 [ −0.30 , +8.23 ]
Qwen2.5-Math-7B
oppd , λ=1
42.53
47.87
+5.34 [ +1.07 , +9.76 ]
Table 2: HumanEval, 164 problems, four samples each, pass@1. Every trained model was trained on mathematics questions only. Δ is trained minus untrained pass@1 on the same problems, with a 95% interval from a paired bootstrap ( 100,000 resamples).
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Accuracy against inference cost on MATH500, with n=250 problems. Sixteen candidates add 12.4 points for the untrained model. oppd obtains 11.6 of these points with a single generation, and running the sampler on the trained model adds only 1.6 more. Both markers on each row are measured values, and no interpolation between them is implied. Exact values in Table 9 .
MATH500
GSM8K
pass@1
best-of-4
pass@1
best-of-4
sampling at temperature 1
seed 0
73.25
88.5
86.50
95.5
seed 1
71.50
88.5
85.88
95.5
seed 2
74.63
89.5
89.50
95.5
mean
73.13
88.8
87.29
95.5
Appendix
Table 3: oppd , λ=0 , self-distilled on Qwen2.5-Math-7B with three training seeds, 300 steps each. Seed 0 is the model of Table 1 .
proposal distribution
ESS/N
D^2
own
base
ratio
MATH@1
untrained
0.561
0.579
baseline
0.138
0.292
0.47
62.25
oppd , λ=1
0.602
0.508
oppd , λ=1
0.127
0.273
0.47
64.00
oppd , λ=0
0.635
0.454
oppd , λ=0
0.096
0.296
0.32
73.25
Appendix
Table 4: Left: ESS/N with each model as the proposal distribution, and the divergence D^2 to the power distribution that it implies through ( 4 ). Right: the untrained model’s token entropy along the answers each trained model writes (own), along its own answers (base), and the ratio of the two. The baseline is ordinary on-policy distillation.
Figure 4: Replication across three models, each sampled with its default generation configuration: temperature 1 , and for Llama-3.1-8B-Instruct temperature 0.6 with nucleus (top- p ) sampling at p=0.9 . oppd improves on both the untrained model and ordinary on-policy distillation on every model and both benchmarks. Exact values in Table 11 ; best-of-4 and the temperature- 1/α rows are in Appendix L .
MATH500
GSM8K
AIME
pass@1
best-of-4
pass@1
best-of-4
pass@1
best-of-4
sampling at temperature 1
untrained
65.50
86.5
87.62
97.5
8.75
25.0
oppd
84.00
93.0
95.88
97.0
16.25
28.3
sampling at temperature 1/α
untrained
82.00
89.0
92.38
96.0
12.92
20.0
Appendix
Table 5: Qwen3-14B-Base , 300 steps at λ=0 , evaluated at both decoding temperatures. Sampled at temperature 1 , oppd scores above the untrained model at either temperature in pass@1 on all three benchmarks, and in best-of-4 on MATH500 and AIME.
sampling
task
untrained
oppd
Δ [ 95% interval]
temperature 1
MATH500
47.12
51.50
+4.38 [ +0.88 , +7.88 ]
GSM8K
85.38
87.75
+2.38 [ +0.00 , +4.75 ]
temperature 1/α
MATH500
47.12
49.88
+2.75 [ −0.38 , +6.00 ]
GSM8K
86.00
86.88
+0.88 [ −0.75 , +2.38 ]
Appendix
Table 6: deepseek-math-7b-rl , 300 steps at λ=1 , 200 problems and four samples per problem. Δ is the pass@1 of oppd minus that of the untrained model on the same problems, with a 95% interval from a paired bootstrap over problems ( 100,000 resamples).
MATH500 pass@1
GSM8K pass@1
untrained model
50.25
59.25
after GRPO with verified rewards
64.00
71.00
GRPO, then oppd , λ=1
72.13
79.75
GRPO, then oppd , λ=0
73.25
79.38
Appendix
Table 7: oppd applied to a checkpoint that has already been trained with verified rewards. The second row is that checkpoint. GRPO moves the untrained model from 50.25 to 64.00 on MATH500, and oppd adds a further 9.25 points.
Figure 5: AIME 2024 and 2025 pooled, n=60 problems, four answers each. oppd reaches the same pass@1 as training with verified rewards without using reference answers, and applying it after that training exceeds either method alone. Best-of-4 rises to 31.67 , so the combined model solves problems that the untrained model does not solve in four attempts. The same evaluation on Qwen3-8B-Base gives 7.08 untrained and 14.17 with oppd . Exact values in Table 12 .
Figure 6: oppd and GRPO trained from the same Qwen3-8B-Base checkpoint with an equal budget. Solid bar: pass@1; pale extension: best-of-4; dashed line: untrained pass@1.
MATH500
GSM8K
teacher
pass@1
best-of-4
pass@1
best-of-4
none (untrained)
49.75
77.0
57.88
88.5
Qwen2.5-Math-7B
58.25
82.5
67.63
89.5
the student itself
59.88
82.5
73.25
91.0
Appendix
Table 8: Qwen2.5-Math-1.5B trained twice from the same starting point, 300 steps each, differing only in which model scores the candidates. Scoring against the student’s own distribution matches or exceeds scoring against a 4× larger teacher.
proposal distribution
N=1
N=16
gain from candidates
best candidate
ESS/N
untrained
66.8
79.2
+12.4
82.4
0.561
oppd , λ=1
71.2
76.4
+5.2
79.6
0.602
oppd , λ=0
78.4
80.0
+1.6
81.6
0.635
Appendix
Table 9: Each model used as the proposal distribution of the sampler, on MATH500 with n=250 and every candidate graded. Generating a single answer, oppd at λ=0 comes within 0.8 points of the untrained model with sixteen candidates.
αeff(pbase→pϕ)
IQR
contexts from base
own answers
ratio
oppd , λ=0
2.019
[1.713,2.356]
baseline
1.105
1.144
1.04
oppd , λ=0.5
1.283
[1.204,1.372]
oppd , λ=0
1.207
2.019
1.67
oppd , λ=1
1.243
[1.172,1.319]
oppd , λ=2
1.193
[1.144,1.286]
baseline
1.144
[1.110,1.178]
Appendix
Table 10: Left: the absorbed exponent decreases as λ increases, across five settings, with Qwen2.5-Math-1.5B as the student, α=4 , and contexts taken from the answers each model writes itself. Right: the same estimator on two sets of contexts. Only the model trained with the sequence-level term measures higher on its own answers. The baseline is ordinary on-policy distillation.
Qwen3-8B-Base
Llama-3.1-8B-Instruct
MATH
GSM
MATH
GSM
@1
@4
@1
@4
@1
@4
@1
@4
untrained
65.50
86.0
78.00
96.5
untrained
49.25
69.5
82.25
92.5
baseline
75.00
86.5
93.00
97.5
baseline
49.75
65.5
85.20
94.0
oppd , λ=0
78.25
87.5
95.13
99.0
oppd , λ=0
47.75
63.0
85.75
91.5
oppd , λ=1
52.50
68.5
85.38
93.0
Appendix
Table 11: Replication at each model’s default generation configuration. Left: Qwen3-8B-Base , temperature 1 . Right: Llama-3.1-8B-Instruct , temperature 0.6 with nucleus sampling at p=0.9 . On Llama-3.1-8B-Instruct , lowering the sampling temperature to 1/α gains only 0.6 points on MATH500, so gains measured here come from training rather than from decoding. The baseline is ordinary on-policy distillation.
pass@1
best-of-4
untrained Qwen2.5-Math-7B
4.17
13.33
after GRPO with verified rewards
10.83
21.67
oppd , no reference answers
10.83
20.00
GRPO, then oppd , λ=0
12.92
31.67
untrained Qwen3-8B-Base
7.08
16.67
Qwen3-8B-Base with oppd
14.17
26.67
Appendix
Table 12: AIME 2024 and 2025 pooled, n=60 problems, four answers each. oppd reaches the same pass@1 as training with verified rewards without using reference answers. Applying it after that training exceeds either method alone, and best-of-4 rises to 31.67 , so the combined model solves problems that the untrained model does not solve in four attempts.
MATH500
GSM8K
AIME
pass@1
best-of-4
pass@1
best-of-4
pass@1
best-of-4
untrained
65.50
86.0
78.00
96.5
7.08
16.67
GRPO, reference answers
74.50
86.5
91.13
97.5
8.75
16.67
oppd , no reference answers
78.25
87.5
95.13
99.0
14.17
26.67
Appendix
Table 13: oppd and GRPO trained from the same Qwen3-8B-Base checkpoint for 300 steps each, generating 9,600 answers each, then evaluated with the same code. GRPO uses a reference answer for every training question, and oppd uses none. The two runs used the same GSM8K prompt set.
MATH500
GSM8K
pass@1
best-of-4
pass@1
best-of-4
temperature 1
untrained
65.50
86.0
78.00
96.5
on-policy distillation
75.00
86.5
93.00
97.5
oppd , λ=0
78.25
87.5
95.13
99.0
temperature 1/α
Appendix
Table 14: Qwen3-8B-Base , both decoding temperatures.
MATH500
GSM8K
pass@1
best-of-4
pass@1
best-of-4
untrained, default configuration
49.25
69.5
82.25
92.5
untrained, temperature 1/α
49.88
66.5
83.50
93.0
on-policy distillation
49.75
65.5
85.20
94.0
oppd , λ=0
47.75
63.0
85.75
91.5
oppd , λ=1
52.50
68.5
85.38
93.0
Appendix
Table 15: Llama-3.1-8B-Instruct . Lowering the decoding temperature gains only +0.6 and +1.25 on this model, so improvements measured at its default configuration (temperature 0.6 , nucleus sampling at p=0.9 ) come from training rather than from decoding.
model
proposal, N
% correct
weight on correct
difference
1.5B
untrained, 16
53.5
53.0
−0.45
trained, 16
49.8
49.7
−0.06
Qwen2.5-Math-7B
untrained, 16
78.9
78.8
−0.10
oppd λ=1 , 16
76.4
76.7
+0.32
oppd λ=0 , 16
80.0
80.1
+0.15
Qwen3-8B-Base
untrained, 4
81.20
81.04
−0.16
Appendix
Table 16: Percentage of candidates that are correct, against the total normalized importance weight on correct candidates. If the teacher’s probabilities ranked finished answers usefully, the weight on correct candidates would exceed the percentage correct.
Standard on-policy distillation (OPD) for large language models estimates the reverse-KL objective using student-sampled tokens, yielding an unbiased single-sample Monte Carlo estimator that avoids vocabulary-wide computation. However, we show that this estimator suffers from severe training pathologies in practice: sample inefficiency, unstable generation dynamics, and a substantial performance gap compared to exact full-vocabulary OPD. Reward-level diagnosis traces these pathologies to the log-ratio reward, which is unbounded by construction, producing extremely high-variance gradients concentrated at early positions and persisting throughout training; standard post-hoc scaling fail as they operate only after this distortion occurs. To solve this problem, we propose PowerOPD: a family of natively bounded, sign-consistent rewards from the Box-Cox power transformation, parameterized by alpha > 0, of which the log-ratio is the degenerate alpha -> 0 limit. Across six mathematical reasoning benchmarks and four Qwen3 teacher-student pairs, PowerOPD achieves benchmark-averaged Avg@8/Pass@8 gains of up to +6.37/+5.71 over vanilla OPD, +3.01/+3.54 over post-hoc stabilization, and +2.59/+8.90 over full-vocabulary OPD, while reducing wall-clock time by 59.2% and peak GPU memory by 23.1%. Larger alpha generally improves accuracy, consistently shortens responses, and keeps gradient norms more than 3,000x smaller than vanilla OPD.
Anhao Zhao, Junlong Tong, Yingqi Fan +3
Eastern Institute of Technology, Ningbo · Shanghai Jiao Tong University · University of Waterloo +1
On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E2-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E2-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E2-OPSD remains simple, requiring no additional forward passes or networks.
Yifei Liu, Minghao Fang, Xinyu Gu +6
The Chinese University of Hong Kong · Zhejiang University · Shanghai Artificial Intelligence Laboratory +5
Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond the sampled continuation. Teacher-weighted forward KL targets alternatives the student underestimates, with clipping to limit individual vocabulary contributions. Supervision is restricted to the highest-entropy half of student positions, concentrating learning where predictions are uncertain. Experiments on three competition-math benchmarks with Qwen3-1.7B, 4B, and 8B show improvements over the base models without role prompts at inference. Forward KL achieves the highest macro-averaged accuracy among the three evaluated divergences at every scale. Code is available at https://github.com/zhansan114514/OPSRD.
Weijie Ren, Yanwen Zhang, Hao Li +3
Zhejiang University · University of Electronic Science and Technology of China · University of Science and Technology of China