Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have become two main paradigms for post-training reasoning models. RLVR gives each response a single outcome label, leaving the steps inside it without separate credit. OPD provides token-level guidance at student-visited prefixes, but its pointwise signal does not directly reflect the pattern of teacher--student disagreement across the vocabulary. Dense, unbounded log-ratio supervision can amplify the teacher's influence, yet a strong solver is not necessarily a suitable guide when the student's solution paths depart from the teacher's. We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage. The guidance term has zero mean within each response, so the verifier advantage remains the response's mean label and the teacher only redistributes credit among the steps within it. \CoRA{} further updates a teacher LoRA with verifier advantages on the same scored student batch and uses the updated teacher in the next iteration's residual, adapting guidance to the student's attempts. With Qwen3-1.7B-Base and Qwen3-4B-Base students and a Qwen3-8B teacher, \RA{} combined with GRPO or REINFORCE++ improves the underlying sequence-advantage algorithm in all 24 comparisons on three mathematical benchmarks, raising macro Avg@8 by 1.7--3.6 points and Pass@8 by 3.9--6.3 points. Both combinations surpass teacher-only OPD, and \CoRA{} adds a further 1.0--1.5 Avg@8 points.
Figures & tables
Figure 1: Local teacher guidance redistributes the verifier’s response-level credit. (a) At one prefix, the reward Qd(s,v)=qs(v)−ps(v) over the student’s likeliest tokens, its value Vdp(s) under the student policy, and the advantage Asd(a) of the sampled token, here the student’s third choice and the teacher’s first. (b) Labels along one response, with blue and coral stems for positive and negative guidance terms δt ; each label is AV plus a centered guidance term, so the labels’ weighted mean is AV .
Figure 2: One iteration of RA and Co- RA . Three schematic responses, only the middle one correct, receive verifier advantages AiV . RA adds a centered guidance term to every token label. Token strips show one cell per position, blue for At>0 and rose for At<0 , darker for larger ∣At∣ ; the last cell is the response mean, which centering pins to AiV . Insets show (a) the student value baseline at one prefix and (b) response centering. In Co- RA , a LoRA adapter on the frozen teacher is trained with the verifier advantages of the same scored batch and supplies the next iteration’s residual.
Method
MATH-500
AIME24
AIME25
Macro
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Qwen3-1.7B-Base
36.83
76.67
2.08
10.00
1.53
8.89
13.48
31.85
OPD (teacher only)
67.78
85.00
8.61
23.33
5.97
15.56
27.45
41.30
REINFORCE++
65.76
83.87
7.78
20.00
3.47
14.44
25.67
39.44
+ RA
69.69
85.60
13.06
30.00
5.00
20.00
29.25
45.20
Table 1: Residual guidance improves both RLVR estimators on every benchmark and metric, and the adapted teacher raises accuracy further. Three-seed means in %. Student rows are untrained, rows marked + add RA or Co- RA to the estimator above, bold marks the best per column and student, and the last row is the Qwen3-8B teacher; per-seed scores and SD are in Appendix D .
Variant (weight)
Macro Avg@8 at weight
Macro Pass@8 at weight
5
10
20
5
10
20
Full RA ( λ )
28.84
29.25
29.29
44.94
45.20
45.29
w/o student value baseline ( γx )
26.96
27.16
26.53
41.93
41.62
40.76
w/o response centering ( λ )
28.37
28.71
28.66
44.46
44.74
44.67
Raw residual (w/o both, γraw )
27.48
27.52
26.96
42.63
42.43
40.96
Direct L2 loss ( α )
27.83
28.02
27.71
43.39
43.44
42.96
Table 2: Every residual variant improves on verifier-only training at every tested weight, and full RA stays above every control at every matched weight. Qwen3-1.7B-Base with REINFORCE++; three-seed macro means in %. The weight is λ , γx , γraw , or α as marked, with 10 the reference weight; per-seed values and sample SD are in Table 23 .
Variant
Reward
Baselines
Weight
Avg@8
Pass@8
Full RA
q−p
both
10
29.25 ± 0.25
45.20 ± 1.81
Clipped centered log ratio ( c=5 )
log(q/p)
both
0.6†
28.04 ± 0.28
43.50 ± 0.59
Centered log ratio
log(q/p)
both
0.6†
27.85 ± 0.37
43.07 ± 0.21
Raw residual (w/o both)
q−p
neither
10
27.52 ± 0.43
42.43 ± 0.79
Raw log ratio
log(q/p)
neither
0.2†
27.28 ± 0.37
42.50 ± 0.68
Verifier only (REINFORCE++)
–
–
–
25.67 ± 0.66
39.44 ± 1.65
Table 3: The probability residual with both baselines is the best construction, and no log-ratio variant reaches it at any swept weight. Three-seed macro means in % with SD; weights marked † were chosen on the validation split (Appendix B.5 ), and Tables 23 and 24 give all swept weights.
Effect of the teacher-peak branch
Sampled OPD
REINFORCE++
RA +REINFORCE++
no weight
λ=5
λ=10
λ=20
λ=5
λ=10
λ=20
Δ macro Avg@8
+0.30
−0.38
−0.88
−1.76
−2.05
−2.07
−2.16
Δ macro Pass@8
+0.64
−0.46
−0.86
−2.11
−3.91
−3.41
−3.89
Table 4: The same teacher-peak branch slightly helps sampled OPD but lowers both students trained with the verifier. Change in three-seed macro score when the branch is added, Qwen3-1.7B-Base, in points. The weight λ multiplies the branch’s teacher part and, with RA , also the sampled guidance. Absolute, benchmark, and per-seed scores are in Appendix D.3 .
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
Role
Quantity
Value
Input
Student p
(0.1,0.3,0.6)
Input
Teacher q
(0.2,0.1,0.7)
Input
Sampled action a
1
Output
Local advantage Ad(a)
0.09
Output
Expected update Jpd
(0.009,−0.063,0.054)
Output
Sampled update ga
(0.081,−0.027,−0.054)
Appendix
Table 5: A sampled update that opposes the expected update. The probability vectors and sampled action are inputs; the remaining entries follow by substitution.
Figure 3: The same sampled reward can yield local advantages of opposite signs. Both teachers assign the sampled token probability 0.28 against the student’s 0.30 , giving the same log ratio ( −0.069 ) and residual ( −0.02 ). Their residual vectors have the same squared norm, 0.0398 , but the student values differ. The resulting Ad(2) is +0.021 for teacher I and −0.063 for teacher II.
Quantity
Teacher I
Teacher II
Student p
(0.20,0.30,0.50)
(0.20,0.30,0.50)
Teacher q
(0.35,0.28,0.37)
(0.07,0.28,0.65)
Residual d=q−p
(+0.150,−0.020,−0.130)
(−0.130,−0.020,+0.150)
Sampled action a
2
2
Sampled residual d(a)
-0.020
-0.020
Squared residual norm ∥d∥22
0.0398
0.0398
Appendix
Table 6: The two teachers of Figure 3 . The student distribution and sampled action are shared. The last row pairs the two teachers as positions of one response with equal position weights and unit guidance weight.
Role
Quantity
Value
Input
Student p
(0.99,0.01)
Input
Teacher q
(0.01,0.99)
Output
Common-action update
(−0.000196,0.000196)
Output
Expected update
(−0.019404,0.019404)
Appendix
Table 7: A confident student. The sampled and expected student-centered logit updates are computed from the stipulated probability vectors.
Teacher
Expected logit update Jpd
First-position guidance term with an agreement position
I
(0.0382,0.0063,−0.0445)
0.0105λ
II
(−0.0346,−0.0189,0.0535)
−0.0315λ
Appendix
Table 8: Further quantities for the two teachers. The teachers are those of Figure 3 . The final column pairs each teacher’s position with an agreement position at equal weights c=(1/2,1/2) , keeping the symbolic weight λ .
Setting
Configuration
Students
Qwen3-1.7B-Base and Qwen3-4B-Base
Teacher backbone / inference mode
Qwen3-8B (post-trained) / non-thinking
Training data / prompt count
DAPO-Math-17K / 17,398
Validation prompts drawn from the training file / problems used
256 / 64
Validation samples per problem / temperature
8 / 1.0
Training epochs / prompts per rollout batch
2 / 128
Appendix
Table 9: Training and evaluation configuration. Lengths count response tokens. The correct-only linear length penalty applies to verifier runs; all runs share the response cap. Evaluation uses unshaped answer correctness.
Student
Comparison (macro, pp)
Δ Avg@8
Δ Pass@8
1.7B
RA +REINFORCE++ − REINFORCE++
+3.58
+5.76
1.7B
Co- RA +REINFORCE++ − RA +REINFORCE++
+1.45
+1.18
1.7B
RA +REINFORCE++ − OPD
+1.80
+3.90
1.7B
Co- RA +REINFORCE++ − OPD
+3.25
+5.08
1.7B
RA +REINFORCE++ with branch − RA +REINFORCE++
-2.07
-3.41
1.7B
RA +GRPO − GRPO
+2.37
+4.35
Appendix
Table 10: Macro differences among runs. Differences of three-seed macro means from Tables 1 , 19 , and 4 , in percentage points. Each mean and difference is computed from raw counts before rounding; subtracting the displayed macro columns can differ by 0.01 .
Table 11: Probability, actor, and reporting settings. These settings supplement Table 9 .
Setting
Configuration
GRPO group size / standard deviation
8 responses / sample standard deviation
GRPO denominator stabilizer
10−6
REINFORCE++ discount / variance floor
1 / 10−8
REINFORCE++ normalization measure
Response tokens
Appendix
Table 12: Verifier advantage normalization settings. Normalization follows reward shaping, with no reward-side KL.
Operation at iteration k
RA
Co- RA
Teacher distribution qt
Fixed q0,t
Current qϕk,t
Local reward field
dt=q0,t−pk,t
dt=qϕk,t−pk,t
Student value baseline
⟨pk,t,dt⟩
⟨pk,t,dt⟩
Sampled local advantage
Atd=dt(at)−⟨pk,t,dt⟩
Centered guidance term
λ(Atd−Ad)
λ(Atd−Ad)
Appendix
Table 13: Residual guidance with a fixed or adapted teacher. Each method uses one teacher distribution, the student value baseline, and response centering. The weight is λ=10 in both reference configurations.
Setting
Configuration
Adapter rank / scaling
16 / 32
Dropout / initial effective update
0 / 0
Target matrices
Attention and MLP projections
Separate optimizer / learning rate
AdamW / 10−5
Optimizer betas
(0.9,0.999)
Weight decay / gradient norm limit
0 / 1
Appendix
Table 14: Adapter configuration and update schedule.
Figure 4: Training accuracy of Qwen3-4B-Base on DAPO-Math-17K. (A) Full two-epoch response budgets: 276,480 generations for GRPO runs and 34,560 for the one-response runs. (B) The first 34,560 generations, the shaded region in (A), show the early trajectories at the same response budget. Curves show logged training accuracy; Co- RA is omitted.
Figure 5: Training diagnostics for the 4B student with RA +REINFORCE++. The guidance weight is fixed at λ=10 , and the horizontal axis counts rollout batches. (a) Recorded mean absolute guidance term ∣δt∣ , averaged over the tokens of each batch. (b) Mean response length separated by verifier correctness. The curves are from a single training seed.
Tokens with
Share of tokens
Share of RA guidance
Share of centered log-ratio guidance
p(at)≥0.9
60.8%
0.4% ( 0.01× )
0.6% ( 0.01× )
0.5≤p(at)<0.9
11.9%
17.8% ( 1.5× )
7.1% ( 0.6× )
0.1≤p(at)<0.5
11.8%
58.9% ( 5.0× )
31.6% ( 2.7× )
p(at)<0.1
15.5%
22.9% ( 1.5× )
60.7% ( 3.9× )
q(at)<10−4
3.8%
4.1% ( 1.1× )
50.1% ( 13× )
Appendix
Table 15: Where the guidance lands. Share of tokens and of the summed absolute guidance term per bucket of the student’s probability of the sampled token, for full RA and the centered log-ratio control, both computed offline on the same 256 responses of the untrained 1.7B student, with the ratio of the two shares in parentheses. The last row overlaps the others.
Responses
n
λAd
Raw residual
Mean
SD
Mean
SD
All
256
+0.03
0.30
−0.32
0.73
M<1,000
216
+0.04
0.32
−0.31
0.78
1,000≤M≤4,000
34
+0.00
0.03
−0.38
0.29
M>4,000
6
+0.00
0.01
−0.05
0.24
Appendix
Table 16: Response mean of the guidance without centering. Mean and SD across responses of λAd (local advantage, with the student value baseline) and of λdt(at) (raw residual), on the 256 offline responses of Table 15 , with λ=10 ; M is the number of tokens and n the number of responses in each group.
Statistic
Full RA
Uncentered
Direct L2
Cosine of sampled and analytic gT
0.994
0.992
1.000
Conditional noise of gT
0.85
1.08
0
cos(gT,gV) , sampled tokens
+0.18
+0.13
0.00
cos(gT,gV) , resampled tokens
+0.02
0.00
–
Outcome-compatible teacher mass
64%
60%
49%
Batches with cos(gT,gV)<0
26%
33%
50%
Appendix
Table 17: The sampled label reproduces the direct loss’s expected teacher gradient but shares the update with the verifier. Qwen3-1.7B-Base with REINFORCE++ at weight 10 ; aggregates over three seeds.
Method
Avg@8
Pass@8
S1
S2
S3
Mean ± SD
S1
S2
S3
Mean ± SD
Qwen3-1.7B-Base
OPD
27.61
27.17
27.58
27.45 ± 0.25
41.60
40.69
41.60
41.30 ± 0.53
REINFORCE++
26.12
24.91
25.97
25.67 ± 0.66
41.20
37.93
39.18
39.44 ± 1.65
RA +REINFORCE++
29.36
28.96
29.43
29.25 ± 0.25
46.11
43.11
46.38
45.20 ± 1.81
Co- RA +REINFORCE++
30.13
30.65
31.33
30.70 ± 0.61
44.29
47.36
47.49
46.38 ± 1.81
Appendix
Table 18: Macro performance across three training seeds. Scores in %; ± denotes sample SD across the three run-level macro scores. The middle block lists the 1.7B REINFORCE++ controls of Section 4 , whose reference is the RA +REINFORCE++ run of the first block; log-ratio weights are in parentheses.
Variant
MATH-500
AIME24
AIME25
Macro
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Verifier only (REINFORCE++)
65.76
83.87
7.78
20.00
3.47
14.44
25.67
39.44
Teacher only (OPD)
67.78
85.00
8.61
23.33
5.97
15.56
27.45
41.30
Full RA ( λ=10 )
69.69
85.60
13.06
30.00
5.00
20.00
29.25
45.20
Probability residual, baselines removed
w/o student value baseline
67.17
84.87
9.72
24.44
4.58
15.56
27.16
41.62
Appendix
Table 19: Benchmark-level results for the control runs of Section 4 . Qwen3-1.7B-Base with REINFORCE++; three-seed means in %. The macro columns repeat Tables 2 and 3 ; sample SD are reported in Table 18 .
Control
Seed
MATH-500
AIME24
AIME25
Macro
S1
66.65 / 84.80
10.83 / 23.33
4.17 / 16.67
27.22 / 41.60
S2
67.73 / 86.00
10.42 / 23.33
4.17 / 20.00
27.44 / 43.11
w/o student value baseline
S3
67.13 / 83.80
7.92 / 26.67
5.42 / 10.00
26.82 / 40.16
S1
68.48 / 85.00
13.33 / 26.67
4.58 / 16.67
28.80 / 42.78
S2
69.50 / 85.60
11.25 / 30.00
4.17 / 20.00
28.31 / 45.20
w/o response centering
S3
69.13 / 85.40
12.50 / 30.00
5.42 / 23.33
29.01 / 46.24
Appendix
Table 20: Per-seed benchmark results for seven controls. Qwen3-1.7B-Base with REINFORCE++; entries are Avg@8 / Pass@8 in %. Weights: no student value baseline γx=10 , no response centering λ=10 , raw residual γraw=10 , clipped ( c=5 ) and centered log ratio γlog=0.6 , raw log ratio β=0.2 , and direct L2α=10 .
Base method + branch
MATH-500
AIME24
AIME25
Macro
λ
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Sampled OPD
–
63.23
83.27
7.50
20.00
4.58
14.44
25.10 ± 1.00
39.24 ± 2.53
REINFORCE++
5
65.58
83.60
6.94
18.89
3.33
14.44
25.29 ± 1.10
38.98 ± 2.22
REINFORCE++
10
65.06
82.40
6.81
21.11
2.50
12.22
24.79 ± 1.17
38.58 ± 3.02
REINFORCE++
20
64.36
82.00
4.86
16.67
2.50
13.33
23.91 ± 1.91
37.33 ± 3.63
RA +REINFORCE++
5
66.35
84.20
9.31
24.44
4.72
14.44
26.79 ± 1.22
41.03 ± 2.07
Appendix
Table 21: Benchmark-level results for the teacher-peak branch. Qwen3-1.7B-Base; three-seed means in %, with sample SD for the macro scores.
Base method + branch
λ
S1
S2
S3
Sampled OPD
–
25.63 / 41.07
23.94 / 36.36
25.74 / 40.29
REINFORCE++
5
25.93 / 40.16
24.02 / 36.42
25.91 / 40.36
REINFORCE++
10
24.78 / 39.36
23.62 / 35.24
25.96 / 41.13
REINFORCE++
20
24.39 / 37.40
25.53 / 40.93
21.80 / 33.67
RA +REINFORCE++
5
27.43 / 42.71
25.39 / 38.71
27.56 / 41.67
RA +REINFORCE++
10
28.19 / 44.16
26.20 / 38.84
27.14 / 42.38
Appendix
Table 22: Per-seed macro scores for the teacher-peak branch. Entries are macro Avg@8 / Pass@8 in %.
Variant
Weight
S1
S2
S3
Avg@8
Pass@8
5
28.62 / 45.27
29.00 / 44.42
28.89 / 45.13
28.84 ± 0.20
44.94 ± 0.45
10
29.36 / 46.11
28.96 / 43.11
29.43 / 46.38
29.25 ± 0.25
45.20 ± 1.81
Full RA
20
29.67 / 45.33
30.75 / 48.09
27.45 / 42.44
29.29 ± 1.68
45.29 ± 2.82
5
27.70 / 42.84
29.19 / 45.40
28.22 / 45.13
28.37 ± 0.76
44.46 ± 1.40
10
28.80 / 42.78
28.31 / 45.20
29.01 / 46.24
28.71 ± 0.36
44.74 ± 1.78
w/o response centering
20
29.46 / 46.38
30.33 / 46.58
26.18 / 41.07
28.66 ± 2.19
44.67 ± 3.13
Appendix
Table 23: Weight sweeps with the probability residual. Qwen3-1.7B-Base with REINFORCE++; S1–S3 entries are macro Avg@8 / Pass@8 in %. The shaded row is the reference configuration of Table 1 , and † marks the direct- L2 weight selected on the validation split and reported in Table 3 . The weight is λ for full RA and the uncentered control, γx and γraw for the two controls without the student value baseline, and α for direct L2 .
Variant
Weight
S1
S2
S3
Avg@8
Pass@8
0.05
26.70 / 41.40
26.36 / 41.73
26.71 / 41.67
26.59 ± 0.20
41.60 ± 0.18
0.2†
27.51 / 42.71
26.85 / 41.73
27.47 / 43.04
27.28 ± 0.37
42.50 ± 0.68
0.6
27.02 / 42.04
26.53 / 41.53
27.04 / 41.33
26.86 ± 0.29
41.64 ± 0.37
1.2
25.27 / 38.78
25.05 / 39.18
25.25 / 38.98
25.19 ± 0.12
38.98 ± 0.20
Raw log ratio ( β )
2
24.79 / 37.40
23.71 / 37.80
24.55 / 37.67
24.35 ± 0.57
37.62 ± 0.20
0.05
26.61 / 41.47
26.30 / 40.29
26.59 / 40.96
26.50 ± 0.17
40.90 ± 0.59
Appendix
Table 24: Weight sweeps with the log ratio. Qwen3-1.7B-Base with REINFORCE++; S1–S3 entries are macro Avg@8 / Pass@8 in %. For each control, † marks the weight selected on the validation split (Appendix B.5 ) and reported in Table 3 .
Variant
Weight
Seed
MATH-500
AIME24
AIME25
Macro
S1
65.45 / 83.60
8.75 / 20.00
4.17 / 20.00
26.12 / 41.20
S2
65.58 / 83.80
7.92 / 20.00
1.25 / 10.00
24.91 / 37.93
Verifier only
–
S3
66.25 / 84.20
6.67 / 20.00
5.00 / 13.33
25.97 / 39.18
S1
68.35 / 85.80
12.50 / 30.00
5.00 / 20.00
28.62 / 45.27
S2
69.50 / 86.60
11.67 / 26.67
5.83 / 20.00
29.00 / 44.42
Full RA
5
S3
68.35 / 85.40
12.92 / 30.00
5.42 / 20.00
28.89 / 45.13
Appendix
Table 25: Per-seed benchmark results for the residual weight sweeps. Qwen3-1.7B-Base with REINFORCE++; entries are Avg@8 / Pass@8 in %. The shaded rows are the reference runs of Table 1 . The weight is λ for full RA and the uncentered control, γx for the control without the student value baseline, and γraw for the raw residual. Macro values average the three benchmark percentages within each run.
Run
MATH-500
AIME24
AIME25
Macro
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
OPD (two epochs)
67.78
85.00
8.61
23.33
5.97
15.56
27.45 ± 0.25
41.30 ± 0.53
REINFORCE++
65.76
83.87
7.78
20.00
3.47
14.44
25.67 ± 0.66
39.44 ± 1.65
OPD, then REINFORCE++
68.65
85.27
9.72
26.67
6.81
18.89
28.39 ± 0.57
43.61 ± 1.88
REINFORCE++ + reverse KL
66.21
84.27
8.33
21.11
3.89
14.44
26.14 ± 0.97
39.94 ± 1.97
RA +REINFORCE++
69.69
85.60
13.06
30.00
5.00
20.00
29.25 ± 0.25
45.20 ± 1.81
Appendix
Table 26: Sequential and joint combinations of OPD and REINFORCE++ against RA +REINFORCE++ at a matched response budget. Qwen3-1.7B-Base; three-seed means in %, with sample SD for the macro scores and the per-seed scores of the two combinations below; the reverse-KL weight is β=0.2 . The other rows repeat Table 1 ; bold marks the best mean per column.
Run
MATH-500
AIME24
AIME25
Macro
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
OPD, full vocabulary
67.78
85.00
8.61
23.33
5.97
15.56
27.45 ± 0.25
41.30 ± 0.53
OPD, sampled token
62.31
82.47
7.64
20.00
4.44
13.33
24.80 ± 0.79
38.60 ± 2.42
REINFORCE++
65.76
83.87
7.78
20.00
3.47
14.44
25.67 ± 0.66
39.44 ± 1.65
RA +REINFORCE++
69.69
85.60
13.06
30.00
5.00
20.00
29.25 ± 0.25
45.20 ± 1.81
OPD, sampled token, by seed
Appendix
Table 27: Full-vocabulary and sampled OPD. Qwen3-1.7B-Base; three-seed means in %, with sample SD for the macro scores and the per-seed scores of the sampled run below. The other rows repeat Table 1 .
Method
K=8
K=16
K=32
K=64
K=128
K=256
K=512
Qwen3-1.7B-Base (untrained)
31.85
36.07
39.73
41.79
46.12
47.87
48.95
increment per doubling
–
+4.22
+3.66
+2.06
+4.33
+1.75
+1.08
REINFORCE++
39.44
43.16
46.03
49.57
52.30
54.21
55.50
increment per doubling
–
+3.72
+2.87
+3.54
+2.73
+1.91
+1.29
OPD (teacher only)
41.30
46.35
50.37
53.50
55.10
57.79
59.12
increment per doubling
–
+5.05
+4.02
+3.13
+1.60
+2.69
+1.33
Appendix
Table 28: Macro Pass@K up to K=512 , Qwen3-1.7B-Base with REINFORCE++. Three-seed means in %, with the increment from the previous K in gray.
Figure 6: The margin over the untrained student does not close up to Pass@512. Macro Pass@K of Table 28 and each trained run’s margin over the untrained student, in points.
Model
MATH-500
AIME24
AIME25
Macro
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
1.7B-Base
36.83 ± 0.17
76.67 ± 1.10
2.08 ± 1.10
10.00 ± 3.33
1.53 ± 0.64
8.89 ± 1.92
13.48 ± 0.56
31.85 ± 1.00
4B-Base
51.79 ± 0.30
89.60 ± 0.40
7.22 ± 0.64
22.22 ± 1.92
4.72 ± 0.24
20.00 ± 3.33
21.25 ± 0.13
43.94 ± 1.17
8B teacher
84.23 ± 0.33
95.53 ± 0.70
26.25 ± 1.25
43.33 ± 0.00
21.39 ± 2.29
38.89 ± 6.94
43.95 ± 1.23
59.25 ± 2.47
Appendix
Table 29: Untrained students and the teacher under the evaluation protocol. Qwen3-1.7B-Base and Qwen3-4B-Base before any training, and the Qwen3-8B teacher in non-thinking mode. Scores in %; ± is the sample SD across three evaluation runs.
Teacher
MATH-500
AIME24
AIME25
Macro
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Avg@8
Pass@8
Frozen q0
84.23
95.53
26.25
43.33
21.39
38.89
43.95 ± 1.23
59.25 ± 2.47
Adapter, 1.7B + REINFORCE++
82.75
94.27
24.03
41.11
19.31
35.56
42.03 ± 0.77
56.98 ± 2.39
Adapter, 1.7B + GRPO
82.56
93.93
23.61
38.89
18.75
34.44
41.64 ± 0.65
55.76 ± 2.08
Adapter, 4B + REINFORCE++
82.97
94.53
26.39
44.44
20.14
36.67
43.16 ± 0.53
58.55 ± 2.46
Adapter, 4B + GRPO
82.90
94.40
24.44
40.00
21.53
40.00
42.96 ± 0.80
58.13 ± 2.32
Appendix
Table 30: The adapted teacher is a worse solver than the frozen teacher. Qwen3-8B in non-thinking mode with the Co- RA adapter of each setting, under the evaluation protocol; scores in %. The frozen row repeats Table 29 , where ± is the SD across three evaluation runs; in the adapted rows ± is the sample SD across the three seeds’ adapters.
Adapter trained with
S1
S2
S3
Qwen3-1.7B + REINFORCE++
42.29 / 57.09
41.16 / 54.53
42.63 / 59.31
Qwen3-1.7B + GRPO
42.01 / 56.89
40.89 / 53.36
42.02 / 57.02
Qwen3-4B + REINFORCE++
43.44 / 59.38
42.55 / 55.78
43.50 / 60.49
Qwen3-4B + GRPO
43.08 / 58.20
42.10 / 55.78
43.69 / 60.42
Appendix
Table 31: Per-seed macro results of the adapted teachers. Macro Avg@8 / Pass@8 in % of the adapter from each Co- RA training seed.
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple two-stage scheme, OPD-then-RL, consistently outperforms pure OPD, pure RLVR, and all such joint baselines across logic and math reasoning benchmarks. Beyond the empirical results, we further provide a systematic understanding of this through pass@k behavior, learning dynamics, and parameter updates, yielding a consistent explanation: OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support, while jointly optimizing the two signals causes them to interfere.To provide a practical recipe, we find that the OPD validation score is the key signal for when to switch to RL, and that OPD is a better cold start for RL than SFT. Together, our results establish OPD-then-RL as a simple yet strong way to combine the two methods, turning two entangled signals into complementary stages.
Boyan Li, Bingsen Chen, Chenghao Yang +3
University of Alberta · New York University · NYU Shanghai +3
Reinforcement learning with verifiable rewards (RLVR) supervises mathematical reasoning through final-answer correctness, but provides little guidance on individual tokens. On-policy distillation (OPD) supplies dense feedback on student-generated responses, yet teacher preference need not reflect correctness. Recent hybrids combine OPD and verifier-derived advantages or reweight task credit using teacher ratios. However, teacher guidance enters after verifier-based group normalization, and token reweighting need not preserve the total task credit assigned to each response. We introduce Unified Entropy-Calibrated Credit Redistribution for GRPO (UECR-GRPO), which integrates verifier and teacher signals within a single GRPO-style update at both the response and token levels. \emph{Path-Utility Unification} (PUU) combines verifier reward and a teacher-to-anchor path log-ratio in a single KL-regularized objective. Its on-policy implementation uses a length-normalized teacher score and combines both rewards before group normalization and PPO clipping, allowing teacher evidence to influence the response ranking. \emph{Entropy-Calibrated Redistribution} (ECR) then uses the signed teacher--old-policy token gap to redistribute the verifier-derived component. Full-vocabulary teacher entropy attenuates uncertain guidance, while a response-wise zero-sum projection preserves the total task credit and its token-wise sign before clipping. Across five mathematical reasoning benchmarks, UECR-GRPO achieves average Avg@12 accuracies of 17.21% and 65.09% with Qwen3-1.7B and Qwen3-4B students, respectively, exceeding the strongest baseline at each scale by 0.89 and 0.56 percentage points.
Jie Zhang, Jingxiao Yang, Zhehao Huang +2
Shanghai Jiao Tong University · Zhejiang University
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.
Wenze Lin, Jiale Zhao, Xitai Jiang +5
LeapLab, Tsinghua University · Qiuzhen College, Tsinghua University · Beihang University +2