Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
Organizations: Harbin Engineering University · Tencent · Harbin Institute of Technology
Abstract
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have become two main paradigms for post-training reasoning models. RLVR gives each response a single outcome label, leaving the steps inside it without separate credit. OPD provides token-level guidance at student-visited prefixes, but its pointwise signal does not directly reflect the pattern of teacher--student disagreement across the vocabulary. Dense, unbounded log-ratio supervision can amplify the teacher's influence, yet a strong solver is not necessarily a suitable guide when the student's solution paths depart from the teacher's. We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage. The guidance term has zero mean within each response, so the verifier advantage remains the response's mean label and the teacher only redistributes credit among the steps within it. \CoRA{} further updates a teacher LoRA with verifier advantages on the same scored student batch and uses the updated teacher in the next iteration's residual, adapting guidance to the student's attempts. With Qwen3-1.7B-Base and Qwen3-4B-Base students and a Qwen3-8B teacher, \RA{} combined with GRPO or REINFORCE++ improves the underlying sequence-advantage algorithm in all 24 comparisons on three mathematical benchmarks, raising macro Avg@8 by 1.7--3.6 points and Pass@8 by 3.9--6.3 points. Both combinations surpass teacher-only OPD, and \CoRA{} adds a further 1.0--1.5 Avg@8 points.
Figures & tables
| Method | MATH-500 | AIME24 | AIME25 | Macro | ||||
|---|---|---|---|---|---|---|---|---|
| Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | |
| Qwen3-1.7B-Base | 36.83 | 76.67 | 2.08 | 10.00 | 1.53 | 8.89 | 13.48 | 31.85 |
| OPD (teacher only) | 67.78 | 85.00 | 8.61 | 23.33 | 5.97 | 15.56 | 27.45 | 41.30 |
| REINFORCE++ | 65.76 | 83.87 | 7.78 | 20.00 | 3.47 | 14.44 | 25.67 | 39.44 |
| + RA | 69.69 | 85.60 | 13.06 | 30.00 | 5.00 | 20.00 | 29.25 | 45.20 |
| Variant (weight) | Macro Avg@8 at weight | Macro Pass@8 at weight | ||||
|---|---|---|---|---|---|---|
| 5 | 10 | 20 | 5 | 10 | 20 | |
| Full RA ( ) | 28.84 | 29.25 | 29.29 | 44.94 | 45.20 | 45.29 |
| w/o student value baseline ( ) | 26.96 | 27.16 | 26.53 | 41.93 | 41.62 | 40.76 |
| w/o response centering ( ) | 28.37 | 28.71 | 28.66 | 44.46 | 44.74 | 44.67 |
| Raw residual (w/o both, ) | 27.48 | 27.52 | 26.96 | 42.63 | 42.43 | 40.96 |
| Direct loss ( ) | 27.83 | 28.02 | 27.71 | 43.39 | 43.44 | 42.96 |
| Variant | Reward | Baselines | Weight | Avg@8 | Pass@8 |
|---|---|---|---|---|---|
| Full RA | both | 10 | 29.25 0.25 | 45.20 1.81 | |
| Clipped centered log ratio ( ) | both | 28.04 0.28 | 43.50 0.59 | ||
| Centered log ratio | both | 27.85 0.37 | 43.07 0.21 | ||
| Raw residual (w/o both) | neither | 10 | 27.52 0.43 | 42.43 0.79 | |
| Raw log ratio | neither | 27.28 0.37 | 42.50 0.68 | ||
| Verifier only (REINFORCE++) | – | – | – | 25.67 0.66 | 39.44 1.65 |
| Effect of the teacher-peak branch | |||||||
|---|---|---|---|---|---|---|---|
| Sampled OPD | REINFORCE++ | RA +REINFORCE++ | |||||
| no weight | |||||||
| macro Avg@8 | |||||||
| macro Pass@8 | |||||||
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Role | Quantity | Value |
|---|---|---|
| Input | Student | |
| Input | Teacher | |
| Input | Sampled action | |
| Output | Local advantage | |
| Output | Expected update | |
| Output | Sampled update |
| Quantity | Teacher I | Teacher II |
|---|---|---|
| Student | ||
| Teacher | ||
| Residual | ||
| Sampled action | 2 | 2 |
| Sampled residual | -0.020 | -0.020 |
| Squared residual norm | 0.0398 | 0.0398 |
| Role | Quantity | Value |
|---|---|---|
| Input | Student | |
| Input | Teacher | |
| Output | Common-action update | |
| Output | Expected update |
| Teacher | Expected logit update | First-position guidance term with an agreement position |
|---|---|---|
| I | ||
| II |
| Setting | Configuration |
|---|---|
| Students | Qwen3-1.7B-Base and Qwen3-4B-Base |
| Teacher backbone / inference mode | Qwen3-8B (post-trained) / non-thinking |
| Training data / prompt count | DAPO-Math-17K / 17,398 |
| Validation prompts drawn from the training file / problems used | 256 / 64 |
| Validation samples per problem / temperature | 8 / 1.0 |
| Training epochs / prompts per rollout batch | 2 / 128 |
| Student | Comparison (macro, pp) | Avg@8 | Pass@8 |
|---|---|---|---|
| 1.7B | RA +REINFORCE++ REINFORCE++ | +3.58 | +5.76 |
| 1.7B | Co- RA +REINFORCE++ RA +REINFORCE++ | +1.45 | +1.18 |
| 1.7B | RA +REINFORCE++ OPD | +1.80 | +3.90 |
| 1.7B | Co- RA +REINFORCE++ OPD | +3.25 | +5.08 |
| 1.7B | RA +REINFORCE++ with branch RA +REINFORCE++ | -2.07 | -3.41 |
| 1.7B | RA +GRPO GRPO | +2.37 | +4.35 |
| Setting | Configuration |
|---|---|
| Aligned model-output width / tokenizer nominal size | / |
| Probability, baseline, and label accumulation | At least float32 |
| Actor reduction / additional label whitening | Token mean / none |
| PPO clipping width / epochs per rollout batch | / |
| REINFORCE++ optimizer steps per rollout batch | |
| Student minibatch size | responses |
| Setting | Configuration |
|---|---|
| GRPO group size / standard deviation | responses / sample standard deviation |
| GRPO denominator stabilizer | |
| REINFORCE++ discount / variance floor | / |
| REINFORCE++ normalization measure | Response tokens |
| Operation at iteration | RA | Co- RA |
|---|---|---|
| Teacher distribution | Fixed | Current |
| Local reward field | ||
| Student value baseline | ||
| Sampled local advantage | ||
| Centered guidance term | ||
| Setting | Configuration |
|---|---|
| Adapter rank / scaling | / |
| Dropout / initial effective update | / |
| Target matrices | Attention and MLP projections |
| Separate optimizer / learning rate | AdamW / |
| Optimizer betas | |
| Weight decay / gradient norm limit | / |
| Tokens with | Share of tokens | Share of RA guidance | Share of centered log-ratio guidance |
|---|---|---|---|
| 60.8% | 0.4% ( ) | 0.6% ( ) | |
| 11.9% | 17.8% ( ) | 7.1% ( ) | |
| 11.8% | 58.9% ( ) | 31.6% ( ) | |
| 15.5% | 22.9% ( ) | 60.7% ( ) | |
| 3.8% | 4.1% ( ) | 50.1% ( ) |
| Responses | Raw residual | ||||
|---|---|---|---|---|---|
| Mean | SD | Mean | SD | ||
| All | 256 | ||||
| 216 | |||||
| 34 | |||||
| 6 | |||||
| Statistic | Full RA | Uncentered | Direct |
| Cosine of sampled and analytic | 0.994 | 0.992 | 1.000 |
| Conditional noise of | 0.85 | 1.08 | 0 |
| , sampled tokens | |||
| , resampled tokens | – | ||
| Outcome-compatible teacher mass | 64% | 60% | 49% |
| Batches with | 26% | 33% | 50% |
| Method | Avg@8 | Pass@8 | |||||||
| S1 | S2 | S3 | Mean SD | S1 | S2 | S3 | Mean SD | ||
| Qwen3-1.7B-Base | |||||||||
| OPD | 27.61 | 27.17 | 27.58 | 27.45 0.25 | 41.60 | 40.69 | 41.60 | 41.30 0.53 | |
| REINFORCE++ | 26.12 | 24.91 | 25.97 | 25.67 0.66 | 41.20 | 37.93 | 39.18 | 39.44 1.65 | |
| RA +REINFORCE++ | 29.36 | 28.96 | 29.43 | 29.25 0.25 | 46.11 | 43.11 | 46.38 | 45.20 1.81 | |
| Co- RA +REINFORCE++ | 30.13 | 30.65 | 31.33 | 30.70 0.61 | 44.29 | 47.36 | 47.49 | 46.38 1.81 | |
| Variant | MATH-500 | AIME24 | AIME25 | Macro | ||||
| Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | |
| Verifier only (REINFORCE++) | 65.76 | 83.87 | 7.78 | 20.00 | 3.47 | 14.44 | 25.67 | 39.44 |
| Teacher only (OPD) | 67.78 | 85.00 | 8.61 | 23.33 | 5.97 | 15.56 | 27.45 | 41.30 |
| Full RA ( ) | 69.69 | 85.60 | 13.06 | 30.00 | 5.00 | 20.00 | 29.25 | 45.20 |
| Probability residual, baselines removed | ||||||||
| w/o student value baseline | 67.17 | 84.87 | 9.72 | 24.44 | 4.58 | 15.56 | 27.16 | 41.62 |
| Control | Seed | MATH-500 | AIME24 | AIME25 | Macro |
|---|---|---|---|---|---|
| S1 | 66.65 / 84.80 | 10.83 / 23.33 | 4.17 / 16.67 | 27.22 / 41.60 | |
| S2 | 67.73 / 86.00 | 10.42 / 23.33 | 4.17 / 20.00 | 27.44 / 43.11 | |
| w/o student value baseline | S3 | 67.13 / 83.80 | 7.92 / 26.67 | 5.42 / 10.00 | 26.82 / 40.16 |
| S1 | 68.48 / 85.00 | 13.33 / 26.67 | 4.58 / 16.67 | 28.80 / 42.78 | |
| S2 | 69.50 / 85.60 | 11.25 / 30.00 | 4.17 / 20.00 | 28.31 / 45.20 | |
| w/o response centering | S3 | 69.13 / 85.40 | 12.50 / 30.00 | 5.42 / 23.33 | 29.01 / 46.24 |
| Base method + branch | MATH-500 | AIME24 | AIME25 | Macro | |||||
|---|---|---|---|---|---|---|---|---|---|
| Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | ||
| Sampled OPD | – | 63.23 | 83.27 | 7.50 | 20.00 | 4.58 | 14.44 | 25.10 1.00 | 39.24 2.53 |
| REINFORCE++ | 5 | 65.58 | 83.60 | 6.94 | 18.89 | 3.33 | 14.44 | 25.29 1.10 | 38.98 2.22 |
| REINFORCE++ | 10 | 65.06 | 82.40 | 6.81 | 21.11 | 2.50 | 12.22 | 24.79 1.17 | 38.58 3.02 |
| REINFORCE++ | 20 | 64.36 | 82.00 | 4.86 | 16.67 | 2.50 | 13.33 | 23.91 1.91 | 37.33 3.63 |
| RA +REINFORCE++ | 5 | 66.35 | 84.20 | 9.31 | 24.44 | 4.72 | 14.44 | 26.79 1.22 | 41.03 2.07 |
| Base method + branch | S1 | S2 | S3 | |
|---|---|---|---|---|
| Sampled OPD | – | 25.63 / 41.07 | 23.94 / 36.36 | 25.74 / 40.29 |
| REINFORCE++ | 5 | 25.93 / 40.16 | 24.02 / 36.42 | 25.91 / 40.36 |
| REINFORCE++ | 10 | 24.78 / 39.36 | 23.62 / 35.24 | 25.96 / 41.13 |
| REINFORCE++ | 20 | 24.39 / 37.40 | 25.53 / 40.93 | 21.80 / 33.67 |
| RA +REINFORCE++ | 5 | 27.43 / 42.71 | 25.39 / 38.71 | 27.56 / 41.67 |
| RA +REINFORCE++ | 10 | 28.19 / 44.16 | 26.20 / 38.84 | 27.14 / 42.38 |
| Variant | Weight | S1 | S2 | S3 | Avg@8 | Pass@8 |
|---|---|---|---|---|---|---|
| 5 | 28.62 / 45.27 | 29.00 / 44.42 | 28.89 / 45.13 | 28.84 0.20 | 44.94 0.45 | |
| 10 | 29.36 / 46.11 | 28.96 / 43.11 | 29.43 / 46.38 | 29.25 0.25 | 45.20 1.81 | |
| Full RA | 20 | 29.67 / 45.33 | 30.75 / 48.09 | 27.45 / 42.44 | 29.29 1.68 | 45.29 2.82 |
| 5 | 27.70 / 42.84 | 29.19 / 45.40 | 28.22 / 45.13 | 28.37 0.76 | 44.46 1.40 | |
| 10 | 28.80 / 42.78 | 28.31 / 45.20 | 29.01 / 46.24 | 28.71 0.36 | 44.74 1.78 | |
| w/o response centering | 20 | 29.46 / 46.38 | 30.33 / 46.58 | 26.18 / 41.07 | 28.66 2.19 | 44.67 3.13 |
| Variant | Weight | S1 | S2 | S3 | Avg@8 | Pass@8 |
|---|---|---|---|---|---|---|
| 0.05 | 26.70 / 41.40 | 26.36 / 41.73 | 26.71 / 41.67 | 26.59 0.20 | 41.60 0.18 | |
| 27.51 / 42.71 | 26.85 / 41.73 | 27.47 / 43.04 | 27.28 0.37 | 42.50 0.68 | ||
| 0.6 | 27.02 / 42.04 | 26.53 / 41.53 | 27.04 / 41.33 | 26.86 0.29 | 41.64 0.37 | |
| 1.2 | 25.27 / 38.78 | 25.05 / 39.18 | 25.25 / 38.98 | 25.19 0.12 | 38.98 0.20 | |
| Raw log ratio ( ) | 2 | 24.79 / 37.40 | 23.71 / 37.80 | 24.55 / 37.67 | 24.35 0.57 | 37.62 0.20 |
| 0.05 | 26.61 / 41.47 | 26.30 / 40.29 | 26.59 / 40.96 | 26.50 0.17 | 40.90 0.59 |
| Variant | Weight | Seed | MATH-500 | AIME24 | AIME25 | Macro |
|---|---|---|---|---|---|---|
| S1 | 65.45 / 83.60 | 8.75 / 20.00 | 4.17 / 20.00 | 26.12 / 41.20 | ||
| S2 | 65.58 / 83.80 | 7.92 / 20.00 | 1.25 / 10.00 | 24.91 / 37.93 | ||
| Verifier only | – | S3 | 66.25 / 84.20 | 6.67 / 20.00 | 5.00 / 13.33 | 25.97 / 39.18 |
| S1 | 68.35 / 85.80 | 12.50 / 30.00 | 5.00 / 20.00 | 28.62 / 45.27 | ||
| S2 | 69.50 / 86.60 | 11.67 / 26.67 | 5.83 / 20.00 | 29.00 / 44.42 | ||
| Full RA | 5 | S3 | 68.35 / 85.40 | 12.92 / 30.00 | 5.42 / 20.00 | 28.89 / 45.13 |
| Run | MATH-500 | AIME24 | AIME25 | Macro | ||||
|---|---|---|---|---|---|---|---|---|
| Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | |
| OPD (two epochs) | 67.78 | 85.00 | 8.61 | 23.33 | 5.97 | 15.56 | 27.45 0.25 | 41.30 0.53 |
| REINFORCE++ | 65.76 | 83.87 | 7.78 | 20.00 | 3.47 | 14.44 | 25.67 0.66 | 39.44 1.65 |
| OPD, then REINFORCE++ | 68.65 | 85.27 | 9.72 | 26.67 | 6.81 | 18.89 | 28.39 0.57 | 43.61 1.88 |
| REINFORCE++ + reverse KL | 66.21 | 84.27 | 8.33 | 21.11 | 3.89 | 14.44 | 26.14 0.97 | 39.94 1.97 |
| RA +REINFORCE++ | 69.69 | 85.60 | 13.06 | 30.00 | 5.00 | 20.00 | 29.25 0.25 | 45.20 1.81 |
| Run | MATH-500 | AIME24 | AIME25 | Macro | ||||
|---|---|---|---|---|---|---|---|---|
| Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | |
| OPD, full vocabulary | 67.78 | 85.00 | 8.61 | 23.33 | 5.97 | 15.56 | 27.45 0.25 | 41.30 0.53 |
| OPD, sampled token | 62.31 | 82.47 | 7.64 | 20.00 | 4.44 | 13.33 | 24.80 0.79 | 38.60 2.42 |
| REINFORCE++ | 65.76 | 83.87 | 7.78 | 20.00 | 3.47 | 14.44 | 25.67 0.66 | 39.44 1.65 |
| RA +REINFORCE++ | 69.69 | 85.60 | 13.06 | 30.00 | 5.00 | 20.00 | 29.25 0.25 | 45.20 1.81 |
| OPD, sampled token, by seed | ||||||||
| Method | |||||||
|---|---|---|---|---|---|---|---|
| Qwen3-1.7B-Base (untrained) | 31.85 | 36.07 | 39.73 | 41.79 | 46.12 | 47.87 | 48.95 |
| increment per doubling | – | +4.22 | +3.66 | +2.06 | +4.33 | +1.75 | +1.08 |
| REINFORCE++ | 39.44 | 43.16 | 46.03 | 49.57 | 52.30 | 54.21 | 55.50 |
| increment per doubling | – | +3.72 | +2.87 | +3.54 | +2.73 | +1.91 | +1.29 |
| OPD (teacher only) | 41.30 | 46.35 | 50.37 | 53.50 | 55.10 | 57.79 | 59.12 |
| increment per doubling | – | +5.05 | +4.02 | +3.13 | +1.60 | +2.69 | +1.33 |
| Model | MATH-500 | AIME24 | AIME25 | Macro | ||||
|---|---|---|---|---|---|---|---|---|
| Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | |
| 1.7B-Base | 36.83 0.17 | 76.67 1.10 | 2.08 1.10 | 10.00 3.33 | 1.53 0.64 | 8.89 1.92 | 13.48 0.56 | 31.85 1.00 |
| 4B-Base | 51.79 0.30 | 89.60 0.40 | 7.22 0.64 | 22.22 1.92 | 4.72 0.24 | 20.00 3.33 | 21.25 0.13 | 43.94 1.17 |
| 8B teacher | 84.23 0.33 | 95.53 0.70 | 26.25 1.25 | 43.33 0.00 | 21.39 2.29 | 38.89 6.94 | 43.95 1.23 | 59.25 2.47 |
| Teacher | MATH-500 | AIME24 | AIME25 | Macro | ||||
|---|---|---|---|---|---|---|---|---|
| Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | Avg@8 | Pass@8 | |
| Frozen | 84.23 | 95.53 | 26.25 | 43.33 | 21.39 | 38.89 | 43.95 1.23 | 59.25 2.47 |
| Adapter, 1.7B + REINFORCE++ | 82.75 | 94.27 | 24.03 | 41.11 | 19.31 | 35.56 | 42.03 0.77 | 56.98 2.39 |
| Adapter, 1.7B + GRPO | 82.56 | 93.93 | 23.61 | 38.89 | 18.75 | 34.44 | 41.64 0.65 | 55.76 2.08 |
| Adapter, 4B + REINFORCE++ | 82.97 | 94.53 | 26.39 | 44.44 | 20.14 | 36.67 | 43.16 0.53 | 58.55 2.46 |
| Adapter, 4B + GRPO | 82.90 | 94.40 | 24.44 | 40.00 | 21.53 | 40.00 | 42.96 0.80 | 58.13 2.32 |
| Adapter trained with | S1 | S2 | S3 |
|---|---|---|---|
| Qwen3-1.7B + REINFORCE++ | 42.29 / 57.09 | 41.16 / 54.53 | 42.63 / 59.31 |
| Qwen3-1.7B + GRPO | 42.01 / 56.89 | 40.89 / 53.36 | 42.02 / 57.02 |
| Qwen3-4B + REINFORCE++ | 43.44 / 59.38 | 42.55 / 55.78 | 43.50 / 60.49 |
| Qwen3-4B + GRPO | 43.08 / 58.20 | 42.10 / 55.78 | 43.69 / 60.42 |