On the Off-Policy Teacher in On-Policy Distillation
Organizations: Washington University in St. Louis · AWS AI Labs · Carnegie Mellon University · Georgia Institute of Technology
Abstract
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.
Figures & tables
| Method | Mathematical Reasoning | Code Generation | |||||||||
| AIME24 Avg@32 | AIME25 Avg@32 | AMC23 Avg@32 | HMMT Avg@32 | Olym. Avg@4 | MATH500 Avg@8 | Mean | LCB v5 Avg@4 | HE+ Avg@16 | MBPP Avg@8 | Mean | |
| Qwen3-1.7B | |||||||||||
| Qwen3-4B-Ins | |||||||||||
| GRPO | |||||||||||
| OPD | |||||||||||
| ESR | |||||||||||
| Method | Mathematical Reasoning | ||||||
| AIME24 Avg@32 | AIME25 Avg@32 | AMC23 Avg@32 | HMMT Avg@32 | Olym. Avg@4 | MATH500 Avg@8 | Mean | |
| Qwen3-1.7B | |||||||
| Qwen3-8B-DAPO | |||||||
| GRPO | |||||||
| OPD | |||||||
| ESR | |||||||
| Method | Mathematical Reasoning | ||||||
| AIME24 Avg@32 | AIME25 Avg@32 | AMC23 Avg@32 | HMMT Avg@32 | Olym. Avg@4 | MATH500 Avg@8 | Mean | |
| DeepSeek-R1-Distill-Qwen-1.5B | |||||||
| Skywork-OR1-Math-7B | |||||||
| GRPO † | |||||||
| OPD | |||||||
| ESR | |||||||
| Task | Base | w/o SCOUT | + SCOUT | Gain |
| Math | OPD | 49.2 | 51.4 | +2.2 |
| OPTR | 49.8 | 52.3 | +2.5 | |
| Code | OPD | 56.6 | 59.7 | +3.1 |
| OPTR | 59.3 | 60.5 | +1.2 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Teacher Model | Student Model | Task | Student LR | Teacher LR |
| Qwen3-4B-Instruct | Qwen3-1.7B | Math | ||
| Qwen3-8B-DAPO | Qwen3-1.7B | Math | ||
| Skywork-OR1-MATH-7B | DeepSeek-R1-Distill-Qwen-1.5B | Math | ||
| Qwen3-4B-Instruct | Qwen3-1.7B | Code |
| Benchmark | Reasoning | Code Generation | |||||||
| AIME24 | AIME25 | AMC23 | HMMT | Olympiad | MATH500 | LCB v5 | HumanEval+ | MBPP | |
| # Questions | 30 | 30 | 40 | 30 | 581 | 500 | 880 | 164 | 500 |
| # Repeats | 32 | 32 | 32 | 32 | 4 | 8 | 4 | 16 | 8 |
| Method | Run | AIME24 | AIME25 | AMC23 | HMMT | Olymp. | MATH500 | Avg6 |
| Avg.@ | 32 | 32 | 32 | 32 | 4 | 8 | – | |
| GRPO | Run 1 | |||||||
| Run 2 | ||||||||
| Run 3 | ||||||||
| Mean | ||||||||
| OPD | Run 1 |
| Method | Run | LCB v5 | HumanEval+ | MBPP | Avg. |
| Avg.@ | 4 | 16 | 8 | – | |
| GRPO | Run 1 | ||||
| Run 2 | |||||
| Run 3 | |||||
| Mean | |||||
| OPD | Run 1 |
| Method | Run | AIME24 | AIME25 | AMC23 | HMMT | Olymp. | MATH500 | Avg6 |
| Avg.@ | 32 | 32 | 32 | 32 | 4 | 8 | – | |
| GRPO | Run 1 | |||||||
| Run 2 | ||||||||
| Run 3 | ||||||||
| Mean | ||||||||
| OPD | Run 1 |
| Method | Run | AIME24 | AIME25 | AMC23 | HMMT | Olymp. | MATH500 | Avg6 |
| Avg.@ | 32 | 32 | 32 | 32 | 4 | 8 | – | |
| OPD | Run 1 | |||||||
| Run 2 | ||||||||
| Run 3 | ||||||||
| Mean | ||||||||
| ESR | Run 1 |
| Method | Run | AIME24 | AIME25 | AMC23 | HMMT | Olymp. | MATH500 | Avg. |
| Avg.@ | 32 | 32 | 32 | 32 | 4 | 8 | – | |
| (a) Qwen3-4B-Instruct-2507 Qwen3-1.7B | ||||||||
| OPD | Mean | |||||||
| OPD + Teacher GRPO | Run 1 | 47.51 | ||||||
| Run 2 | 49.57 | |||||||
| Run 3 | 49.35 | |||||||
| Method | Run | LCB v5 | HumanEval+ | MBPP | Avg. |
| Avg.@ | 4 | 16 | 8 | – | |
| OPD | Mean | ||||
| OPD + Teacher GRPO | Run 1 | 56.19 | |||
| Run 2 | 59.13 | ||||
| Run 3 | 55.72 | ||||
| Mean |
| Model | AIME24 | AIME25 | AMC23 | HMMT | Olympiad | MATH-500 | Mean |
| Qwen3-1.7B | |||||||
| Qwen3-8B-DAPO | |||||||
| OPD | |||||||
| SCOUT ( ) | |||||||
| SCOUT ( ) | |||||||
| SCOUT ( ) |
| Method | Run | AIME24 | AIME25 | AMC23 | HMMT | Olymp. | MATH500 | Avg. |
| Avg.@ | 32 | 32 | 32 | 32 | 4 | 8 | – | |
| OPD | Mean | |||||||
| SCOUT | Mean | |||||||
| OPTR | Run 1 | 49.87 | ||||||
| Run 2 | 49.20 | |||||||
| Run 3 | 50.32 |
| Method | Run | LCB v5 | HumanEval+ | MBPP | Avg. |
| Avg.@ | 4 | 16 | 8 | – | |
| OPD | Mean | ||||
| SCOUT | Mean | ||||
| OPTR | Run 1 | 59.62 | |||
| Run 2 | 59.18 | ||||
| Run 3 | 59.22 |