OPSRD: On-Policy Self-Role Distillation
Organizations: Zhejiang University · University of Electronic Science and Technology of China · University of Science and Technology of China
Abstract
Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond the sampled continuation. Teacher-weighted forward KL targets alternatives the student underestimates, with clipping to limit individual vocabulary contributions. Supervision is restricted to the highest-entropy half of student positions, concentrating learning where predictions are uncertain. Experiments on three competition-math benchmarks with Qwen3-1.7B, 4B, and 8B show improvements over the base models without role prompts at inference. Forward KL achieves the highest macro-averaged accuracy among the three evaluated divergences at every scale. Code is available at https://github.com/zhansan114514/OPSRD.
Figures & tables
| Method | AIME24 | AIME25 | HMMT25 | Macro | Base |
| Base | 46.67 | 36.11 | 24.44 | 35.74 | |
| Answer OPSD | 51.39 | 39.17 | 25.56 | 38.70 | |
| Base with expert role | 46.94 | 37.22 | 19.44 | 34.54 | |
| Base with neutral prompt | 45.83 | 35.83 | 21.94 | 34.54 | |
| Role full | 55.56 | 39.72 | 25.00 | 40.09 | |
| OPSRD | 57.78 | 41.67 | 26.67 | 42.04 |
| Model | Evaluation | Base | Answer OPSD | OPSRD | Base | Answer |
| Qwen3-1.7B | Primary | 35.74 | 38.70 | 42.04 | ||
| Qwen3-4B | Mean of 4 | 60.86 | 61.50 | 62.52 | ||
| Qwen3-8B | Primary | 62.78 | 64.26 | 65.83 |
| Training role | Test role | Macro |
| Teacher only | No role | 42.04 |
| Teacher only | Expert role | 39.44 |
| Shared role | No role | 38.98 |
| Shared role | Expert role | 36.94 |
| Student only | No role | 38.24 |
| Student only | Expert role | 36.30 |
| Model | Repeat | Method | Macro |
| 1.7B | Training | Role full | |
| 1.7B | Training | OPSRD | |
| 4B | Evaluation | Base | |
| 4B | Evaluation | Answer OPSD | |
| 4B | Evaluation | OPSRD |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Setting |
| Optimizer / steps | AdamW / 100 |
| Learning rate / maximum gradient norm | / 0.1 |
| Effective batch size / rollouts per problem | 36 / 1 |
| LoRA rank / scale | 64 / 128 |
| LoRA projections | q, k, v, o, gate, up, down |
| Training temperature / top- / top- | 1.1 / 0.95 / 20 |
| Model | Method | AIME24 | AIME25 | HMMT25 | Macro | Base |
| Qwen3-1.7B | Base | 46.67 | 36.11 | 24.44 | 35.74 | |
| Qwen3-1.7B | Answer OPSD | 51.39 | 39.17 | 25.56 | 38.70 | |
| Qwen3-1.7B | OPSRD | 57.78 | 41.67 | 26.67 | 42.04 | |
| Qwen3-4B | Base | 73.61 | 66.11 | 41.11 | 60.28 | |
| Qwen3-4B | Answer OPSD | 75.00 | 68.61 | 43.89 | 62.50 | |
| Qwen3-4B | OPSRD | 72.22 | 67.78 | 44.44 | 61.48 |
| Model | Objective | GPU | AIME24 | AIME25 | HMMT25 | Macro |
| 1.7B | Forward KL | A100 | 57.78 | 41.67 | 26.67 | 42.04 |
| 1.7B | Reverse KL | A100 | 50.56 | 35.56 | 25.56 | 37.22 |
| 1.7B | JSD | H100 | 51.39 | 38.61 | 22.22 | 37.41 |
| 4B | Forward KL | A100 | 72.22 | 67.78 | 44.44 | 61.48 |
| 4B | Reverse KL | A100 | 74.72 | 67.50 | 40.56 | 60.93 |
| 4B | JSD | H100 | 71.94 | 65.28 | 43.06 | 60.09 |
| Training role | Test role | AIME24 | AIME25 | HMMT25 | Macro | |
| Teacher only | No role | 57.78 | 41.67 | 26.67 | 42.04 | |
| Teacher only | Expert role | 51.11 | 41.94 | 25.28 | 39.44 | |
| Shared role | No role | 52.22 | 40.83 | 23.89 | 38.98 | |
| Shared role | Expert role | 47.50 | 39.44 | 23.89 | 36.94 | |
| Student only | No role | 50.28 | 39.72 | 24.72 | 38.24 | |
| Student only | Expert role | 47.22 | 38.33 | 23.33 | 36.30 |
| Quantity | Mean |
| Selected valid-token fraction | 0.500 |
| Selected-position entropy | 0.874 |
| Omitted-position entropy | 0.023 |
| Unclipped selected divergence | 0.194 |
| Fraction | Avg@4 | 50% |
| 25% | 43.06 | |
| 50% | 43.06 | |
| 75% | 38.89 |