EOPSA: Efficient On-Policy Self-Distilled Safety Alignment
Organizations: The State Key Laboratory of Blockchain and Data Security, Zhejiang University · Ant Group
Abstract
On-Policy Self-Distillation (OPSD) has emerged as a promising paradigm for safety alignment, delivering dense, token-level supervision by distilling from a teacher conditioned on refusal-oriented privileged prompts. However, we reveal that this paradigm suffers from critical inefficiencies that degrade both training efficiency and general reasoning capabilities. Specifically, we diagnose two fundamental bottlenecks: (1) supervisory collapse over extended rollouts, where the teacher's corrective efficacy degrades precipitously as the student's generation prefix lengthens, injecting noisy gradients into late-stage tokens; and (2) gradient dilution from stylistic shifts, where the distillation objective is dominated by safety-irrelevant stylistic discrepancies induced by privileged prompting, washing out genuine safety signals and impairing base reasoning. To resolve these issues, we propose Efficient On-Policy Self-Distilled Safety Alignment (EOPSA), which concentrates computational and gradient budgets exclusively on reliably supervised, safety-critical tokens. EOPSA incorporates two coordinated mechanisms: (i) Adaptive Rollout Scheduling, which dynamically bounds the generation horizon guided by a novel Teacher Rescue Rate (TRR) metric to operate strictly within reliable supervision regimes; and (ii) Selective Distillation, which filters out safety-neutral tokens to restrict gradient updates exclusively to safety-pivotal transitions. Extensive evaluations across reasoning models up to 32B parameters demonstrate that EOPSA slashes rollout computation by 50% and backpropagates through merely 2% of tokens, consistently outperforming full-token distillation baselines in both safety compliance and reasoning retention.
Figures & tables
| Model | Safety ( ) | Over-Refusal ( ) | Reasoning ( ) | Efficiency | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HarmB. | WildC. | WildJ. | StrongR. | Avg | XSTest | OKTest | Avg | MATH-500 | GPQA-D | HumanEval | LCBench | Avg | Toks/Smp( ) | |
| Qwen3-1.7B (Student: Thinking, Teacher: Thinking) | ||||||||||||||
| Base | 57.50 | 45.68 | 51.20 | 63.26 | 54.41 | 0.40 | 0.00 | 0.20 | 91.20 ±0.72 | 40.24 ±1.54 | 86.38 ±1.53 | 33.33 ±0.35 | 62.79 | - |
| GRPO | 96.50 | 64.05 | 77.60 | 96.49 | 83.66 | 10.00 | 16.80 | 13.40 | 90.40 ±0.53 | 33.84 ±2.67 | 83.94 ±0.35 | 31.93 ±1.04 | 60.03 | 933.55 |
| STAR-1 | 99.00 | 72.97 | 81.60 | 99.68 | 88.31 | 27.60 | 18.07 | 22.84 | 89.13 ±1.30 | 30.98 ±2.04 | 80.69 ±2.75 | 28.31 ±2.09 | 57.28 | 361.60 |
| ThinkSafe | 97.50 | 66.22 | 67.20 | 86.90 | 79.46 | 1.60 | 3.64 | 2.62 | 90.87 ±1.15 | 37.71 ±3.55 | 80.89 ±1.76 | 31.53 ±2.43 | 60.25 | 1070.35 |
| Ablation | Safety ( ) | Over-Refusal ( ) | Reasoning ( ) | Efficiency | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HarmB. | WildC. | WildJ. | StrongR. | Avg | XSTest | OKTest | Avg | MATH-500 | GPQA-D | HumanEval | LCBench | Avg | Toks/Smp( ) | |
| w/o ARS & TF | 96.50 | 71.08 | 74.00 | 92.33 | 83.48 | 4.40 | 9.20 | 6.80 | 89.93 ±0.58 | 36.36 ±2.02 | 84.55 ±1.86 | 31.12 ±1.39 | 60.49 | 653.69 |
| w/o TF | 97.00 | 75.68 | 77.20 | 94.57 | 86.11 | 2.40 | 11.20 | 6.80 | 90.93 ±1.14 | 35.35 ±3.54 | 84.15 ±2.20 | 31.53 ±0.35 | 60.49 | 472.75 |
| w/o ARS | 99.50 | 80.00 | 82.40 | 94.89 | 89.20 | 2.80 | 12.00 | 7.40 | 91.07 ±0.81 | 36.36 ±3.64 | 84.76 ±1.61 | 32.53 ±1.04 | 61.18 | 8.82 |
| EOPSA | 99.50 | 80.27 | 88.00 | 96.81 | 91.15 | 3.20 | 9.20 | 6.20 | 91.60 ±0.20 | 36.20 ±2.78 | 85.37 ±1.61 | 33.73 ±1.59 | 61.73 | 7.41 |
| Strategy | Safety DSR (%) | Efficiency | |||||
|---|---|---|---|---|---|---|---|
| HarmB. | WildC. | WildJ. | StrongR. | Avg | Tokens | Time (h) | |
| Length = 32 | 95.50 | 65.68 | 68.00 | 92.01 | 80.30 | 32.0 | 0.18 |
| Length = 64 | 97.00 | 73.24 | 69.60 | 93.29 | 83.28 | 64.0 | 0.23 |
| Length = 128 | 97.50 | 69.73 | 74.40 | 95.21 | 84.21 | 128.0 | 0.26 |
| Length = 512 | 99.00 | 70.81 | 75.20 | 95.21 | 85.06 | 472.8 | 0.31 |
| Length = 1024 | 96.50 | 71.35 | 70.80 | 92.65 | 82.83 | 607.2 | 0.42 |
| Mode | Safety ( ) | Reasoning ( ) | Efficiency | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| HarmB. | WildC. | WildJ. | StrongR. | Safe-Avg | MATH-500 | LCBench | Avg | Toks/Smp ( ) | S/Tok ( ) | |
| Consistent | 94.00 | 66.76 | 67.20 | 89.14 | 79.27 | 90.07 | 30.92 | 60.49 | 115.17 | 0.22 |
| Function | 88.00 | 59.46 | 64.00 | 91.05 | 75.63 | 91.40 | 30.52 | 60.96 | 1.79 | 11.85 |
| Other | 94.00 | 67.84 | 63.60 | 91.05 | 79.12 | 90.53 | 30.92 | 60.73 | 7.13 | 3.47 |
| Pivot | 93.00 | 57.03 | 69.20 | 92.65 | 77.97 | 92.07 | 32.33 | 62.20 | 0.33 | 71.39 |
| Intent | 96.50 | 78.92 | 84.40 | 97.44 | 89.32 | 91.40 | 32.73 | 62.06 | 0.45 | 77.58 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Dictionary Set | Representative Tokens |
|---|---|
| but , however , maybe , still , wait , which , yet , | |
| acknowledge , address , asking , avoid , confirm , consider , focus , guidance , intent , recognize , check , safety , | |
| actionable , clarify , create , describe , explain , generate , goal , provide , understand , use , start , | |
| cannot , crime , dangerous , ethical , harm , harmful , illegal , refuse , risk , safe , threat , violence , | |
| , , . , : , ; , and , are , for , in , is , of , the , to , with , |
| Model | Safety ( ) | Over-Refusal ( ) | Reasoning ( ) | Efficiency | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HarmB. | WildC. | WildJ. | StrongR. | Avg | XSTest | OKTest | Avg | MATH-500 | GPQA-D | HumanEval | LCBench | Avg | Toks/Smp( ) | |
| Qwen3-14B | 84.50 | 68.38 | 68.00 | 97.44 | 79.58 | 0.00 | 3.20 | 1.60 | 97.13 ±0.12 | 63.30 ±1.17 | 95.53 ±0.70 | 65.66 ±3.61 | 80.41 | - |
| EOPSA | 100.00 | 97.03 | 99.60 | 100.00 | 99.16 | 4.00 | 9.20 | 6.60 | 97.67 ±0.50 | 63.64 ±1.25 | 94.11 ±1.96 | 64.06 ±3.03 | 79.87 | 8.66 |
| Qwen3-32B | 73.50 | 66.22 | 62.40 | 93.93 | 74.01 | 0.40 | 2.40 | 1.40 | 97.53 ±0.64 | 66.67 ±1.01 | 97.36 ±0.70 | 66.47 ±1.84 | 82.01 | - |
| EOPSA | 100.00 | 98.92 | 100.00 | 100.00 | 99.73 | 4.40 | 10.40 | 7.40 | 97.60 ±0.35 | 65.66 ±1.12 | 96.95 ±1.06 | 66.47 ±2.51 | 81.67 | 11.79 |
| Model | Method | PAIR ( ) | GCG ( ) | TAP ( ) | Avg ( ) |
|---|---|---|---|---|---|
| Qwen3-4B | Base | 72.0 | 79.0 | 74.0 | 75.0 |
| ThinkSafe | 95.0 | 88.0 | 90.0 | 91.0 | |
| STAR-1 | 97.0 | 55.0 | 100.0 | 84.0 | |
| OPSA | 98.0 | 63.0 | 96.0 | 85.7 | |
| EOPSA | 100.0 | 83.0 | 100.0 | 94.3 |
| Strategy | Safety ( ) | Reasoning ( ) | ||||||
|---|---|---|---|---|---|---|---|---|
| HarmB. | WildC. | WildJ. | StrongR. | Avg | MATH-500 | LCBench | Avg | |
| Random Mask | 94.50 | 64.59 | 70.00 | 92.33 | 80.36 | 90.53 | 32.13 | 61.33 |
| Top-KL | 95.50 | 74.05 | 77.60 | 93.93 | 85.27 | 90.93 | 32.33 | 61.63 |
| Top-1 Pair (Ours) | 99.50 | 81.35 | 87.20 | 92.01 | 90.02 | 90.60 | 34.94 | 62.77 |
| Teacher Strategy | Safety ( ) | Reasoning ( ) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| HarmB. | WildC. | WildJ. | StrongR. | Avg | MATH-500 | GPQA-D | HumanEval | LCBench | Avg | |
| Static Teacher | 96.50 | 71.08 | 74.00 | 92.33 | 83.48 | 89.93 | 36.36 | 84.55 | 31.12 | 60.49 |
| Synchronized (every 100 steps) | 100.00 | 83.78 | 85.20 | 97.44 | 91.61 | 86.80 | 32.83 | 76.22 | 25.90 | 55.44 |
| Category | SafeChain | STAR-1 | Mixed | Mean | Std Dev ( ) |
|---|---|---|---|---|---|
| Pivot | 149 | 156 | 147 | 150.7 | 4.73 |
| Intent | 179 | 211 | 208 | 199.3 | 17.67 |
| Risk | 219 | 228 | 216 | 221.0 | 6.24 |
| Configuration | Safety ( ) | ||||
|---|---|---|---|---|---|
| HarmB. | WildC. | WildJ. | StrongR. | Avg | |
| Base | 57.50 | 45.68 | 51.20 | 63.26 | 54.41 |
| EOPSA (SafeChain) | 99.50 | 80.27 | 88.00 | 96.81 | 91.15 |
| EOPSA (STAR-1) | 98.00 | 74.05 | 84.00 | 96.17 | 88.06 |
| EOPSA (Mixed) | 97.50 | 81.08 | 85.60 | 96.17 | 90.09 |
| (A) Classifier–Human Agreement | ||||
|---|---|---|---|---|
| Category / Setting | Acc. | Prec. | Recall | F1 |
| Pivot | – | 94.1 | 91.2 | 92.6 |
| Intent | – | 86.5 | 84.3 | 85.4 |
| Risk | – | 93.2 | 95.0 | 94.1 |
| Function | – | 96.4 | 97.1 | 96.7 |
| Consistent | – | 100.0 | 100.0 | 100.0 |