cs.AISep 28, 2026

EOPSA: Efficient On-Policy Self-Distilled Safety Alignment

Authors: Qirui Liu, Yichen Sun, Yan Wang, Yu Mi, Wei Cao, Yue Shen, Zhixuan Chu, Kui Ren

Organizations: The State Key Laboratory of Blockchain and Data Security, Zhejiang University · Ant Group

Abstract

On-Policy Self-Distillation (OPSD) has emerged as a promising paradigm for safety alignment, delivering dense, token-level supervision by distilling from a teacher conditioned on refusal-oriented privileged prompts. However, we reveal that this paradigm suffers from critical inefficiencies that degrade both training efficiency and general reasoning capabilities. Specifically, we diagnose two fundamental bottlenecks: (1) supervisory collapse over extended rollouts, where the teacher's corrective efficacy degrades precipitously as the student's generation prefix lengthens, injecting noisy gradients into late-stage tokens; and (2) gradient dilution from stylistic shifts, where the distillation objective is dominated by safety-irrelevant stylistic discrepancies induced by privileged prompting, washing out genuine safety signals and impairing base reasoning. To resolve these issues, we propose Efficient On-Policy Self-Distilled Safety Alignment (EOPSA), which concentrates computational and gradient budgets exclusively on reliably supervised, safety-critical tokens. EOPSA incorporates two coordinated mechanisms: (i) Adaptive Rollout Scheduling, which dynamically bounds the generation horizon guided by a novel Teacher Rescue Rate (TRR) metric to operate strictly within reliable supervision regimes; and (ii) Selective Distillation, which filters out safety-neutral tokens to restrict gradient updates exclusively to safety-pivotal transitions. Extensive evaluations across reasoning models up to 32B parameters demonstrate that EOPSA slashes rollout computation by ∼\sim50% and backpropagates through merely ∼\sim2% of tokens, consistently outperforming full-token distillation baselines in both safety compliance and reasoning retention.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation

    May 14, 2026Yu Fu, Longxuan Yu, Haz Sameen Shahgir +4Unsupervised On-Policy Self-DistillationSafety Alignment

  2. Constitutional On-Policy Safe Distillation

    Jun 2, 2026Ming Wen, Yuxuan Liu, Kun Yang +8Unsupervised On-Policy Self-Distillation

  3. Outcome-Guided On-Policy Self-Distillation

    Oct 4, 2026ZheXu Wang, Mao-Lin Luo, Yankun Hong +6Unsupervised On-Policy Self-Distillation