cs.LGOct 6, 2026

Does On-Policy Distillation for Safety Pose Backdoor Risks?

Authors: Jian Luo, Kehan Qi, Qingqiao Hu, Meilong Xu, Jiacheng Qiu, Weimin Lyu, Jiawei Zhou, Chao Chen

Organizations: Stony Brook University · Amazon

Abstract

On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Recent studies further explore OPD as a tool for improving large language model safety with promising results. However, these approaches typically assume that the teacher and training data are trustworthy. In this paper, we uncover an overlooked threat to OPD for safety: a safety-aligned but backdoored teacher can propagate its hidden malicious behavior to an initially clean student. Under our threat model, a poisoning rate as low as 3% results in an attack success rate (ASR) of up to 70% on the distilled student. We further identify two training choices that can amplify this risk. First, increasing the number of training epochs can lead to high ASR even at low poisoning rates. With only 10 poisoned samples, ASR reaches 67% after 16 epochs. Second, the commonly used top-k KL can accelerate backdoor transfer, causing trigger-conditioned harmful behavior to emerge earlier than sampled-token KL in most settings. Alongside these findings, we explore a simple mitigation, Lazy Defense, which clips KL rewards to make student updates less aggressive, limiting aggressive updates and slowing backdoor learning. Experiments show that Lazy Defense delays backdoor transfer in low poisoning rate settings. Together, our findings reveal that OPD can propagate backdoors, highlighting the need to address the safety risks of OPD.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes

    May 11, 2026Siqi Zhu, Xuyan Ye, Hongyu Lu +2Kullback-Leibler Divergence

  2. EOPSA: Efficient On-Policy Self-Distilled Safety Alignment

    Sep 28, 2026Qirui Liu, Yichen Sun, Yan Wang +5Safety AlignmentCollapse

  3. Constitutional On-Policy Safe Distillation

    Jun 2, 2026Ming Wen, Yuxuan Liu, Kun Yang +8Unsupervised On-Policy Self-Distillation