cs.LGSep 28, 2026

Distillation Defenses Easily Break After Reinforcement Learning

Authors: Shidan Javaheri, Alexander Panfilov, Oliver Britton, Yarin Gal, Yonatan Gideoni

Organizations: University of Oxford · ELLIS Institute Tübingen & MPI for Intelligent Systems

Abstract

Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models

    Apr 25, 2026Max Hartman, Vidhata Jayaraman, Moulik Choraria +2Effective DistillationReasoning Traces

  2. What Does It Mean to Break a Distillation Defense?

    Jun 23, 2026Lena Libon, Pura Peetathawatchai, Michael Aerni +2Realistic Threat ModelLLM Defense Mechanisms

  3. The Distillation Game: Adaptive Attacks & Efficient Defenses

    May 21, 2026Youssef Allouah, Mahdi Haghifam, Sanmi Koyejo +1Distribution Matching DistillationLLM Defense Mechanisms