cs.LGOct 4, 2026

Learning without Overwriting: A Theory of Self-Distillation and Supervised Fine-Tuning in Continual Reasoning

Authors: Shinichi Uemura, Taiji Suzuki

Organizations: Department of Mathematics Informatics, The University of Tokyo · Center for Advanced Intelligence Project, RIKEN, Tokyo, Japan

Abstract

On-policy self-distillation (OPSD) of large language models (LLMs) has demonstrated the ability to improve reasoning capabilities while preserving previously acquired knowledge. Despite substantial empirical success, the dynamics of OPSD in continual reasoning remain incompletely understood. Modeling LLM reasoning as search over a directed acyclic graph, we provide a unified theoretical analysis of both the dynamics of post-training---OPSD and supervised fine-tuning (SFT) in continual learning---and the impact of pre-training on subsequent performance. Our findings establish three key insights with an optimization guarantee: (i) OPSD with hints from correct outputs enables continual learning without forgetting by sparse yet effective gradient descent updates induced by the hint structure. (ii) SFT on correct reasoning paths can lead to catastrophic forgetting due to dense updates along the training paths, which overwrite the information previously acquired. (iii) Diversity in pre-training is crucial for enabling a post-trained model to reach a correct output when a rollout starts from an intermediate state. Our results, supported by theoretical analysis, show that reliable continual reasoning depends on how post-training updates interact with the reasoning structure established during pre-training.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

    Jul 2, 2026Zhanming Shen, Jintao Tong, Shaotian Yan +9Unsupervised On-Policy Self-DistillationTeacher

  2. Beyond Imitation: Reflective On-Policy Self-Distillation for LLM Reasoning

    May 27, 2026Ziqi Zhao, Xinyu Ma, Liu Yang +6Unsupervised On-Policy Self-DistillationLLM Reasoning Strategies

  3. Negative Self-Distillation: Learning to Reason by Avoiding Flaws

    Sep 10, 2026Rongcan Pei, Zhepei Wei, Shuyao Xu +3Unsupervised On-Policy Self-DistillationSelf-Distillation Framework