cs.LGOct 5, 2026

Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning

Authors: Chen Henry Wu, Thomas Zhang, Aditi Raghunathan

Organizations: Carnegie Mellon University

Abstract

A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-policy training is a prerequisite for continual learning. In practice, however, data containing new knowledge or capabilities are often off-policy. While methods such as on-policy self-distillation (OPSD) try to bridge this gap by converting off-policy data into on-policy signal, they have been shown to cause reasoning collapse. In this paper, we show that off-policy merging beats OPSD for continual learning. We first show that SFT learns a useful signal from new data, but naively applying its update interferes with existing capabilities. We reduce this interference with a simple recipe we term grafting, which changes where the update is learned and how it is applied: (1) learning the update on an earlier donor checkpoint, ideally even before the end of pretraining, and applying the weight update to the post-trained model; (2) scaling the weight update, equivalent to a form of model merging; and (3) optionally, masking the most sensitive update directions when the new data distribution is far from the post-trained model. Across continual learning settings including (1) distilling from expert traces, (2) self-improvement with STaR and Pedagogical RL, and (3) injecting knowledge after pretraining cutoff, grafting Pareto-dominates both SFT and OPSD in new-task and old-task performance, while avoiding expensive on-policy sampling. Therefore, our work challenges on-policy training as a necessity for continual learning on RL-trained models.

Figures & tables

Explore similar work

CardsList
  1. Denser ≠\neq Better: Limits of On-Policy Self-Distillation for Continual Post-Training

    Jul 2, 2026Meng Wang, Haohan Zhao, Wenzhuo Liu +7Unsupervised On-Policy Self-DistillationReplay-Based Continual Learning

  2. Learning without Overwriting: A Theory of Self-Distillation and Supervised Fine-Tuning in Continual Reasoning

    Oct 4, 2026Shinichi Uemura, Taiji SuzukiUnsupervised On-Policy Self-DistillationSupervised Finetuning

  3. On-Policy Replay for Continual Supervised Fine-Tuning

    May 28, 2026Yan Chen, Taojie Zhu, Meng Zhang +4Supervised FinetuningModel Fine-Tuning