A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-policy training is a prerequisite for continual learning. In practice, however, data containing new knowledge or capabilities are often off-policy. While methods such as on-policy self-distillation (OPSD) try to bridge this gap by converting off-policy data into on-policy signal, they have been shown to cause reasoning collapse. In this paper, we show that off-policy merging beats OPSD for continual learning. We first show that SFT learns a useful signal from new data, but naively applying its update interferes with existing capabilities. We reduce this interference with a simple recipe we term grafting, which changes where the update is learned and how it is applied: (1) learning the update on an earlier donor checkpoint, ideally even before the end of pretraining, and applying the weight update to the post-trained model; (2) scaling the weight update, equivalent to a form of model merging; and (3) optionally, masking the most sensitive update directions when the new data distribution is far from the post-trained model. Across continual learning settings including (1) distilling from expert traces, (2) self-improvement with STaR and Pedagogical RL, and (3) injecting knowledge after pretraining cutoff, grafting Pareto-dominates both SFT and OPSD in new-task and old-task performance, while avoiding expensive on-policy sampling. Therefore, our work challenges on-policy training as a necessity for continual learning on RL-trained models.
Figures & tables
Figure 1: (Left) Model merging alone greatly improve generalization; Grafting dominates SFT and OPSD. Model: Qwen3-1.7B (thinking). (Right) The best donor model for grafting is even before the end of pretraining . Model: OLMo-3-7B (thinking). Only frontier points are plotted.
Figure 2: Model merging improves both old and new task performance.
Figure 3: Grafting from different θdonor with OLMo-3-7B (thinking). (Left) θdonor is varied along the post-training trajectory: base model SFT SFT checkpoint DPO DPO checkpoint RL OLMo-3-7B (thinking). (Right) θdonor is moved back along pretraining, up to 12 k steps before the base model.
Figure 4: Distillation from teacher traces across different models.
Figure 5: Self-improvement with rejection sampling (STaR) on Qwen3-1.7B (thinking).
Figure 6: Self-improvement with Pedagogical RL on Qwen3-1.7B (thinking).
Figure 7: (Left) Knowledge injection results. (Right) Disentangling the contribution of grafting’s components. The protection mask ( ρ sweep) dominates no mask ( λ sweep).
Figure 8: (Left) Grafting is much faster to train than OPSD. We show the per-step sampling and weight update time for Qwen3-8B (thinking). (Right) Grafting writes less to sensitive coordinates than SFT. We show the weighted delta ∑jFret(j)Δj2 summed across the 1% of coordinates with the largest Fret .