Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.
Figures & tables
Figure 1: Composed with sampling, SFT can generalize better and forget less than on-policy learning. Left: we compare SFT composed with our sampling algorithm for off-policy data against SFT on the original dataset and an on-policy learning algorithm that integrates off-policy data into self-distillation (OPSD) ( Shenfeld et al., 2026 ) . We illustrate this on the task of scientific reasoning for chemistry problems and plot performance on prior capabilities introduced during pretraining (MMLU and AMC). Right: we compare the three along two axes: new task accuracy (chemistry) and prior task retention, which measures the percent of base model performance the finetuned models are able to maintain. SFT composed with sampling performs the best on both axes.
Figure 2: Toy schematic. Our sampling algorithm progressively shifts the original data policy πdata to a target policy πtarget that is closer to the base model distribution πbase .
πk(x0:kB)∝pC(x0:kB).
Algorithm 1 Projection Sampling with Autoregressive Models
Model
Algorithm
New Task Domain
Prior Capabilities
Chemistry
MMLU
GPQA
AMC
MATH500
GSM8K
Avg.
Base
0.343
0.687
0.253
0.422
0.734
0.918
0.597
SFT
0.618
0.586
0.242
0.277
0.622
0.871
0.520
Qwen2.5-7B-Instruct
Rewrite SFT
0.613
0.640
0.273
0.313
0.676
0.887
0.558
OPSD (on-policy)
0.618
0.651
0.222
0.349
0.716
0.901
0.568
Sampling SFT (Ours)
0.660
0.692
0.303
0.374
0.670
0.891
0.586
Table 1: Sampling enables SFT to generalize better while forgetting less. Top: Chemistry on Qwen2.5-7B-Instruct. Middle: Math on Qwen2.5-3B. Bottom: Medical on Qwen2.5-7B-Instruct. We bold the highest score in each column, and for math, we also bold if one of our listed approaches outperforms the other baselines, with a color gradient to emphasize numerical gaps. Across model sizes and evaluation tasks, sampling consistently enables SFT to generalize and retain prior capabilities at least as well as, if not better than, its on-policy posttraining counterparts.
Figure 3: Dataset likelihoods for projection sampling. We plot likelihoods of the boosted dataset traces for both the chemistry and math task under their corresponding base models (Qwen2.5-7B-Instruct and Qwen2.5-3B).
Figure 4: Pass@ k performance on Chemistry (Olmo-3-7B-Instruct) . We plot the pass@ k accuracy (correct if at least one of k samples is accurate) of SFT with projection sampling (ours) as well as on-policy learning with privileged information (OPSD) relative to the base model. Our performance curve is strictly better than both OPSD and the base model, and our pass rate at high k exceeds both those of the base model and OPSD.
Figure 5: KL divergence vs. accuracy vs. MCMC steps (Qwen2.5-7B-Instruct). For k∈[0,2,4,6,8,10] , we plot both the KL divergence of the boosted data distribution with the base model (Qwen2.5-7B-Instruct) as well as the accuracy of finetuning on the science task. As sampling compute scales, the KL gap decreases while accuracy improves.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Model
Algorithm
New Task Domain
Prior Capabilities
Chemistry
MMLU
GPQA
AMC
MATH500
GSM8K
Avg.
Base
0.328
0.562
0.419
0.361
0.738
0.925
0.601
SFT
0.567
0.568
0.338
0.422
0.722
0.902
0.590
Olmo-3-7B-Instruct
Rewrite SFT
0.563
0.570
0.359
0.410
0.764
0.916
0.604
OPSD (on-policy)
0.597
0.573
0.379
0.374
0.754
0.921
0.600
Sampling SFT (Ours)
0.583
0.575
0.389
0.446
0.762
0.916
0.617
Appendix
Table 2: Results on Chemistry using Olmo-3-7B-Instruct. Chemistry measures generalization to the new task domain, while MMLU, GPQA, AMC, MATH500, and GSM8K measure retention of prior capabilities. We report the average performance across prior-capability evaluations in the final column.
Large language model post-training methods such as supervised fine-tuning (SFT), reinforcement learning (RL), and distillation are often analyzed through their loss functions: maximum likelihood, policy gradients, forward KL, reverse KL, or related objective-level variants. We study a complementary factor: the state distribution on which supervision is applied. For an autoregressive policy, a state is a prompt plus generated prefix. SFT trains on fixed dataset states, while RL and on-policy distillation (OPD) train on states induced by the current learner. We formalize post-training as state-distribution shaping and run a controlled smallscale study using Qwen3-0.6B-Base on GSM8K, with TruthfulQA and MMLU as retention evaluations. Our results show three phenomena. First, a mild SFT run improves GSM8K with little forgetting, while a stress SFT run causes substantial retention loss. Second, OPD from a degraded SFT teacher surpasses that teacher on GSM8K, TruthfulQA, and MMLU, despite using the teacher as its only supervision source. Third, a lightweight on-policy RL run improves GSM8K while preserving retention. These results support a state-centric view of post-training: the source and locality of training states can be as important as the form of the supervision signal.
Adapting language models (LMs) to new tasks via post-training carries the risk of degrading existing capabilities -- a phenomenon classically known as catastrophic forgetting. In this paper, toward identifying guidelines for mitigating this phenomenon, we systematically compare the forgetting patterns of two widely adopted post-training methods: supervised fine-tuning (SFT) and reinforcement learning (RL). Our experiments reveal a consistent trend across LM families (Llama, Qwen) and tasks (instruction following, general knowledge, and arithmetic reasoning): RL leads to less forgetting than SFT while achieving comparable or higher target task performance. To investigate the cause for this difference, we consider a simplified setting in which the LM is modeled as a mixture of two distributions, one corresponding to prior knowledge and the other to the target task. We identify that the mode-seeking nature of RL, which stems from its use of on-policy data, enables keeping prior knowledge intact when learning the target task. We then verify this insight by demonstrating that the use on-policy data underlies the robustness of RL to forgetting in practical settings, as opposed to other algorithmic choices such as the KL regularization or advantage estimation. Lastly, as a practical implication, our results highlight the potential of mitigating forgetting using approximately on-policy data, which can be substantially more efficient to obtain than fully on-policy data.
Howard Chen, Noam Razin, Karthik Narasimhan +1
Princeton Language and Intelligence, Princeton University.
Large language models (LLMs) have achieved remarkable progress, with post-training playing a crucial role in enhancing their reasoning capabilities. Among post-training paradigms, supervised fine-tuning (SFT) is widely used: it leverages external data to provide dense supervision and enables efficient training. However, directly fine-tuning on expert data can hurt generalization when the data distribution is mismatched with the target model's own distribution. In this work, we propose Data Adaptation for Reasoning Tuning (DART), which formulates the use of a fixed, potentially distributionally misaligned SFT dataset as an optimization problem over demonstration transformations. DART trains a mapper model with reinforcement learning to convert original SFT data into model-adapted supervision that better matches the target model's distribution and learning preferences. The transformed data are then used for SFT, allowing the target model to better exploit external supervision. Experiments across multiple models and datasets show that DART improves generalization, achieves higher training efficiency than direct RL, and helps models surpass standard SFT. Our code is available at https://anonymous.4open.science/r/DART525E50D.
Lisong Sun, Li Wang, Chen Zhang +4
Beihang University · Tsinghua University · Nanyang Technological University +2