On-policy self-distillation (OPSD) of large language models (LLMs) has demonstrated the ability to improve reasoning capabilities while preserving previously acquired knowledge. Despite substantial empirical success, the dynamics of OPSD in continual reasoning remain incompletely understood. Modeling LLM reasoning as search over a directed acyclic graph, we provide a unified theoretical analysis of both the dynamics of post-training---OPSD and supervised fine-tuning (SFT) in continual learning---and the impact of pre-training on subsequent performance. Our findings establish three key insights with an optimization guarantee: (i) OPSD with hints from correct outputs enables continual learning without forgetting by sparse yet effective gradient descent updates induced by the hint structure. (ii) SFT on correct reasoning paths can lead to catastrophic forgetting due to dense updates along the training paths, which overwrite the information previously acquired. (iii) Diversity in pre-training is crucial for enabling a post-trained model to reach a correct output when a rollout starts from an intermediate state. Our results, supported by theoretical analysis, show that reliable continual reasoning depends on how post-training updates interact with the reasoning structure established during pre-training.
Figures & tables
Figure 1: Example of the CoT graph we considered in Section 2 . Each node represents a reasoning state, and edges correspond to reasoning transitions. Black edges indicate easy intra-cluster transitions, while red and green ones mean difficult inter-cluster transitions. The goal of training is to improve difficult transitions and to reach the correct terminal. Therefore, the model is required to learn the inter-cluster transitions while preserving the diversity within clusters.
Figure 2: Training dynamics under SD (top row) and SFT (bottom row) for Tasks 1–4. Solid curves show the mean correct-generation probability over ten seeds, and shaded regions indicate the minimum–maximum range. Dash-dotted and dashed black curves represent the SD population learning references and conditional retention lower bounds. Dashed and dotted black curves represent the SFT conditional retention lower bounds and population upper references, respectively. Population references use B→∞ for stabilizing the curves. Purple dash-dotted lines show the concentration-profile term Cold(1−θtotal) . Vertical gray lines indicate task transitions.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Table 3Table 4
Task
SFT
SD
Task 1
0.002799[0.002701,0.002955]
0.980386[0.980386,0.980386]
Task 2
0.100312[0.065468,0.149365]
0.980386[0.980386,0.980386]
Task 3
0.304232[0.263488,0.334231]
0.980386[0.980386,0.980386]
Task 4
0.999998[0.999998,0.999998]
0.980386[0.980386,0.980386]
Appendix
Table 5: Final correct-generation probabilities after all four tasks. Entries are means [minimum, maximum] over ten seeds, rounded to six decimal places.
Figure 3: Training dynamics under SD for Tasks 1–4. The population learning lower reference is plotted during own-task training. At its end, the conditional retention lower bound is anchored at the observed endpoint and applies throughout subsequent-task training.
Figure 4: Training dynamics under SFT for Tasks 1–4. Black curves show conditional retention lower bounds and upper references. The concentration-profile term Cold(1−θtotal) reflects the cumulative conflict and, for Task 1, the accumulation of forgetting across subsequent tasks.