cs.LGSep 27, 2026

Towards Eliminating Catastrophic Forgetting in the Curriculum Learning of Math Reasoning Tasks

Authors: Zengyan Yang, Yangyang Wu, Kai Huang, Pengfei Lyu, Tianyi Zhang, Mengying Zhu

Organizations: Zhejiang University · Ant Group

Abstract

Curriculum learning has found broad application across numerous domains. Nevertheless, its effectiveness is intrinsically curtailed by catastrophic forgetting, driven by the shifts in model parameter distributions between curriculum tasks. In this paper, we investigate the phenomenon of catastrophic forgetting in this training paradigm, building on the established efficacy of curriculum learning. Our theoretical analyses of parameter update dynamics demonstrate that catastrophic forgetting in curriculum learning stems from the divergence of task optima, which is generally essential to the faster convergence of curriculum learning; therefore, forgetting cannot be completely eliminated. Based on this finding, we augment the training process and propose IV-EWC, which incorporates Elastic Weight Consolidation (EWC) into the curriculum learning objective to curb catastrophic forgetting in mathematical reasoning, a prototypical curriculum learning scenario. IV-EWC employs the influence function to construct a representative validation set from the curriculum's training data, which is used to drive dynamic regularization during training. We further present an extended theoretical analysis to show that EWC-based regularization methods mitigate catastrophic forgetting in curriculum learning, thereby providing theoretical support for IV-EWC. Empirical evaluations on three backbone models and three benchmarks indicate that curriculum learning exhibits catastrophic forgetting. IV-EWC alleviates this issue, reducing forgetting by 162% on average relative to vanilla curriculum learning and yielding positive backward transfer, as evidenced by improved performance on easier tasks after subsequent training on challenging tasks.

Figures & tables

Explore similar work

Aug 30, 2026cs.LG

The Sequential Price of Continual Learning

Sequential task updates are fundamental to continual learning, but their recency bias can impose a lasting performance cost. We study this cost in an overparameterized linear-regression model with i.i.d. task sampling. We prove that distribution-level forgetting and population loss converge to the same stationary limit. We quantify the additional loss incurred by sequential exact fitting, or the sequential price. In more homogeneous task geometries, it equals the intrinsic loss asymptotically attained by joint training, making the total loss twice as large. We further analyze fixed-strength elastic weight consolidation (EWC) under general task curvatures and characterize its stationary sequential price at every regularization strength. Under strong regularization, the price decays inversely with EWC strength while the mean-square coupling horizon grows proportionally. Experiments on Jester and Rotated MNIST support the predicted sequential price and its reduction by EWC, with quantitative agreement on real-world tasks satisfying the theory's assumptions and qualitative agreement under nonlinear finite-step training.
Jun 23, 2026cs.LG

The Gentle Collapse: Distributional Metrics for Continual Learning

Accuracy degradation is the standard metric for Catastrophic Forgetting (CF), however, it records only whether forgetting occurred or not. It saturates at the extremes and collapses discretely at task boundaries, hiding the internal structure of what is being forgotten. We introduce six softmax-derived metrics spanning true-label rank (TLR), predictive confidence, and distributional divergence that characterize forgetting continuously, each normalized to [0, 1] with no modification to training. On CIFAR-100, these metrics carry information where accuracy does not: at 0% accuracy, the Confusion Margin spans an IQR of [0.32, 0.50] across classes that accuracy treats identically. We demonstrate that this richer signal is actionable in mitigating catastrophic forgetting. Per-sample metric scores used as loss weights reduce forgetting by 1.3 percentage points over uniform experience replay (ER) on CIFAR-100. Furthermore, the slope of a metric over a small window provides a stable sampling criterion: at a small-window size (e.g. 3 epochs), accuracy-trend degrades to 34.79% (std. = 2.32) while log-TLR achieves 41.07% (std. = 0.57). This gap is structural since reliable small-window trend estimation requires a continuous signal. On TinyImageNet, log-TLR trend sampling reduces forgetting by 7.7 percentage points over the ER baseline.
Jun 30, 2026cs.LG

ISM:Self-Improving Strategy Memory for Continual Mathematical Reasoning

We propose Intelligent Schema Memory (ISM), a self-evolving memory-augmented system that improves mathematical reasoning for a frozen LLM under continual learning with hard episodic resets. ISM maintains a compact, self-refined bank of strategy schemas learned from both successful and failed episodes, with symbolic tools that check intermediate steps and certify answers. Without updating model parameters, ISM outperforms passive, retrieval, and reflection baselines on MATH-Hard and OlympiadBench, using 64% and 86% fewer schemas respectively than the strongest passive baseline. These results show that small, actively maintained, and verified strategy memories can support reliable continual mathematical reasoning under strict episodic isolation. The codebase is available at https://github.com/pdx97/ISM .