From Order to Distribution: An Exact Operator Framework for Forgetting in Continual Learning
Authors: Zonghuan Xu, Xingjun Ma
Organizations: Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China · Shanghai Key Laboratory of Multimodal Embodied AI, Shanghai, China
A central challenge in continual learning is forgetting: the loss of performance on previously learned tasks after learning new ones. Prior theory has analyzed forgetting under random orderings of fixed task collections in overparameterized linear regression. We shift the focus from task order to task distribution, asking how its structure determines forgetting. In the linear setting with a shared solution, i.i.d. task sampling, and sequential exact fitting, we derive an exact operator identity expressing historical forgetting directly in terms of the task distribution. Building on this identity, we establish an exponential decay guarantee for expected historical forgetting under every fixed task distribution in finite dimensions, characterize its asymptotic behavior, and relate decay to the distribution's coverage of observable directions. For an individual learned task, we show that subsequent tasks can collectively support recovery without exact revisits. We derive a lower bound on recovery time and construct a task distribution attaining its inverse-coverage scaling.
Figures & tables
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1 : Synthetic experiments in the realizable exact-fit setting. In the first task distribution, panel (a) compares empirical forgetting and the exact finite-horizon expectation from the spectral expansion with the explicit upper bound and a coarse projector-based O(1/k) baseline. Both bounds are combined with the trivial bound 1/r . Panel (b) compares empirical and exact finite-horizon local decay rates with the analytic ρΠ . In the second task distribution, panel (c) sweeps L over 30 values: larger L decreases ρΠ and lowers long-horizon forgetting. Shaded regions show 95% bootstrap intervals.
Task family
Representations
Tasks per bank
Larger-bank n
MNIST classes
Pixels, ResNet-18
10
16
MNIST four rotations
Pixels
4
20
MNIST gradual rotations
Pixels
7
20
MNIST permutations
Pixels
8
20
CIFAR-10 classes
Pixels, ResNet-18
10
16
CIFAR-100 class pairs
Pixels, ResNet-18
10
16
Appendix
Table 1 : Image-task constructions. Every combination uses five banks at n=3 and n=8 images per task. Larger banks use ten pixel banks or five feature banks. Each bank is evaluated under its original grouping and random regrouping.
Figure 2 : Historical loss under original grouping and random regrouping of the same labeled images. Each task contains eight images. Bold curves use bank seed zero, and faint curves show the other four banks; all thirty curves are shown. Losses use the original scale, and vertical scales differ across panels.
n=8
Larger source banks
Task family
Representation
Peak
F(64)
Peak [min, max]
F(64)
MNIST classes
Pixels
2.42
3.41
2.74 [2.58, 2.92]
3.93
MNIST four rotations
Pixels
0.61
0.13
0.93 [0.83, 1.03]
0.47
MNIST gradual rotations
Pixels
0.82
0.85
1.16 [1.05, 1.39]
1.36
MNIST permutations
Pixels
0.22
0.06
0.33 [0.31, 0.36]
0.05
CIFAR-10 classes
Pixels
1.12
1.33
1.13 [1.10, 1.19]
1.35
Appendix
Table 2 : Paired original-to-random ratios, reported as medians across banks. The full-bank peak column also gives the range. Values summarize five banks at n=8 , ten larger pixel banks, and five larger feature banks.
Figure 3 : Paired grouping effects across every family and representation. Each point is one image bank; ratios above one indicate more forgetting under original grouping. The three marker types show n=3 , n=8 , and the larger source banks. Peak height and loss at k=64 are compared separately. The vertical reference line marks equality.
Sequential task updates are fundamental to continual learning, but their recency bias can impose a lasting performance cost. We study this cost in an overparameterized linear-regression model with i.i.d. task sampling. We prove that distribution-level forgetting and population loss converge to the same stationary limit. We quantify the additional loss incurred by sequential exact fitting, or the sequential price. In more homogeneous task geometries, it equals the intrinsic loss asymptotically attained by joint training, making the total loss twice as large. We further analyze fixed-strength elastic weight consolidation (EWC) under general task curvatures and characterize its stationary sequential price at every regularization strength. Under strong regularization, the price decays inversely with EWC strength while the mean-square coupling horizon grows proportionally. Experiments on Jester and Rotated MNIST support the predicted sequential price and its reduction by EWC, with quantitative agreement on real-world tasks satisfying the theory's assumptions and qualitative agreement under nonlinear finite-step training.
Zonghuan Xu, Xingjun Ma
Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China · Shanghai Key Laboratory of Multimodal Embodied AI, Shanghai, China
Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting. Yet its generalization behavior is shaped by two coupled effects that existing analyses fold into a single hypothesis-level quantity: finite memory replaces each past distribution with an empirical proxy, and repeated reuse couples the buffer, the current data, and the final hypothesis through a shared optimization trajectory. We develop a layer-wise information-theoretic framework that separates these effects at every depth. Our main result decomposes the expected generalization gap into a replay-induced representation drift and an optimization-dependence term, the latter further resolved into stability, plasticity, interaction, and residual-coupling components. Two refinements make the framework operational. A Wasserstein relaxation of the drift term, valid under support mismatch, yields a depth-dependent drift--sensitivity trade-off whose minimizer identifies which interior layer to stabilize. An SGLD instantiation of the optimization term reduces it to a trajectory-level log-determinant budget, exposing a curvature-aware gradient-alignment statistic that serves as an online diagnostic of task-wise forgetting. Controlled and benchmark experiments confirm the predicted memory scaling, the interior funnel, and the alignment signal's link to forgetting.
Tieliang Gong, Zhongbo Zhang, Wen Wen +1
Tieliang Gong, Zhongbo Zhang and Wen Wen are with the School of Computer Science and Technology, Xi’an Jiaotong University, China · Yong-Jin Liu is with Department of Computer Science and Technology, Tsinghua University, China
Continual learning commonly relies on post-hoc mechanisms such as replay, elastic regularization, or distillation. This work argues that forgetting should instead be modeled directly as interference between tasks. In the frozen-feature regime, forgetting from learning a new task is exactly the interference energy induced on the old task. In deep networks, the same quantity is recovered through path-averaged curvature with minimal additional forward passes. When task supports are disjoint, forgetting can be eliminated structurally and when task supports overlap in conflicting directions, a non-zero distortion floor is unavoidable. The same geometry optimally merges models through task-aware orthogonalization. From this analysis we derive Interference-Gated Functional Allocation (IGFA), a replay-free, Fisher-free method that shares directions when tasks align and protects them when they conflict. Across benchmarks, IGFA achieves lossless retention when tasks are structurally separable and moves unavoidable cost from irreversible forgetting into deferred but recoverable plasticity when they are not. It matches the strongest replay-free structural baselines on dissimilar-task streams and improves on unconditional projection when similarity makes transfer worth preserving.