From Order to Distribution: An Exact Operator Framework for Forgetting in Continual Learning
Organizations: Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China · Shanghai Key Laboratory of Multimodal Embodied AI, Shanghai, China
Abstract
A central challenge in continual learning is forgetting: the loss of performance on previously learned tasks after learning new ones. Prior theory has analyzed forgetting under random orderings of fixed task collections in overparameterized linear regression. We shift the focus from task order to task distribution, asking how its structure determines forgetting. In the linear setting with a shared solution, i.i.d. task sampling, and sequential exact fitting, we derive an exact operator identity expressing historical forgetting directly in terms of the task distribution. Building on this identity, we establish an exponential decay guarantee for expected historical forgetting under every fixed task distribution in finite dimensions, characterize its asymptotic behavior, and relate decay to the distribution's coverage of observable directions. For an individual learned task, we show that subsequent tasks can collectively support recovery without exact revisits. We derive a lower bound on recovery time and construct a task distribution attaining its inverse-coverage scaling.
Figures & tables
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Task family | Representations | Tasks per bank | Larger-bank |
|---|---|---|---|
| MNIST classes | Pixels, ResNet-18 | 10 | 16 |
| MNIST four rotations | Pixels | 4 | 20 |
| MNIST gradual rotations | Pixels | 7 | 20 |
| MNIST permutations | Pixels | 8 | 20 |
| CIFAR-10 classes | Pixels, ResNet-18 | 10 | 16 |
| CIFAR-100 class pairs | Pixels, ResNet-18 | 10 | 16 |
| Larger source banks | |||||
|---|---|---|---|---|---|
| Task family | Representation | Peak | Peak [min, max] | ||
| MNIST classes | Pixels | 2.42 | 3.41 | 2.74 [2.58, 2.92] | 3.93 |
| MNIST four rotations | Pixels | 0.61 | 0.13 | 0.93 [0.83, 1.03] | 0.47 |
| MNIST gradual rotations | Pixels | 0.82 | 0.85 | 1.16 [1.05, 1.39] | 1.36 |
| MNIST permutations | Pixels | 0.22 | 0.06 | 0.33 [0.31, 0.36] | 0.05 |
| CIFAR-10 classes | Pixels | 1.12 | 1.33 | 1.13 [1.10, 1.19] | 1.35 |