Organizations: Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China · Shanghai Key Laboratory of Multimodal Embodied AI, Shanghai, China
Sequential task updates are fundamental to continual learning, but their recency bias can impose a lasting performance cost. We study this cost in an overparameterized linear-regression model with i.i.d. task sampling. We prove that distribution-level forgetting and population loss converge to the same stationary limit. We quantify the additional loss incurred by sequential exact fitting, or the sequential price. In more homogeneous task geometries, it equals the intrinsic loss asymptotically attained by joint training, making the total loss twice as large. We further analyze fixed-strength elastic weight consolidation (EWC) under general task curvatures and characterize its stationary sequential price at every regularization strength. Under strong regularization, the price decays inversely with EWC strength while the mean-square coupling horizon grows proportionally. Experiments on Jester and Rotated MNIST support the predicted sequential price and its reduction by EWC, with quantitative agreement on real-world tasks satisfying the theory's assumptions and qualitative agreement under nonlinear finite-step training.
Figures & tables
Figure 1: Sequential learning from Jester user preferences. Panel (a) compares the exact forgetting and population-loss curves. Panels (b)–(c) show the stationary population loss and the number of tasks required to reduce the initial-state effect to 1% under EWC. In these two panels, solid curves are exact finite-distribution calculations, dashed curves show the large- λ expressions in Theorem 5.2 , and the dotted line marks R⋆ . Markers show averages over 2,000 independent task streams in panel (a) and 1,000 in panel (b). Error bars denote standard errors and are mostly smaller than the markers.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Additional diagnostics for sequential exact fitting on Jester. Panel (a) displays the intrinsic-loss, mean-bias, and stationary-variance components of G∞ , together with an independent task-stream estimate of the total. Panel (b) compares the stationary mean μj with the population minimizer wj⋆ for every joke; filled markers identify the five largest coordinate differences.
Figure 3: A nonlinear Rotated-MNIST stress test. Panel (a) compares forgetting and population loss under sequential training ( λ=0 ). Curves are five-evaluation moving averages and shaded bands denote standard errors across five seeds. Panel (b) reports stationary population loss as a function of functional-EWC strength. The dotted line is the joint-training loss, and the shaded gap is the empirical sequential price. Panel (c) reports the number of tasks required to enter a 1% neighborhood of the stationary population loss. Error bars in panels (b)–(c) denote standard errors across seeds.
ETH AI Center, Department of Computer Science, ETH Zürich · Mila - Quebec AI Institute, School of Computer Science, McGill University · Mila - Quebec AI Institute +1
Tieliang Gong, Zhongbo Zhang and Wen Wen are with the School of Computer Science and Technology, Xi’an Jiaotong University, China · Yong-Jin Liu is with Department of Computer Science and Technology, Tsinghua University, China