Continual learning models suffer from catastrophic forgetting when trained sequentially on non-stationary data distributions. Previously, this has been addressed through weight regularization. While preconditioning gradients offer a promising alternative to mitigate forgetting, current approaches are myopic. Conversely, standard regularization methods apply rigid, scalar Euclidean penalties that entirely ignore the underlying Riemannian geometry of the parameter space. To overcome this gap, we propose TMLN (Trajectory-Modulatory Landscape Navigation), a normative navigation policy that formalizes continual learning as an optimal control problem over a curved loss landscape. TMLN utilizes a memory-efficient diagonal empirical Fisher Information Matrix (FIM) to define a localized Riemannian manifold. To compensate for the spatial limitations of the diagonal approximation, TMLN dynamically modulates a preconditioner using the normalized historical trajectory of the network's parameter values. By integrating this trajectory-based preconditioning directly into the gradient update, we actively shield historically critical parameter directions without relying on additive penalties. Empirical evaluations on class- and domain-incremental benchmarks demonstrate that our method significantly reduces the loss barrier between consecutive tasks.
Figures & tables
Figure 1: TMLN Framework. (a) On a loss landscape, continual learning tackles both forgetting (in which the model moves out of task 1’s local minima and into task 2’s local minima) and intransigence (where the model struggles to move out of task 1’s minima). (b) Trajectory Modulated Landscape Navigation measures local curvature and historical trajectory, and incorporates it into a preconditioned gradient step. (c) Our method efficiently modulates optimization learning rules, allowing sequential learning while preventing movement in steep directions that harm past knowledge. This figure is produced with the aid of Gemini 3.
Split CIFAR-100
CORe50
Method
ACC ( ↑ )
FM ( ↓ )
INT ( ↓ )
ACC ( ↑ )
FM ( ↓ )
INT ( ↓ )
Joint (oracle)
94.39±0.14
−−
−−
80.03±0.00
−−
−−
Naive SGD
7.78±0.40
66.53±2.40
21.67±2.55
63.72±0.57
41.13±0.69
23.88±0.00
ER
9.23±0.19
73.42±3.37
14.99±3.24
77.36±1.50
23.89±1.68
1.57±0.11
EWC++
6.01±0.06
58.18±0.59
33.64±0.39
76.94±0.52
16.44±0.76
9.51±0.32
RWalk
12.12±1.35
70.01±1.46
15.04±0.38
67.76±0.68
4.25±0.96
32.32±0.20
Table 1: Performance metrics on Split CIFAR-100 and CORe50. Final Accuracy (ACC), Average Forgetting (FM), and Intransigence (INT) are reported. Joint (oracle) serves as an upper bound for accuracy. All methods are averaged over five runs.
Method
Geometry (FIM)
Trajectory
ACC ( ↑ )
Forgetting ( ↓ )
Intransigence ( ↓ )
Standard SGD
–
–
19.04±0.15
89.89±1.26
0.00±0.48
TMLN-Static
✓
–
46.97±2.56
41.39±4.30
14.56±4.14
TMLN-Euclidean
–
✓
50.35±1.00
39.48±1.85
12.27±0.78
TMLN (Full)
✓
✓
51.55±0.56
40.32±3.67
9.94±3.66
Table 2: Ablation results of TMLN. We evaluate the impact of the Riemannian metric tensor (FIM) and historical-trajectory modulation on Split CIFAR-10. Results are averaged over five runs.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Parameter
Value
TMLN (Ours)
λ
0.1
γ
1.0
ER
B
500
EWC++
λ
10000
RWalk
λ
10
α
0.8
Appendix
Table 3: Hyperparameter specifications for baselines and the proposed method.
Table 4: Computational and Memory Overhead on CIFAR-100. Time is normalised relative to Naive SGD. Auxiliary memory indicates the additional state required beyond the standard model parameters ( θ ).
Task 1 Loss
Task 2 Loss
Method
Init
End of Task 1
End of Task 2
Init
End of Task 1
End of Task 2
Naive SGD
2.9338
1.3871
24.5702
11.1482
12.1247
6.3916
ER
2.3030
0.3775
22.5264
10.4751
13.3639
5.2835
TMLN
2.4773
1.1380
9.0663
8.7170
10.6379
1.0983
Appendix
Table 5: Optimization trajectory loss values across sequential tasks for Tasks 1 & 2. Init, End of Task 1, and End of Task 2 columns correspond to the loss landscape visualization markers. Best final loss is in bold.
Method
Configuration
ACC ( ↑ )
FM ( ↓ )
INT ( ↓ )
Joint (oracle)
−−
75.98±0.56
−−
−−
Naive
η=0.001
19.04±0.25
92.15±1.94
−1.55±1.64
η=0.005
19.18±0.05
91.28±0.80
−0.74±0.89
η=0.01
19.04±0.15
89.89±1.26
0.00±0.48
η=0.05
16.78±3.39
86.77±1.70
5.09±4.93
η=0.1
12.14±3.70
84.00±0.69
13.92±4.13
Appendix
Table 6: Continual learning metrics for Split CIFAR-10 at various hyperparameter configurations.
In Online Continual Learning (OCL), a neural network sequentially learns from a non-stationary data stream in a single-pass with access only to a limited memory replay buffer. This contrasts sharply with off-line continual learning where training is multiple epoch dependent on large datasets. The main challenge faced by OCL is to overcome catastrophic forgetting of past tasks (stability) while learning new ones efficiently (plasticity). Existing methods counter forgetting via replay-based rehearsal, output level distillation, fixed regularization, or meta-learning on the current data. However, these methods have limitations: rehearsal introduces a stored sample bias; distillation operates on output-distributions without modulating parameter updates; fixed-regularization penalizes parameters irrespective of sensitivity; stream-only meta-learning lacks a feedback controlled parameter update. We propose Meta-Adaptive Network Gradient Optimization (MANGO), an OCL framework that balances stability-plasticity via gradient-gating and meta-learned regularization. Gradient-gating scales parameter updates based on sensitivity, preventing destructive updates. Meta-learned regularization adapts stability coefficients, evaluating the effect of parameter update on replay. In MANGO, replay acts as both a training signal and a forgetting evaluator. We evaluated our method on three standard OCL benchmark datasets. MANGO outperforms strong baselines, achieving state-of-the-art results with consistent performance across replay sizes. In domain incremental learning on CLEAR-10 and class incremental learning on CIFAR-100 and Tiny-ImageNet, it achieves highest accuracy among all baselines and achieves positive Backward Transfer, overcoming forgetting on CLEAR-10.
In many real-world settings, data streams are nonstationary and arrive sequentially, requiring learning systems to adapt continuously without retraining from scratch. Continual learning (CL) addresses this challenge by incorporating new tasks while mitigating catastrophic forgetting, where learning new information degrades performance on previously acquired knowledge. We introduce a control-theoretic perspective on CL that explicitly regulates the evolution of forgetting, framing adaptation as a controlled process subject to long-term stability constraints. We focus on replay-based CL, where a finite memory buffer stores representative samples from prior tasks. We propose COntinual Learning with Drift-Plus-Penalty (COLD), a continual learning framework based on the Drift-Plus-Penalty (DPP) principle from stochastic optimization. To facilitate analysis, we also consider an oracle variant, COLD-ORACLE, as a reference benchmark. At each task, both methods minimize the current task loss while maintaining a virtual queue that tracks deviations from long-term stability on previously learned tasks, capturing the stability-plasticity trade-off as a regulated dynamical process. We establish stability and convergence guarantees that characterize this trade-off through a tunable control parameter. Experiments on standard benchmarks demonstrate that COLD consistently outperforms a broad range of state-of-the-art CL methods while providing competitive and controllable forgetting behavior through explicit regulation of stability and plasticity.
Parameter-efficient continual learning aims to adapt pre-trained models to sequential tasks without forgetting previously acquired knowledge. Most existing approaches treat continual learning as avoiding interference with past updates, rather than considering what properties make the current task-specific update naturally preserve previously acquired knowledge. From a knowledge-decomposition perspective, we observe that low-rank adaptations exhibit highly imbalanced singular value spectra: a few dominant components absorb most of the adaptation energy, thereby (i) more likely to disrupt previously acquired knowledge and (ii) making the update more vulnerable to interference from subsequent tasks. To enable explicit balance among components, we decouple the magnitude of the task update from its directional structure and formulate it as a constrained optimization problem on a restricted Stiefel manifold. We address this problem using a projected first-order method compatible with standard deep-learning optimizers used in vision-language models. Our method mitigates both backward and forward forgetting, consistently outperforming continual learning baselines. The implementation code is available at https://github.com/haodotgu/EBLoRA.
Hao Gu, Mao-Lin Luo, Zi-Hao Zhou +3
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China