Organizations: School of Future Technology, South China University of Technology, Guangzhou, China · Shien-Ming Wu School of Intelligent Engineering, South China University of Technology, Guangzhou, China
Continual learning (CL) agents incrementally learn from sequentially arriving data and adapt to the dynamic, ever-changing nature of real-world environments. However, many existing CL methods suffer performance degradation in evolving, imbalanced data streams due to limited adaptation to changing class frequencies or ineffective use of mixed data from new and previously observed classes. To deal with these challenges, we propose an analytic imbalance rectifier (AIR) algorithm for real-world CL. AIR is an online exemplar-free approach with a frozen backbone as the feature extractor and a closed-form incremental classifier whose weight equals the joint-learning weight for the same class-weighted ridge objective. AIR addresses class imbalance with an analytic reweighting module (ARM) that calculates a reweighting factor for each class in the loss function to equalize total sample weights across classes. Under long-tailed class-incremental learning, AIR leads 28 baselines in aggregate accuracy and exemplar-free methods in aggregate macro F1, gaining 3.21% accuracy and 2.14% macro F1 over the respective strongest exemplar-free baselines. Under the Si-Blurry setting with recurring classes, AIR leads 15 exemplar-based and exemplar-free baselines, gaining 2.32% aggregate accuracy and 1.27% aggregate macro F1 over the strongest baseline. One-sided paired tests support positive mean absolute gains in these four comparisons (Holm-adjusted p<0.006).
Figures & tables
Method
Score
CIFAR-100 (LT)
ImageNet-R (LT)
Time
Ascending
Descending
Shuffled
Ascending
Descending
Shuffled
Aavg
Alast
Aavg
Alast
Aavg
Alast
Aavg
Alast
Aavg
Alast
Aavg
Alast
DGR [ 11 ]
61.01 ± 0.45
47.49 ± 0.77
57.85 ± 0.24
85.67 ± 0.27
65.07 ± 0.44
67.67 ± 3.45
60.09 ± 1.66
44.14 ± 1.54
54.35 ± 0.77
72.12 ± 0.35
56.55 ± 0.32
63.60 ± 1.31
57.50 ± 0.39
9.19
MVP-R [ 27 ]
61.39 ± 0.66
64.49 ± 0.57
54.57 ± 0.59
83.62 ± 0.25
70.99 ± 0.68
67.60 ± 2.88
59.39 ± 2.88
51.16 ± 0.58
46.42 ± 0.50
69.46 ± 0.22
58.96 ± 0.36
59.59 ± 1.33
50.42 ± 0.70
6.84
CLIB [ 19 ]
62.48 ± 0.55
62.86 ± 1.08
62.36 ± 0.79
84.15 ± 0.30
70.03 ± 0.26
71.63 ± 2.41
65.69 ± 1.50
54.84 ± 0.59
47.56 ± 0.40
66.17 ± 0.29
53.51 ± 0.34
60.34 ± 1.00
50.67 ± 0.64
0.51
PRS [ 17 ]
64.21 ± 0.61
49.02 ± 0.67
66.42 ± 0.77
81.29 ± 0.59
64.45 ± 0.60
70.46 ± 2.53
69.81 ± 1.97
46.39 ± 0.68
56.60 ± 1.72
73.66 ± 0.39
66.45 ± 0.50
64.60 ± 1.30
61.34 ± 0.89
7.20
Table 1: LT-CIL accuracy (%). The first block uses exemplars, with fixed-memory methods before growing-memory methods: fixed memories retain 500 samples on CIFAR-100 or 1,000 on ImageNet-R, and growing memories retain 5/class. The second block and AIR are exemplar-free. Bold / underlined mark the best/second-best exemplar-free results. Values are mean ± standard error over 6 runs. Time is the mean training time in s/1,000 samples across datasets and orders (lower is better).
Method
Score
CIFAR-100 (LT)
ImageNet-R (LT)
Time
Ascending
Descending
Shuffled
Ascending
Descending
Shuffled
Favg
Flast
Favg
Flast
Favg
Flast
Favg
Flast
Favg
Flast
Favg
Flast
DGR [ 11 ]
59.41 ± 0.47
44.42 ± 0.90
55.22 ± 0.27
85.39 ± 0.27
62.13 ± 0.49
65.24 ± 3.64
57.63 ± 1.45
42.68 ± 1.37
52.31 ± 0.78
71.66 ± 0.42
56.38 ± 0.24
63.00 ± 1.44
56.81 ± 0.41
9.19
MVP-R [ 27 ]
61.08 ± 0.74
62.66 ± 0.76
51.87 ± 0.47
83.02 ± 0.25
67.98 ± 0.70
64.74 ± 3.35
57.57 ± 2.91
53.49 ± 0.52
47.10 ± 0.56
71.13 ± 0.27
59.36 ± 0.37
61.12 ± 1.24
52.94 ± 0.77
6.84
CLIB [ 19 ]
63.91 ± 0.63
61.44 ± 1.30
64.25 ± 0.71
84.02 ± 0.30
67.87 ± 0.35
69.49 ± 2.96
65.64 ± 1.51
58.17 ± 0.58
54.47 ± 0.52
68.17 ± 0.36
55.13 ± 0.29
62.63 ± 0.92
55.59 ± 0.55
0.51
PRS [ 17 ]
64.04 ± 0.65
47.39 ± 0.77
66.78 ± 0.66
79.91 ± 0.63
59.03 ± 0.79
68.05 ± 2.85
69.37 ± 1.98
49.75 ± 0.71
61.61 ± 1.39
73.86 ± 0.38
62.43 ± 0.26
66.66 ± 1.17
63.63 ± 0.67
7.20
Table 2: LT-CIL macro F1 (%). The experimental settings, layout, methods, memory grouping, seeds, and timing convention follow Table 1 . Score is the unweighted mean of Favg and Flast over both datasets and all three orders. Bold / underlined mark the best/second-best exemplar-free results; values are mean ± standard error over six runs.
Method
Type
Score
CIFAR-100
ImageNet-R
CORe50
CUB-200-2011
Time
AUC
Avg
Last
AUC
Avg
Last
AUC
Avg
Last
AUC
Avg
Last
O-LoRA [ 35 ]
A
54.95 ± 0.98
68.15 ± 2.52
68.02 ± 1.83
68.20 ± 2.69
61.45 ± 0.98
60.85 ± 1.06
53.54 ± 1.70
47.06 ± 2.17
50.73 ± 3.33
50.84 ± 3.78
40.54 ± 0.71
47.50 ± 1.66
42.60 ± 1.07
5.10
F
50.54 ± 1.14
64.44 ± 3.07
65.09 ± 2.25
65.83 ± 2.98
59.57 ± 1.18
59.49 ± 1.14
53.91 ± 1.50
40.23 ± 2.62
43.44 ± 4.13
44.41 ± 4.56
33.30 ± 0.95
40.97 ± 2.17
35.80 ± 1.43
UniCIL [ 37 ]
A
57.39 ± 4.70
78.00 ± 0.97
72.96 ± 1.45
85.38 ± 0.34
37.34 ± 15.94
38.28 ± 16.41
38.83 ± 16.87
49.06 ± 16.10
49.84 ± 16.21
49.93 ± 16.92
62.77 ± 2.05
63.29 ± 2.97
63.02 ± 4.83
7.95
F
54.42 ± 4.78
75.27 ± 1.47
69.64 ± 1.60
85.35 ± 0.32
35.47 ± 15.85
36.26 ± 16.20
37.91 ± 16.93
46.49 ± 16.42
47.47 ± 16.76
48.78 ± 17.38
56.09 ± 2.53
56.39 ± 3.43
57.90 ± 5.94
EWC++ [ 18 ]
A
72.55 ± 0.48
72.44 ± 1.22
70.95 ± 0.86
69.91 ± 0.79
61.96 ± 0.76
61.62 ± 1.19
54.16 ± 0.58
81.28 ± 0.68
81.81 ± 0.67
82.54 ± 1.29
77.49 ± 1.02
81.10 ± 0.98
75.28 ± 0.51
0.50
Table 3: Si-Blurry accuracy ( A ) and macro F1 ( F ), in % (mean ± standard error, 6 runs). The first block retains samples, ordered by increasing Accuracy Score; the rest are exemplar-free. Fixed memories hold 500/1,000/250/1,000 samples in dataset-column order; growing memories hold 5/class. O-LoRA retains four hard samples only for MAS estimation. Bold / underlined : best/second-best exemplar-free results. Time is the mean training time in s/1,000 samples across datasets (lower is better).
Method
Ascending
Descending
Shuffled
Aavg
Alast
Aavg
Alast
Aavg
Alast
ACIL w/o ARM
73.28 ± 0.71
62.00 ± 0.44
83.99 ± 0.20
62.00 ± 0.44
69.66 ± 1.45
62.00 ± 0.44
ACIL w/ ARM
77.21 ± 0.74 (+3.93 ± 0.09)
72.60 ± 0.57 (+10.60 ± 0.31)
87.93 ± 0.14 (+3.94 ± 0.13)
72.60 ± 0.57 (+10.60 ± 0.31)
77.13 ± 1.64 (+7.47 ± 0.23)
72.60 ± 0.57 (+10.60 ± 0.31)
DS-AL w/o ARM
73.90 ± 0.75
62.66 ± 0.46
84.70 ± 0.21
64.41 ± 0.52
70.97 ± 1.51
63.63 ± 0.56
DS-AL w/ ARM
75.61 ± 0.77 (+1.71 ± 0.04)
69.32 ± 0.57 (+6.66 ± 0.26)
86.73 ± 0.16 (+2.03 ± 0.09)
69.88 ± 0.56 (+5.47 ± 0.21)
74.58 ± 1.65 (+3.61 ± 0.19)
69.60 ± 0.59 (+5.97 ± 0.20)
Table 4: Accuracy (%) after adding ARM to ACIL and DS-AL on CIFAR-100 LT-CIL ( ρ=500 ). Leading values are six-seed means ± standard errors; parentheses on the ARM rows report the paired improvement as mean ± standard error over the same seeds.
Method
A
C
Wk−1
Total
ACIL & GACL
Θ(f2)
-
Θ(fCk)
Θ(f2+fCk)
RanPAC
Θ(f2)
Θ(fCk)
Θ(fCk)
Θ(f2+fCk)
AIR
Θ(f2)
Θ(fCk)
Θ(fCk)
Θ(f2+fCk)
G-AIR
Θ(f2Ck)
Θ(fCk)
Θ(fCk)
Θ(f2Ck+fCk)
Table 5: Storage costs of retained matrices A , C , and Wk−1 .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Setting
Phase-end
comparisons
maxϵW
Final accuracy (%)
CL
Joint
LT-CIL ascending
40
5.02×10−14
76.11
76.11
LT-CIL descending
40
4.58×10−14
76.11
76.11
LT-CIL shuffled
40
6.03×10−14
76.11
76.11
Si-Blurry
10
5.81×10−14
84.31
84.31
Appendix
Table A.1: Joint–continual equivalence on cached CIFAR-100 features. Comparison counts equal phases × two seeds; errors are maxima across comparisons. Final test accuracies average the two seeds for continual (CL) and joint learning.
Continual learning (CL) evaluates adaptability in learning solutions to retain knowledge. Our research addresses the challenge of catastrophic forgetting, where models lose proficiency in previously learned tasks as they acquire new ones. While numerous solutions have been proposed, existing experimental setups often rely on idealized class-incremental learning scenarios. We introduce Realistic Continual Learning (RealCL), a novel CL paradigm where class distributions across tasks are random. We also present CLARE (Continual Learning Approach with pRE-trained models for RealCL scenarios), a pre-trained model-based solution designed to integrate new knowledge while preserving past learning. Our contributions include pioneering RealCL as a generalization of traditional CL setups, proposing CLARE as an adaptable approach for RealCL tasks, and conducting extensive experiments demonstrating its effectiveness across various RealCL scenarios. Notably, CLARE outperforms existing models on RealCL benchmarks, highlighting its versatility in unpredictable learning environments. Code to reproduce all our experiments can be found at https://github.com/gramuah/clare.
Nadia Nasri, Carlos Gutiérrez-Álvarez, Sergio Lafuente-Arroyo +2
Continual learning must balance the learning of new knowledge with the retention of previously learned knowledge to incrementally learn tasks from a data stream without catastrophic forgetting. While leveraging pretrained models has significantly advanced continual learning, existing methods exhibit a scalability bottleneck when trained sequentially on many tasks, suffering from performance degradation due to inter-task interference and loss of plasticity. Inspired by evidence that sparse fine-tuning achieves performance comparable to full fine-tuning, this paper presents a novel sparsity-driven continual learning framework. Our continual learning method, termed CLARE, operates in two stages: it first identifies a sparse, task-critical parameter mask via a sparsity-inducing objective, then performs mask-constrained fine-tuning by only optimizing parameters selected by the mask. This two-stage sparse adapter mechanism enables all tasks to be accumulated within a shared adapter space while reducing destructive interference across tasks. Extensive experiments demonstrate the scalability of CLARE. On the long task-sequence benchmark Omnibenchmark-1k, CLARE outperforms strong baselines in final accuracy by a large margin, e.g, improving EASE by 4.64% and 13.34% after learning 100 tasks, respectively.
Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a large initial task. Existing cold-start methods typically either train the backbone throughout the stream and compensate for semantic drift, or freeze a backbone after the first task, producing features biased toward the initial classes. These choices also create a computational tension: drift-compensation methods require repeated backbone training and increasingly expensive updates as the task horizon grows, while frozen-backbone methods are cheap but weak under cold start. We study a third option: a feature extractor that is never fit to image data at all. We propose CIRCLE, a class-incremental classifier built from fixed bidirectional two-dimensional reservoir features, adapted from BiRC2D for image classification, and streaming linear discriminant analysis heads. CIRCLE groups multiple random reservoir instantiations into feature ensembles and averages the softmax outputs of independent SLDA heads, yielding a tunable bias-variance tradeoff between richer random features and prediction-level ensembling. Because the feature extractor is fixed and the head admits streaming closed-form updates, CIRCLE performs sample-wise training without replay, task-boundary information, or backbone backpropagation. On CIFAR-100, TinyImageNet, ImageNet-Subset, and ImageNet-1k, CIRCLE is competitive at 10-20 task splits and substantially outperforms strong CS-EFCIL baselines at 50, 100, and 500 task splits, while training much faster than trained-backbone drift-compensation methods. Ablations show that the BiRC2D-style extractor, SLDA head, and balanced feature/prediction ensembling each contribute to the final performance.