EFC++: Elastic Feature Consolidation with Prototype Re-balancing for Cold Start Exemplar-free Incremental Learning
Authors: Simone Magistri, Tomaso Trinci, Albin Soutif-Cormerais, Joost van de Weijer, Andrew D. Bagdanov
Organizations: Media Integration and Communication Center (MICC), University of Florence, Italy · Global Optimization Laboratory, University of Florence, Italy · LAMP Team, Computer Vision Center, Barcelona, Spain · LAMP Team, Computer Vision Center, Universitat Autònoma de Barcelona, Spain
Exemplar-free Class Incremental Learning (EFCIL) aims to learn from a sequence of tasks without having access to previous task data. In this paper, we consider the challenging Cold Start scenario in which insufficient data is available in the first task to learn a high-quality backbone. This is especially challenging for EFCIL since it requires high plasticity, resulting in feature drift which is difficult to compensate for in the exemplar-free setting. To address this problem, we propose an effective approach to consolidate feature representations by regularizing drift in directions highly relevant to previous tasks while employing prototypes to reduce task-recency bias. Our approach, which we call Elastic Feature Consolidation++ (EFC++) exploits a tractable second-order approximation of feature drift based on a proposed Empirical Feature Matrix (EFM). The EFM induces a pseudo-metric in feature space which we use to regularize feature drift in important directions and to update Gaussian prototypes. In addition, we introduce a post-training prototype re-balancing phase that updates classifiers to compensate for feature drift. This strategy allows to improve over our previous EFC method by mitigating the misalignment between stored prototypes and the evolving feature space. Extensive experimental results on Tiny-ImageNet, ImageNet-Subset, ImageNet-1K, and DomainNet show that EFC++ achieves a strong stability--plasticity trade-off in Cold Start and outperforms recent exemplar-free baselines. Code is available at https://github.com/simomagi/elastic_feature_consolidation
Figures & tables
Figure 1: Plasticity potential in Cold and Warm Start. We train a ResNet-18 on C0 classes of CIFAR-100 for C0=10,20,…,50 , and evaluate feature quality via linear probing Davari et al. (2022) on all 100 classes. The plasticity potential Δ , defined as the maximum performance gain which can be obtained with a plastic versus a frozen backbone, is quantified as the performance gap between Joint Training and a model frozen on a subset of classes. Note that ΔCold (the Cold Start scenario) is significantly larger than ΔWarm (the Warm Start case). This is due to the inability to learn a strong feature extractor on only C0 classes and hence greater plasticity is required to incrementally learn new classes. Freezing the backbone in such settings limits adaptability and ultimately constrains performance on subsequent tasks.
Figure 2: Elastic Feature Consolidation with Prototype Re-balancing (EFC++). (a) EFC++ leverages the Empirical Feature Matrix (EFM) to mitigate drift in feature representations by identifying important directions for previous tasks to reduce forgetting while enhancing plasticity for learning new tasks (Section 3.3 ). In this phase, the feature extractor ft and the current task classifier with weights Wt−1:t are trained with EFM regularization and cross-entropy loss (Section 4.3 ) (b) After training, EFC++ uses the EFM to update the prototypes of previous task classes based on the drift induced by the most recent task (Section 4.4 ). (c) EFC++ uses Gaussian prototypes, together with current task features, for training previous and current task classifiers with weights Wt=[Wt−1,Wt−1:t] via a prototype re-balancing phase (Section 4.5 ). (d) Before training on the next task, the new EFM and the prototypes of the current task classes are computed.
Figure 3: The regularizing effects of Et on the Cold Start CIFAR-100 - 10 and 20 step scenarios (see Section 5 for details on dataset settings). Left : Perturbing features in the principal directions of E1 results in significant changes in classifier outputs (in blue ), while perturbations in non-principal directions leave the outputs unchanged (in red). Middle : If we continue incremental learning up through task 3 and perturb features from all three tasks in the principal (solid lines) and non-principal (dashed lines) directions of E3 , we see that E3 captures all important directions in feature space up through task 3. Right : At the end of training, we observe the same behavior: the last per-step accuracy (see Eq. 21 ), representing the average accuracy over all tasks after the last training session, decreases only when perturbed in directions of E10 or E20 relevant for previous tasks in the 10-step and 20-step scenarios, respectively.
Figure 4: Accuracy after each incremental step on the Cold Start CIFAR-100 10-step scenario. Left : In EFC, which combines EFM regularization with the asymmetric PR-ACE loss to balance current task data with prototypes during training, older tasks are forgotten more quickly than more recent ones. Right : EFC++, which applies EFM regularization during backbone training and a post-training prototype re-balancing phase, achieves a better plasticity-stability trade-off.
Figure 5: Average drift of the class means in the relevant directions of the EFM before and after training the task in which they are involved, in both EFC and EFC++ on CIFAR-100 (CS) 10-step. EFC++ consistently exhibits less drift than EFC, especially in the initial tasks, where the drift of the classes is more pronounced (double) for EFC. This confirms that EFC++ better controls the drift along relevant directions.
Dataset
Warm Start (WS)
Cold Start (CS)
10 Step
20 Step
10 Step
20 Step
∣C0∣
∣Ci>0∣
∣C0∣
∣Ci>0∣
∣C0∣
∣Ci>0∣
∣C0∣
∣Ci>0∣
CIFAR-100
50
5
40
3
0
10
0
5
Tiny-ImageNet
100
10
100
5
0
20
0
10
ImageNet-Subset
50
5
40
3
0
10
0
5
ImageNet-1K
500
50
400
30
0
100
0
50
Table 1: Class Distribution in Warm Start and Cold Start: ∣C0∣ is the number of classes in the pre-training phase, while ∣Ci>0∣ represents classes introduced per incremental step.
Warm Start (WS)
Cold Start (CS)
AstepK
AincK
AstepK
AincK
Method
10 Step
20 Step
10 Step
20 Step
10 Step
20 Step
10 Step
20 Step
CIFAR-100
EWC
21.08±1.09
13.53±1.11
41.00±1.11
31.79±2.77
31.17±2.94
17.37±2.43
49.14±1.28
31.02±1.15
LwF
20.73±1.47
11.78±0.60
41.95±1.30
28.93±1.62
32.80±3.08
17.44±0.73
53.91±1.67
38.39±1.05
PASS
53.42±0.48
47.51±0.37
63.42±0.69
59.55±0.97
30.45±1.01
17.44±0.69
47.86±1.93
32.86±1.03
Fusion
56.86
51.75
65.10
61.60
−
−
−
−
Table 2: Small-scale Class-IL experiments. Fusion results are those reported in Toldo and Ozay (2022) as no code is provided to reproduce them. The result for ImageNet-Subset Cold Start for ABD and R-DFCIL are those reported in Gao et al. (2022) .
Figure 6: Accuracy on each task after the final training step on CIFAR-100 Warm Start and Cold Start. EFC++ achieves the best stability-plasticity trade-off compared to the nearest competitor in Warm Start (FeCAM) and Cold Start (R-DFCIL). Task 0 represents the large first task in WS, which is absent in CS (see Table 1 )
ImageNet-1K (CS)
DN4IL (CS)
Method
10 Step
20 Step
6 Step
FeTrIL
34.28±0.20
26.64±0.25
30.95±0.17
FeCAM
36.16±0.10
27.24±0.19
36.42±0.12
R-DFCIL
29.36±0.17
22.30±0.19
22.31±0.21
EFC
42.62±0.08
36.32±0.35
38.27±0.16
EFC++
44.72±0.11
37.46±0.03
39.32±0.15
Table 3: Large scale Class- and Domain-IL Experiments on ImageNet-1K (CS) and DN4IL Gowda et al. (2023) (CS) using AstepK .
Figure 7: Accuracy on each domain after the final training step on DN4IL. Methods that freeze the backbone, such as FeCAM and FeTrIL, achieve higher accuracy on the first domain, while EFC and EFC++ excel at learning new domains.
Dataset
Method
10 Step
20 Step
CIFAR-100
DS-AL ( 2024 )
36.83
28.90
ADC ( 2024 )
46.80
34.69
LDC ( 2025 )
46.60
36.76
DPCR ( 2025 )
50.24
38.98
EFC++
49.47
36.85
Tiny-ImageNet
DS-AL ( 2024 )
27.01
21.86
Table 4: Final per-step accuracy AstepK on CIFAR-100, Tiny-ImageNet, and ImageNet-100 in Class-IL Cold Start. Results are obtained using the DPCR codebase He et al. (2025) , based on the PyCIL framework Zhou et al. (2023a) .
CS - 10 Step
CS - 20 Step
Regularizer
FK(↓)
PLK(↑)
AstepK(↑)
FK(↓)
PLK(↑)
AstepK(↑)
CIFAR-100
E-FIM
67.28
84.27
23.72
71.98
83.22
14.83
KD
34.90
70.67
39.26
41.11
65.41
26.39
FD
12.67
51.83
40.44
0 8.78
35.85
29.02
EFM
16.24
62.72
47.52
18.52
50.79
33.71
Table 5: Ablation study on regularizers. We report the final average forgetting ( FK ), average plasticity ( PLK ), and per-step accuracy ( AstepK ) for CIFAR-100 at 10 and 20 steps.
Figure 8: Comparison of regularization methods on CIFAR-100 (CS) 10-step. EFM balances stability and plasticity, reducing forgetting (left) and improving plasticity (right) over FD. KD and EWC exhibit strong initial plasticity but suffer from severe forgetting after the first task.
Dataset
Step
Update Proto (WS)
Update Proto (CS)
✗
✓
✗
✓
CIFAR-100
10
60.26±0.72
62.15±0.48
46.30±1.45
47.52±0.68
20
56.10±0.34
57.55±0.66
31.71±2.26
33.71±1.41
Tiny-ImageNet
10
50.57±0.40
51.67±0.31
36.17±0.41
37.48±0.52
20
46.86±1.19
50.41±0.51
29.04±0.65
32.56±0.44
ImageNet-Subset
10
69.09±0.30
69.28±0.46
52.95±1.00
53.90±1.18
Table 6: Ablation study on prototype update. The performance of EFC++ improves with prototype updates in both Warm and Cold Start.
Figure 9: Inter-class distance-map analysis on CIFAR-100 Cold Start with 20 steps. We compare the absolute error between the cosine-distance maps computed from fixed prototypes ( Dfixed ) or EFM-updated prototypes ( Dupdated ) and the map computed from real final class prototypes ( Dreal ). EFM-updated prototypes better match the final inter-class geometry and reduce the average off-diagonal error, showing that the proposed update better preserves the feature-space structure across tasks.
Figure 10: Distance between real class means and prototypes when kept fixed (solid) or updated via the Empirical Feature Matrix (dashed) in both Warm and Cold Start scenarios on CIFAR-100 in EFC++. In Cold Start, the average distance of the prototypes from real class means change more than in Warm Start. Our prototype update rule effectively mitigates this drift.
Figure 11: The spectrum of the Empirical Feature Matrix across incremental learning steps. For a better visualization of the spectrum in the analysis we considered a Cold Start 10-step scenario on CIFAR-100. The x -axis is truncated at the 120 th eigenvalue.
Figure 12: A grid search testing 25 combinations of λEFM , with the x -axis in log scale for better visualization. Left : The selected combination, represented by the star symbol, achieves the peak per-step accuracy. Middle : Reducing the value of η weakens regularization and increases forgetting (see Eq. 23 , left). Right : For high values of η , plasticity is significantly compromised (see Eq. 23 , right).
Figure 13: Training times. For EFC++ we report the time required for backbone training, EFM and prototype computation, and the prototype re-balancing phase. Overall, EFC++ exhibits a slight increase in computational time compared to EFC and LwF, shows comparable timings to SSRE, and remains more efficient than PASS, R-DFCIL, and ABD. Timings for R-DFCIL and ABD are omitted as the deep inversion model training time alone already significantly exceeds the time required for the other approaches (about 15 minutes on CIFAR-100).
Step
Mean
Covariance
Accuracy
Delta
10
Fixed
Fixed
46.30±1.45
—
EFM
Fixed
47.52±0.68
0 +1.22 ( Ours )
EFM
Real
47.37±0.19
+1.07
Real
Fixed
49.22±0.17
+2.92
Real
Real
49.33±0.20
+3.03
20
Fixed
Fixed
31.71±2.26
—
Table 7: Impact of Mean and Covariance Drift on EFC++ (CIFAR-100 Cold Start).
Step
Mean
Covariance
Accuracy
Delta
10
Fixed
Fixed
43.65±0.04
—
EFM
Fixed
44.72±0.11
+1.07 ( Ours )
EFM
Real
46.50±0.27
+2.85
Real
Fixed
49.95±0.09
+6.30
Real
Real
49.97±0.09
+6.32
20
Fixed
Fixed
34.00±0.03
—
Table 8: Impact of Mean and Covariance Drift on EFC++ (ImageNet-1K Cold Start).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 14: Accuracy on each task after the final training step on Tiny-ImageNet and ImageNet-Subset Warm Start and Cold Start. EFC++ achieves the best stability-plasticity trade-off compared to the nearest competitor in Warm Start (FeCAM) and Cold Start (R-DFCIL).
AstepK
AincK
Method
10 Step
20 Step
10 Step
20 Step
WS
FeTrIL
55.26±0.16
51.10±0.22
63.42±0.08
60.93±0.12
FeCAM
58.73±0.09
55.67±0.16
65.83±0.14
64.22±0.14
R-DFCIL
35.47±0.19
29.27±0.53
46.42±0.18
42.13±0.58
EFC
59.49±0.07
55.94±0.09
66.64±0.10
64.89±0.10
EFC++
59.09±0.15
54.82±0.17
66.43±0.11
64.45±0.20
Appendix
Table 9: Large scale Class-incremental Experiments. We compare with the state-of-the-art on ImageNet-1K.
Cold Start
AstepK
AincK
Method
6 Step
6 Step
DN4IL
EWC
09.81±0.47
25.55±0.23
LwF
10.50±0.89
27.86±0.56
PASS
24.44±0.19
37.14±0.10
FeTrIL
30.95±0.17
40.39±0.11
Appendix
Table 10: Domain- and Class-incremental Experiments. We compare with the state-of-the-art on DN4IL, a balanced subset of DomainNet for incremental learning Gowda et al. (2023) .
Figure 15: Per-step accuracy during the incremental learning on DN4IL. We compare EFC++ with the state-of-the-art in a setting in which each incremental learning step represents a different domain.
Figure 16: A grid search testing 25 combinations of λEFM , with the x -axis in log scale for better visualization. The y-axes represent per-step accuracy, average forgetting, and average plasticity after the final step.
Step
Mean
Covariance
Accuracy
Delta
10
Fixed
Fixed
52.95±1.00
—
EFM
Fixed
53.90±1.18
+0.95
EFM
Real
52.44±1.21
− 0.51
Real
Fixed
56.14±1.04
+3.19
Real
Real
55.69±1.13
+2.74
20
Fixed
Fixed
39.10±0.86
—
Appendix
Table 11: Impact of Mean and Covariance Drift on EFC++ tested on ImageNet-Subset CS.
Step
Mean
Covariance
Accuracy
Delta
10
Fixed
Fixed
36.17±0.48
—
EFM
Fixed
37.48±0.52
+1.31
EFM
Real
37.47±0.27
+1.30
Real
Fixed
39.94±0.43
+3.77
Real
Real
39.95±0.55
+3.78
20
Fixed
Fixed
29.04±0.65
—
Appendix
Table 12: Impact of Mean and Covariance Drift on EFC++ tested on Tiny-ImageNet CS.
Figure 17: Warm Start (WS) per-step accuracy during the incremental learning. The plots compare recent EFCIL methods against EFC++ on three different datasets for different incremental step sequences.
Figure 18: Cold Start (CS) per-step accuracy plots during incremental learning. The plots compare recent EFCIL methods with EFC++ on three different datasets for different incremental step sequences.
Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a large initial task. Existing cold-start methods typically either train the backbone throughout the stream and compensate for semantic drift, or freeze a backbone after the first task, producing features biased toward the initial classes. These choices also create a computational tension: drift-compensation methods require repeated backbone training and increasingly expensive updates as the task horizon grows, while frozen-backbone methods are cheap but weak under cold start. We study a third option: a feature extractor that is never fit to image data at all. We propose CIRCLE, a class-incremental classifier built from fixed bidirectional two-dimensional reservoir features, adapted from BiRC2D for image classification, and streaming linear discriminant analysis heads. CIRCLE groups multiple random reservoir instantiations into feature ensembles and averages the softmax outputs of independent SLDA heads, yielding a tunable bias-variance tradeoff between richer random features and prediction-level ensembling. Because the feature extractor is fixed and the head admits streaming closed-form updates, CIRCLE performs sample-wise training without replay, task-boundary information, or backbone backpropagation. On CIFAR-100, TinyImageNet, ImageNet-Subset, and ImageNet-1k, CIRCLE is competitive at 10-20 task splits and substantially outperforms strong CS-EFCIL baselines at 50, 100, and 500 task splits, while training much faster than trained-backbone drift-compensation methods. Ablations show that the BiRC2D-style extractor, SLDA head, and balanced feature/prediction ensembling each contribute to the final performance.
Augustinas Jučas, Yangchen Pan
Department of Computer Science University of Oxford · Department of Engineering Sciences University of Oxford
Exemplar-free class-incremental learning (EFCIL) aims to acquire new classes over time without storing raw data. Historically, prototype rehearsal, which samples around stored class prototypes and mixes them with current-task data, has been a popular strategy to reduce catastrophic forgetting. However, recent drift-compensation methods that explicitly realign prototypes in the evolving feature space consistently outperform prototype-based rehearsal, raising the question of whether rehearsal itself is fundamentally limited. We argue that the performance gap stems not from the idea of prototype rehearsal per se, but from how it is typically instantiated: existing approaches treat prototypes as isolated class summaries that ignore information from nearby enemy classes, and fail to correct the emerging class imbalance between a handful of synthetic old-class samples and hundreds of real instances from newly introduced classes. Building on this hypothesis, we revisit prototype rehearsal and propose a manifold-aware variant that restores its competitiveness in EFCIL. First, we introduce Constrained Expansive Over-Sampling, which interpolates each old-class prototype toward its nearest enemy features from new classes, generating boundary-aware rehearsal samples that better follow the underlying data manifold while preserving inter-class separation. Second, we design an Adaptive Class-Balanced loss that performs time-based class weighting, amplifying gradients from older prototypes when they are most informative and gradually annealing their influence as richer supervision from later tasks accumulates. Together, these components turn prototype rehearsal into a drift-resilient, imbalance-aware mechanism that closes, and often reverses, the gap to recent drift-compensation methods, achieving state-of-the-art performance across multiple EFCIL benchmarks.
Hongye Xu, Bartosz Krawczyk
Chester F. Carlson Center for Imaging Science Rochester Institute of Technology
Continual learning (CL) seeks models that acquire new skills without erasing prior knowledge. In exemplar-free class-incremental learning (EFCIL), this challenge is amplified because past data cannot be stored, making representation drift for old classes particularly harmful. Prototype-based EFCIL is attractive for its efficiency, yet prototypes drift as the embedding space evolves; therefore, projection-based drift compensation has become a popular remedy. We show, however, that existing one-directional projections introduce systematic bias: they either retroactively distort the current feature geometry or align past classes only locally, leaving cycle inconsistencies that accumulate across tasks. We introduce BiCyc, a bidirectional projector alignment approach with a cycle-consistency objective. BiCyc jointly optimizes two maps, old-to-new and new-to-old, with stop-gradient gating so that transport and representation co-evolve. Analytically, we show that the cycle loss contracts the singular spectrum toward unity in whitened space, and that improved transport of class means and covariances yields smaller perturbations of classification log-odds, preserving old-class decisions and mitigating catastrophic forgetting. Empirically, across standard EFCIL benchmarks, BiCyc substantially reduces forgetting and improves accuracy in from-scratch settings, while remaining competitive in the pretrained fine-grained regime.
Hongye Xu, Bartosz Krawczyk
Chester F. Carlson Center for Imaging Science · Rochester Institute of Technology