Continual learning enables vision systems to adapt to ever-changing data distributions. Despite significant advances, existing approaches fail to capture continuous and concurrent shifts in classes and domains, a critical capability for real-world deployment. This work introduces Online VIL (Online Versatile Incremental Learning), a novel scenario where class concepts and visual domains evolve simultaneously online without explicit boundaries. To better adapt to the challenges of such dynamic environments that more closely resemble real-world conditions, we propose a novel framework TopFlow, Topology preservation with Flow matching representation that contains two complementary mechanisms: Domain-agnostic Flow Matching (DFM) and Global Topology Preservation (GTP). DFM guides the model to have domain-agnostic representations by integrating the geodesic flow kernel into contrastive learning. In contrast, GTP maintains the global structure of the feature space without explicitly storing past examples. Our extensive experiments demonstrate that TopFlow effectively addresses the limitations of existing methods within the Online VIL scenario, achieving state-of-the-art performance in challenging Online VIL. The proposed methods suggest potential directions for building continual learning systems in realistic dynamic environments. Our implementation code is available at https://github.com/KU-VGI/Online-VIL.
Figures & tables
Figure 1 : Conceptual comparison of (a) Class Incremental Learning (CIL), (b) Domain Incremental Learning (DIL), (c) Versatile Incremental Learning (VIL), and (d) Online Versatile Incremental Learning (Online VIL, Ours). Red lines indicate the scopes observable at once, while black dashed lines depict transitions across time.
Figure 2 : Analysis of layer-wise knowledge on pre-trained ViT for data with changing distributions (CORe50). (a) t-SNE visualization of features from certain layers of the pre-trained ViT. The color indicates the domain. (b) The t-SNE visualization of the last layer with class-wise color. (c) Accuracy of linear probing for class and domain classification from different layers.
Figure 3 : Architecture overview. TopFlow comprises two components: Domain-agnostic Flow Matching (DFM) and Global Topology Preservation (GTP). DFM promotes domain-agnostic representation accumulation, while GTP preserves the global feature topology.
Method
iDigits
CORe50
CLEAR100
AAUC
ALast
AAUC
ALast
AAUC
ALast
Upper-bound
-
87.18 ± 0.13
-
91.66 ± 0.25
-
94.36 ± 0.28
Lower-bound
13.45 ± 0.62
12.71 ± 3.34
3.39 ± 0.10
3.24 ± 1.09
2.38 ± 0.32
2.66 ± 0.55
EWC [ 11 ]
20.07 ± 2.84
14.67 ± 3.61
18.60 ± 4.64
16.07 ± 1.56
23.61 ± 3.11
19.93 ± 1.33
LwF [ 16 ]
19.61 ± 3.50
15.38 ± 1.04
25.27 ± 3.77
21.97 ± 4.22
23.80 ± 2.24
21.70 ± 4.19
CODA-P [ 27 ]
23.96 ± 4.74
20.62 ± 3.47
54.06 ± 5.32
48.88 ± 2.92
28.82 ± 5.77
25.61 ± 3.52
Table 1 : Experimental results with proposed Online VIL scenarios. We used bold and underlined as brief indications of the best and the second best, respectively.
Buffer Size
Method
iDigits
CORe50
CLEAR100
AAUC
ALast
AAUC
ALast
AAUC
ALast
500
ER [ 25 ]
59.43 ± 6.24
48.70 ± 2.51
74.77 ± 4.85
72.25 ± 2.27
73.92 ± 3.93
71.49 ± 3.20
RM [ 1 ]
55.02 ± 5.37
51.73 ± 2.05
81.06 ± 3.90
70.41 ± 3.17
72.42 ± 4.64
72.93 ± 2.05
CLIB [ 12 ]
57.38 ± 4.16
52.63 ± 3.38
75.06 ± 5.81
71.93 ± 1.06
68.39 ± 5.25
66.92 ± 1.52
OCM [ 9 ]
57.40 ± 3.60
52.88 ± 2.52
75.29 ± 3.10
72.66 ± 1.93
77.80 ± 3.25
75.10 ± 2.42
CBA [ 33 ]
58.05 ± 4.39
54.28 ± 3.07
81.92 ± 4.04
81.02 ± 1.39
75.26 ± 4.08
74.47 ± 1.82
Table 2 : Results of OnlineVIL scenarios using replay buffer sizes 500 and 2000.
DFM
GTP
AAUC
ALast
Baseline
58.30
52.84
✓
64.13
64.57
✓
63.95
65.28
✓
✓
64.51
66.20
Table 3 : Ablation study for DFM and GTP on CORe50.
Method
ALast
CODA-P
49.13
+ DFM
52.88
+ GTP
53.17
+ DFM, GTP
54.62
PEC
45.99
+ DFM
46.64
Table 6 : Performance comparison of different methods and variants.
Method
iDigits
CORe50
DomainNet
Avg. Acc ↑
Forgetting ↓
Avg. Acc ↑
Forgetting ↓
Avg. Acc ↑
Forgetting ↓
ICON
70.67 ± 2.13
13.49 ± 2.09
79.61 ± 1.73
8.19 ± 1.47
49.08 ± 2.11
17.88 ± 2.40
TopFlow (Ours)
73.33 ± 1.72
12.14 ± 4.20
79.77 ± 1.47
8.03 ± 1.19
49.62 ± 1.82
17.30 ± 3.09
Table 7 : The Effectiveness of TopFlow in Offline VIL Scenarios.
Figure 4 : t-SNE visualization of output features over tasks with GTP and without GTP.
Continual learning enables vision-language models to accumulate knowledge and adapt to evolving tasks without retraining from scratch. However, in multi-domain task-incremental learning, large domain shifts intensify the stability-plasticity dilemma. Most existing methods rely on fixed architectures with statically allocated parameters, which limits adaptation to new domains and aggravates catastrophic forgetting. To address these challenges, we propose DIMoE-Adapters, a Dynamic Incremental Mixture-of-Experts Adapters framework that introduces a dynamic expert evolution paradigm to balance stability and plasticity. This paradigm is implemented through two collaborative components: Self-Calibrated Expert Evolution (SCEE) and Prototype-Guided Expert Selection (PGES). SCEE constructs and evolves a sparse expert pool through expert optimization dynamics, improving plasticity while reducing redundant capacity. PGES controls expert utilization based on the pool shaped by SCEE, improving stability across both previously encountered and unseen tasks. Extensive experiments show that DIMoE-Adapters outperforms previous state-of-the-art methods across various settings.
Mengxin Qin, Xiang Zhang, Xi Wang +3
School of Electronic Engineering, Xidian University Xi’an 710071, China
Continual learning requires models to adapt to new domains and new classes while retaining prior knowledge. Many existing methods rely on continued optimization, using regularization, replay, or parameter expansion to prevent new updates from overwriting previously learned knowledge. Instead, we propose replacing continual training with continual inference: a PFN-based model that is meta-trained, and then frozen, adapting to new classes only by extending an in-context evidence set. Our model, Latent Concept PFN, performs in-context Bayesian inference over a latent concept space that captures semantic structure shared across domains and classes. As each new domain or class arrives, exemplars are added to the memory; adaptation reflects updated posterior beliefs over latent concepts rather than gradient updates. No parameters are changed, reducing forgetting. The same method handles both domain and class incremental continual learning without task identity. Concept annotations are only used during meta-training, acting as a soft anchor on the latent space rather than a fixed bottleneck. Unlike fixed-vocabulary concept methods, the model also handles noisy, ambiguous, or incomplete annotations by combining concept labels with raw input evidence to discover distinctions beyond the predefined concept set. Experiments on class and domain incremental learning datasets demonstrate competitive continual learning performance while learning interpretable latent concepts.
In this work we introduce a novel approach to domain incremental learning, adapting models over time to evolving, non-stationary data. In contrast to other works, we do not attempt to avoid catastrophic forgetting, but rather allow it and exploit it. Our model combines a main task head with a self-supervised masked autoencoder (MAE) head. We then learn domain-specific LoRA adapters during incremental training. Each adapter specializes to its domain, naturally inducing forgetting on other domains in both heads. At inference, we perform online test-time training on the self-supervised MAE head to identify which LoRAs best matches the current input, so the model can `remember' the domain again. Our scheme is especially well-suited to real-world streaming data, such as video, where consecutive samples are highly correlated and domain shifts are gradual. We demonstrate our method on domain-incremental action recognition and semantic segmentation tasks.