Continual learning enables vision systems to adapt to ever-changing data distributions. Despite significant advances, existing approaches fail to capture continuous and concurrent shifts in classes and domains, a critical capability for real-world deployment. This work introduces Online VIL (Online Versatile Incremental Learning), a novel scenario where class concepts and visual domains evolve simultaneously online without explicit boundaries. To better adapt to the challenges of such dynamic environments that more closely resemble real-world conditions, we propose a novel framework TopFlow, Topology preservation with Flow matching representation that contains two complementary mechanisms: Domain-agnostic Flow Matching (DFM) and Global Topology Preservation (GTP). DFM guides the model to have domain-agnostic representations by integrating the geodesic flow kernel into contrastive learning. In contrast, GTP maintains the global structure of the feature space without explicitly storing past examples. Our extensive experiments demonstrate that TopFlow effectively addresses the limitations of existing methods within the Online VIL scenario, achieving state-of-the-art performance in challenging Online VIL. The proposed methods suggest potential directions for building continual learning systems in realistic dynamic environments. Our implementation code is available at https://github.com/KU-VGI/Online-VIL.
Figures & tables
Figure 1 : Conceptual comparison of (a) Class Incremental Learning (CIL), (b) Domain Incremental Learning (DIL), (c) Versatile Incremental Learning (VIL), and (d) Online Versatile Incremental Learning (Online VIL, Ours). Red lines indicate the scopes observable at once, while black dashed lines depict transitions across time.
Figure 2 : Analysis of layer-wise knowledge on pre-trained ViT for data with changing distributions (CORe50). (a) t-SNE visualization of features from certain layers of the pre-trained ViT. The color indicates the domain. (b) The t-SNE visualization of the last layer with class-wise color. (c) Accuracy of linear probing for class and domain classification from different layers.
Figure 3 : Architecture overview. TopFlow comprises two components: Domain-agnostic Flow Matching (DFM) and Global Topology Preservation (GTP). DFM promotes domain-agnostic representation accumulation, while GTP preserves the global feature topology.
Method
iDigits
CORe50
CLEAR100
AAUC
ALast
AAUC
ALast
AAUC
ALast
Upper-bound
-
87.18 ± 0.13
-
91.66 ± 0.25
-
94.36 ± 0.28
Lower-bound
13.45 ± 0.62
12.71 ± 3.34
3.39 ± 0.10
3.24 ± 1.09
2.38 ± 0.32
2.66 ± 0.55
EWC [ 11 ]
20.07 ± 2.84
14.67 ± 3.61
18.60 ± 4.64
16.07 ± 1.56
23.61 ± 3.11
19.93 ± 1.33
LwF [ 16 ]
19.61 ± 3.50
15.38 ± 1.04
25.27 ± 3.77
21.97 ± 4.22
23.80 ± 2.24
21.70 ± 4.19
CODA-P [ 27 ]
23.96 ± 4.74
20.62 ± 3.47
54.06 ± 5.32
48.88 ± 2.92
28.82 ± 5.77
25.61 ± 3.52
Table 1 : Experimental results with proposed Online VIL scenarios. We used bold and underlined as brief indications of the best and the second best, respectively.
Buffer Size
Method
iDigits
CORe50
CLEAR100
AAUC
ALast
AAUC
ALast
AAUC
ALast
500
ER [ 25 ]
59.43 ± 6.24
48.70 ± 2.51
74.77 ± 4.85
72.25 ± 2.27
73.92 ± 3.93
71.49 ± 3.20
RM [ 1 ]
55.02 ± 5.37
51.73 ± 2.05
81.06 ± 3.90
70.41 ± 3.17
72.42 ± 4.64
72.93 ± 2.05
CLIB [ 12 ]
57.38 ± 4.16
52.63 ± 3.38
75.06 ± 5.81
71.93 ± 1.06
68.39 ± 5.25
66.92 ± 1.52
OCM [ 9 ]
57.40 ± 3.60
52.88 ± 2.52
75.29 ± 3.10
72.66 ± 1.93
77.80 ± 3.25
75.10 ± 2.42
CBA [ 33 ]
58.05 ± 4.39
54.28 ± 3.07
81.92 ± 4.04
81.02 ± 1.39
75.26 ± 4.08
74.47 ± 1.82
Table 2 : Results of OnlineVIL scenarios using replay buffer sizes 500 and 2000.
DFM
GTP
AAUC
ALast
Baseline
58.30
52.84
✓
64.13
64.57
✓
63.95
65.28
✓
✓
64.51
66.20
Table 3 : Ablation study for DFM and GTP on CORe50.
Method
ALast
CODA-P
49.13
+ DFM
52.88
+ GTP
53.17
+ DFM, GTP
54.62
PEC
45.99
+ DFM
46.64
Table 6 : Performance comparison of different methods and variants.
Method
iDigits
CORe50
DomainNet
Avg. Acc ↑
Forgetting ↓
Avg. Acc ↑
Forgetting ↓
Avg. Acc ↑
Forgetting ↓
ICON
70.67 ± 2.13
13.49 ± 2.09
79.61 ± 1.73
8.19 ± 1.47
49.08 ± 2.11
17.88 ± 2.40
TopFlow (Ours)
73.33 ± 1.72
12.14 ± 4.20
79.77 ± 1.47
8.03 ± 1.19
49.62 ± 1.82
17.30 ± 3.09
Table 7 : The Effectiveness of TopFlow in Offline VIL Scenarios.
Figure 4 : t-SNE visualization of output features over tasks with GTP and without GTP.