Federated Learning (FL) aims at unburdening the training of deep models by distributing computation across multiple devices (clients) while safeguarding data privacy. On top of that, Federated Continual Learning (FCL) also accounts for data distribution evolving over time, mirroring the dynamic nature of real-world environments. While previous studies have identified Catastrophic Forgetting and Client Drift as major factors of performance degradation in FCL, we shed light on the importance of Incremental Bias and Federated Bias, which cause models to prioritize classes that are recently introduced or locally predominant, respectively. Our proposal constrains both biases to the last layer by efficiently fine-tuning a pre-trained backbone using learnable prompts, resulting in clients that produce less biased representations and more biased classifiers. Therefore, instead of solely relying on parameter aggregation, we leverage generative prototypes to effectively balance the predictions of the global model. Our proposed methodology significantly improves the current state of the art across six datasets, each including three different scenarios.
Figures & tables
Figure 1 : Federated bias. Histogram of the clients’ responses on the global test set ( β=0.05 ) (left). Entropy of the response histograms, averaged on all clients, compared with FL performance (right).
Figure 2 : Classifier Rebalancing procedure through hierarchical sampling.
CIFAR-100
ImageNet-R
ImageNet-A
Joint
92.75
84.02
54.64
Partition β
0.5
0.1
0.05
0.5
0.1
0.05
1.0
0.5
0.2
EWC
78.46
72.42
64.51
58.93
48.15
43.68
10.86
10.07
8.89
LwF
62.87
55.56
47.09
54.03
41.02
46.07
8.89
8.89
7.90
FisherAVG
76.10
74.43
65.31
58.68
50.82
47.33
11.59
11.06
10.14
RegMean
59.80
45.88
39.08
61.18
57.00
55.80
8.56
6.22
4.34
Table 1 : CIFAR-100, ImageNet-R and ImageNet-A . Results in terms of FAA [↑] . Best are highlighted in bold, second-best underlined.
EuroSAT
Cars-196
CUB-200
Joint
98.42
85.62
86.04
Partition β
1.0
0.5
0.2
1.0
0.5
0.2
1.0
0.5
0.2
EWC
64.12
59.30
56.52
19.55
18.02
18.29
31.46
29.60
27.89
LwF
31.91
21.26
31.42
20.84
22.72
31.76
25.25
21.11
18.54
FisherAVG
58.84
59.94
55.86
26.03
24.60
21.58
30.45
28.39
25.06
RegMean
48.74
51.73
45.27
21.83
20.36
15.92
35.57
32.84
32.83
Table 2 : EuroSAT, Cars-196 and CUB-200 . Results in terms of FAA [↑] . Best are highlighted in bold, second-best underlined.
Figure 3 : Prompting vs. fine-tuning. Average pairwise distance of the local prototypes on all clients (left). FL performance before and after Classifier Rebalancing (CR), for β=0.05 (right).
Prompt
CR old
CR cur
C-100
IN-R
CUB-200
✗
✗
✗
30.58
26.42
25.70
✗
✓
✗
62.78
59.42
61.09
✗
✗
✓
65.30
60.38
56.33
✗
✓
✓
82.52
64.91
70.59
✓
✗
✗
52.29
30.28
26.74
✓
✓
✗
81.91
66.70
69.51
Table 3 : HGP components . Evaluation of the impact of each component of HGP. Presented in terms of FAA [↑] .
Figure 4 : FAA [%] in relation with the communication cost [MB] for all tested approaches on ImageNet-R (left) and CIFAR-100 (center) and CUB-200 (right).
Figure 5 : Real training image (left) and its reconstructions leveraging either the input image (center) or the generative prototype of the related class (right).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
CIFAR-100
ImageNet-R
ImageNet-A
Distrib. β
0.5
0.1
0.05
0.5
0.1
0.05
1.0
0.5
0.2
EWC
1.38
2.32
2.67
1.69
1.14
1.14
1.34
0.75
1.38
LwF
1.15
2.38
4.25
1.91
0.93
1.38
0.30
0.27
1.61
FisherAVG
0.40
3.67
3.18
0.77
0.47
0.93
1.59
1.36
1.49
RegMean
0.01
1.78
4.87
1.10
0.85
1.44
0.59
1.10
0.70
CCVR
1.13
2.49
1.39
0.57
0.49
2.29
1.19
0.63
2.26
Appendix
Table A : Standard deviations (reported in FAA) for CIFAR-100, ImageNet-R and ImageNet-A.
EuroSAT
Cars-196
CUB-200
Distrib. β
1.0
0.5
0.2
1.0
0.5
0.2
1.0
0.5
0.2
EWC
7.33
5.78
6.60
1.72
0.46
1.00
0.55
0.94
1.68
LwF
3.32
4.58
5.12
1.35
2.81
2.01
2.07
2.15
2.02
FisherAVG
5.43
6.07
3.40
2.10
1.78
0.78
0.26
1.35
2.21
RegMean
3.99
7.21
5.68
0.23
1.11
1.80
1.79
2.63
2.27
CCVR
9.01
7.16
6.26
1.49
0.87
2.08
1.22
1.69
1.73
Appendix
Table B : Standard deviations (reported in FAA) for EuroSAT, Cars-196, and CUB-200.
Figure A : Federated bias. Entropy of the response histograms, averaged on all clients, compared with FL performance, for ImageNet-R (left) and ImageNet-A (right). The experimental setting follows Section 2 .
Figure B : Accuracies on past vs. current tasks for CIFAR-100 (left) and ImageNet-R (right).
Figure C : Accuracies on past vs. current tasks for ImageNet-A (left) and EuroSAT (right).
Figure D : Accuracies on past vs. current tasks for Cars-196 (left) and CUB200 (right).
Attack
Mean AUROC
Weighted AUROC
Mean AUPRC
Mean Lift
TPR@1%FPR
Gaussian
0.762
0.763
0.542
2.522
0.220
Mahalanobis
0.762
0.763
0.542
2.522
0.220
Cosine
0.846
0.843
0.647
3.036
0.301
Appendix
Table C : ImageNet-A Defense Summary. Despite some ranking signal, strict extraction (TPR@1%FPR) remains difficult.
Attack
Mean AUROC
Weighted AUROC
Mean AUPRC
Mean Lift
TPR@1%FPR
Gaussian
0.558
0.546
0.154
1.439
0.034
Mahalanobis
0.558
0.546
0.154
1.439
0.034
Cosine
0.541
0.535
0.144
1.346
0.027
Appendix
Table D : EuroSAT Defense Summary. HGP reduces attacker metrics to near-random levels.
Attack
Mean AUROC
Weighted AUROC
Mean AUPRC
Mean Lift
TPR@1%FPR
Gaussian
0.569
0.541
0.198
1.372
0.025
Mahalanobis
0.569
0.541
0.198
1.372
0.025
Cosine
0.551
0.536
0.197
1.351
0.024
Appendix
Table E : CIFAR-100 Defense Summary. Attacks fail to recover meaningful membership signal under strict budgets.
Covariance
CIFAR-100
ImageNet-R
ImageNet-A
EuroSAT
Cars-196
CUB-200
Diagonal (ours)
90.39
72.64
41.61
87.97
52.13
79.71
Full (Tikhonov)
90.03
72.77
40.11
88.53
51.73
78.89
Appendix
Table F : Diagonal vs. full covariance. FAA (reported for β equal to the least imbalanced setting of each dataset) of HGP when each generative prototype Nm,c uses a diagonal covariance (default) or a full covariance matrix estimated with strong Tikhonov regularization. The two variants are comparable across all datasets, while the full covariance incurs substantially higher memory and compute.
Figure E : Feature visualizations for CIFAR-100. t-SNE (left) and UMAP (right) projections comparing real ViT features with synthetic features generated via diagonal covariance matrices.
Figure F : Feature visualizations for ImageNet-A. t-SNE (left) and UMAP (right) projections comparing real ViT features with synthetic features generated via diagonal covariance matrices.
Figure G : Feature visualizations for EuroSAT. t-SNE (left) and UMAP (right) projections comparing real ViT features with synthetic features generated via diagonal covariance matrices.
Federated Learning (FL) enables collaborative training of distributed clients while protecting privacy. To enhance generalization capability in FL, prototype-based FL is in the spotlight, since shared global prototypes offer semantic anchors for aligning client-specific local prototypes. However, existing methods update global prototypes at the prototype-level via averaging local prototypes or refining global anchors, which often leads to semantic drift across clients and subsequently yields a misaligned global signal. To alleviate this issue, we introduce hyper-prototypes, defined by a set of learnable global class-wise prototypes to preserve underlying semantic knowledge across clients. The hyper-prototypes are optimized via gradient matching to align with class-relevant characteristics distilled directly from clients' real samples, rather than prototype-level descriptors. We further propose FedHPro, a Federated Hyper-Prototype Learning framework, to leverage hyper-prototypes to promote inter-class separability via mutual-contrastive learning with client-specific margin, while encouraging intra-class uniformity through a consistency penalty. Comprehensive experiments under diverse heterogeneous scenarios confirm that 1) hyper-prototypes produce a more semantically consistent global signal, and 2) FedHPro achieves state-of-the-art performance on several benchmark datasets. Code is available at \href{https://github.com/mala-lab/FedHPro}{https://github.com/mala-lab/FedHPro}.
Huan Wang, Jun Shen, Haoran Li +6
School of Computing and Information Technology, University of Wollongong, Wollongong, Australia · School of Computing and Information Systems, Singapore Management University, Singapore, Singapore · Monash University, Australia +3
Federated continual learning (FCL) enables collaborative model training across distributed clients on sequentially arriving tasks without revisiting past data. However, existing approaches often suffer from catastrophic forgetting, rely on replay buffers or generative models that may violate privacy constraints, or assume knowledge of task identities during inference. We propose FedProTIP (Federated Projection-based Continual Learning with Task Identity Prediction), a replay-free FCL framework that maintains shared task-specific feature subspaces across clients. Each client extracts low-rank core bases from intermediate activations using randomized singular value decomposition, capturing dominant feature directions associated with the current task. These bases are transmitted to the server and aggregated to construct global task subspaces that capture shared feature directions across clients without requiring data sharing. During training, client updates are projected onto the orthogonal complement of previously learned subspaces to reduce cross-task interference and mitigate catastrophic forgetting. The learned subspaces are also reused during inference to estimate task identity via subspace relevance, enabling task-agnostic prediction without requiring explicit task labels. Experiments on CIFAR100, ImageNet-R, and DomainNet demonstrate that FedProTIP consistently outperforms state-of-the-art federated continual learning baselines while maintaining lower training time, memory footprint, and communication cost.
Seohyeon Cha, Huancheng Chen, Haris Vikalo
Department of Electrical and Computer Engineering The University of Texas at Austin
Federated Learning (FL) emerged as a promising distributed machine learning paradigm. However, extending FL to the class incremental learning scenarios introduces unique challenges: 1) Capacity conflict and catastrophic forgetting from the shared model overloading, 2) Heterogeneity from Non-Independent and Identically Distributed (Non-IID) data, and 3) Synchronized class misalignment. In this paper, we propose \textbf{F}isher-Routed \textbf{M}i\textbf{X}ture of Experts for \textbf{Fed}erated Class-Incremental Learning (\textsc{FedFMX}), a novel framework to address these challenges via adaptive expert specialization across clients. The crucial insight is to route each sample to an expert subset that jointly optimizes knowledge acquisition and retention. Specifically, we introduce a Fisher-Routed Expert Scoring (FRES) module to estimate expert importance via Fisher-based stability cost and gradient-based plasticity gain. Then, we design an Adaptive Expert Selection (AES) module by quantifying marginal contributions for adaptive expert subset determination. Finally, by the routing-aware regularization (RAR), we achieve load balance and efficient FL training. We theoretically prove the O(T−1) convergence rate. Extensive experiments on multiple benchmarks compared with state-of-the-art methods demonstrate the superiority of \textsc{FedFMX}.
Wenhao Yuan, Chenchen Lin, Jian Chen +3
Department of Electrical and Computer Engineering, The University of Hong Kong, Hong Kong SAR, China · School of Artificial Intelligence, Sun Yat-sen University, China.