Organizations: Department of Computer Science, Hong Kong Baptist University, Hong Kong SAR, China · School of Computer Science and Engineering, Sun Yat-Sen University, Guangzhou, China
Multimodal federated learning (MFL) has emerged as a pivotal paradigm for leveraging distributed data to enhance model performance. However, existing methods predominantly rely on idealized assumptions of model homogeneity and balanced modality distributions, rendering them ill-suited for practical scenarios characterized by heterogeneous client architectures and severe modality imbalance. To address these challenges, we propose a \textbf{M}ultimodal \textbf{Fed}erated learning Prototype-guided Bilateral Alignment (MFedPBA) framework. MFedPBA facilitates robust knowledge synergy through a dual alignment mechanism: (i) at the feature level, it aligns heterogeneous feature spaces via a projection encoder optimized by contrastive learning and the Gromov-Wasserstein distance; (ii) at the decision level, it employs an entropy-weighted aggregation of naturally aligned logit prototypes. This novel design achieves robust MFL by jointly tackling heterogeneous feature spaces and collectively aggregating decisions. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines under conditions of model heterogeneity and modality imbalance.
Figures & tables
Figure 1 : (a) Illustration of the multimodal federated learning setting with complex architectural and data inconsistencies. (b) Visualization of feature space incompatibility across heterogeneous encoders, where direct prototype aggregation induces drift towards ambiguous regions, resulting in performance inferior to local training.
Figure 2 : The framework of the MFedPBA.
Datasets
Types
# S
# C
# M
Input modalities
Caltech101
{I}
9144
102
3
{I1}, {I2}, and {I3}
Reuters
{L}
18758
6
5
{L1}, {L2}, {L3} {L4}, and {L5}
NUS-WIDE
{I, L}
5000
10
4
{I1}, {I2}, {I3}, and {L}
Youtube
{I, A, T}
2000
10
6
{I1}, {I2}, {T} {A1}, {A2}, and {A3}
Table 1 : Dataset statistics. Acronyms for modality types: I (Image), L (Language), A (Audio), T (Time-series), X1 (Style-one X), X2 (Style-two X). The # S: Samples, # C: Classes, # M: Modalities.
Figure 3 : Example of dataset distribution partitioning.
Dataset
Local
FedProto
FedTGP
FedPall
Harmony
FedMVP
FedMobile
MFedPBA
Caltech101
M2
42.65 ±1.47
38.72 ±0.42
40.78 ±1.31
42.29 ±0.27
45.29 ±0.38
40.23 ±0.65
43.51 ±0.54
49.04 ±0.21
M1+
40.83 ±1.12
36.81 ±0.35
40.71 ±1.23
41.8 ±0.77
43.04 ±0.64
38.91 ±1.25
42.29 ±0.6
46.82 ±0.54
M1
38.47 ±0.45
31.57 ±1.23
36.18 ±0.72
40.77 ±0.36
40.1 ±0.69
38.92 ±0.62
40.09 ±0.13
43.64 ±0.11
K50
29.38 ±1.63
25.17 ±1.78
28.77 ±1.43
30.59 ±1.26
30.58 ±0.65
27.74 ±1.5
30.88 ±0.81
31.19 ±0.92
Reuters
M2
70.61 ±0.28
73.39 ±0.57
69.73 ±0.88
73.08 ±1.04
72.18 ±0.41
74.07 ±1.08
73.19 ±0.26
75.16 ±0.45
M1+
70.85 ±0.59
68.54 ±1.46
69.8 ±0.59
71.09 ±1.56
72.28 ±0.47
68.09 ±0.82
72.79 ±0.44
75.01 ±0.22
Table 2 : Comparison of average performance of different methods in simulations with different data distributions. M#: Data partition (K=M); K#: Number of large clients (in M1+). Best and second-best are bold and underlined respectively.
Figure 4 : Test accuracy (%) on four datasets under the M2 data distribution setting with model heterogeneity.
Figure 5 : Feature visualization results of Local and MFedPBA.
Figure 8
dD
Local
Proto
TGP
Pall
Harm
MVP
Mobile
PBA
24-d
41.19
30.12
41.19
42.56
42.06
43.85
40.91
44.43
48-d
42.34
39.05
43.85
44.63
44.43
45.21
44.57
47.02
64-d
43.02
37.82
45.22
42.27
43.13
43.21
44.24
46.46
128-d
42.64
31.24
45.01
42.15
43.24
44.33
44.33
45.68
256-d
41.95
26.32
43.22
41.75
42.74
42.84
42.54
43.84
Table 3 : The test accuracy (%) on Youtube in the M1+ setting. “Fed” is omitted in the method name due to limited space.
Figure 8 : Hyperparameter experiments for λ1 , λ2 , and S .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9 : The distribution of these datasets is set in the M2 scenario. A larger circle means a larger sample size.
Figure 10 : The distribution of these datasets is set in the M1+ scenario. A larger circle means a larger sample size.
Figure 11 : The distribution of these datasets is set in the M1 scenario. A larger circle means a larger sample size.
Figure 12 : In scenarios involving a larger number of clients, the data distribution across these four datasets within the M1+ scenario.
Table 4 : Heterogeneous model architectures for modality-specific feature extraction.
Descriptions
Caltech101
Reuters
NUS-WIDE
Youtube
Total rounds T
400 / 500
200
400
400
Server
Server optimizers
SGD
Server learning rate ηs
0.005 ∼ 0.01
Server training epoch S
10
Client
Local training epoch E
2
Training batch size B
12
Appendix
Table 5 : List of Hyperparameters.
Dataset
Local
FedProto
FedTGP
FedPall
Harmony
FedMVP
FedMobile
MFedPBA
Caltech101
k1
40.48 ±0.86
37.96 ±0.75
39.71 ±1.14
39.81 ±0.96
42.14 ±0.17
35.01 ±0.96
41.89 ±1.07
47.35 ±0.19
k2
43.28 ±0.42
38.56 ±0.21
39.86 ±0.94
44.18 ±0.16
46.44 ±0.15
41.28 ±0.42
44.72 ±0.28
48.19 ±0.18
k3
44.24 ±0.46
39.89 ±0.15
43.52 ±1.65
43.3 ±0.15
47.23 ±0.23
44.8 ±0.49
43.52 ±0.14
52.44 ±0.27
Avg
42.65 ±1.47
38.72 ±0.42
40.78 ±1.31
42.29 ±0.27
45.29 ±0.38
40.23 ±0.65
43.51 ±0.54
49.04 ±0.21
Reuters
k1
77.22 ±1.58
82.06 ±0.32
78.63 ±0.78
76.92 ±0.62
76.94 ±0.91
82.53 ±1.67
80.8 ±1.07
82.55 ±0.79
k2
66.82 ±1.42
64.49 ±0.65
62.23 ±1.46
66.88 ±0.61
68.12 ±1.05
65.62 ±1.11
68.11 ±0.38
67.83 ±0.78
Appendix
Table 6 : Performance comparison (%) of all compared methods on Caltech101, Reuters, NUS-WIDE, and Youtube using M2 data partitioning, where the number of clients K is equal to the number of modalities M . The k1 represents the client with ID 1.
Figure 13 : The test accuracy and convergence process of each method in M2 scenario.
Dataset
Local
FedProto
FedTGP
FedPall
Harmony
FedMVP
FedMobile
MFedPBA
Caltech101
k1
34.51 ±0.67
30.74 ±1.09
34.7 ±1.53
35.17 ±0.79
39.34 ±0.65
29.58 ±1.00
36.58 ±0.61
41.18 ±0.24
k2
40.63 ±0.87
35.02 ±0.39
38.57 ±0.77
41.68 ±0.83
41.04 ±1.03
33.69 ±0.54
41.05 ±0.58
45.46 ±0.47
k3
43.49 ±0.58
40.46 ±0.24
44.59 ±2.03
44.53 ±1.22
45.92 ±0.56
46.3 ±1.1
46.45 ±1.57
50.02 ±0.59
Avg
40.83 ±1.12
36.81 ±0.35
40.71 ±1.23
41.8 ±0.77
43.04 ±0.64
38.91 ±1.25
42.29 ±0.6
46.82 ±0.54
Reuters
k1
72.28 ±1.92
75.00 ±2.54
74.61 ±2.58
74.8 ±1.76
74.71 ±0.85
61.12 ±0.88
75.12 ±1.25
80.82 ±1.05
k2
82.11 ±0.77
81.2 ±0.81
80.04 ±1.56
81.9 ±1.17
80.76 ±0.29
81.63 ±0.23
83.32 ±0.18
83.72 ±0.64
Appendix
Table 7 : Performance comparison (%) of all compared methods on Caltech101, Reuters, NUS-WIDE, and Youtube using M1+ data partitioning, where the number of clients K is equal to the number of modalities M . The k1 represents the client with ID 1.
Figure 14 : The test accuracy and convergence process of each method in M1+ scenario.
Dataset
Local
FedProto
FedTGP
FedPall
Harmony
FedMVP
FedMobile
MFedPBA
Caltech101
I1
39.47 ±0.13
38.27 ±1.4
41.16 ±0.67
42 ±0.71
39.84 ±0.62
38.62 ±0.5
41.87 ±0.13
48.73 ±0.9
I2
36.89 ±0.28
25.56 ±2.04
31.47 ±1.73
37.02 ±0.56
39.78 ±1.3
38.09 ±1.16
38.17 ±0.73
40.78 ±0.37
I3
39.07 ±1.22
30.89 ±1.34
35.91 ±0.66
43.29 ±0.2
40.69 ±0.83
40.04 ±1.65
40.22 ±0.28
41.4 ±1.05
Avg
38.47 ±0.45
31.57 ±1.23
36.18 ±0.72
40.77 ±0.36
40.1 ±0.69
38.92 ±0.62
40.09 ±0.13
43.64 ±0.11
Reuters
L1
64.47 ±1.07
58.87 ±0.61
60.39 ±0.93
65.85 ±0.55
64.39 ±0.55
63.85 ±1.59
64.93 ±0.43
67.2 ±1.33
L2
69.81 ±1.4
79.52 ±0.67
77.58 ±0.88
74.31 ±2.03
76.44 ±0.18
76.41 ±0.25
75.37 ±0.38
77.87 ±0.85
Appendix
Table 8 : Performance comparison (%) of all compared methods on Caltech101, Reuters, NUS-WIDE, and Youtube using M1 data partitioning, where the number of clients K is equal to the number of modalities M . The X1 represents a client that has the X1 modal enabled.
Figure 15 : The test accuracy and convergence process of each method in M1 scenario.
Figure 16 : The testing accuracy and convergence process of each method in the M1+ scenario, with a large number of clients participating.
Multimodal Federated Learning (MMFL) enables privacy-preserving collaborative learning across decentralized clients with heterogeneous data and modality availability. However, most existing MMFL methods cast multimodal training as a joint optimization problem, overlooking a key bottleneck: modality competition, where dominant modalities suppress weaker ones and lead to suboptimal global models. To address this, we propose FedMChain, a balanced MMFL framework that structures federated multimodal training as a chain of modality-wise phases. This phase-wise design gives each modality a dedicated local optimization window on multimodal clients to mitigate modality competition, and further promotes cross-modal complementarity via an error-compensated regularizer. On the server side, we employ a sparse sign-guided aggregation strategy that leverages directional sign agreement for robust intra-modality aggregation, avoids destructive averaging, and supports less frequent synchronization to reduce communication overhead. Extensive experiments on multimodal benchmarks demonstrate that FedMChain consistently improves predictive performance while requiring less frequent communication than baselines.
Zixin Zhang, Fan Qi, Shuai Li +2
College of Computer Science, Inner Mongolia University, Hohhot, Inner Mongolia, China · School of Computer Science and Engineering, Tianjin University of Technology, Tianjin, China · Institute of Automation, Chinese Academy of Sciences, Beijing, China
Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as GPT-4v and LLaVA, which demonstrate their exceptional proficiency in multimodal tasks, such as image captioning and multimodal question answering. We introduce a novel federated learning framework, named Multimodal Large Language Model Assisted Federated Learning (MLLM-LLaVA-FL), which employs powerful MLLMs at the server end to address the heterogeneous and long-tailed challenges. Owing to the advanced cross-modality representation capabilities and the extensive open-vocabulary prior knowledge of MLLMs, our framework is adept at harnessing the extensive, yet previously underexploited, open-source data accessible from websites and powerful server-side computational resources. Hence, the MLLM-LLaVA-FL not only enhances the performance but also avoids increasing the risk of privacy leakage and the computational burden on local devices, distinguishing it from prior methodologies. Our framework has three key stages. Initially, we conduct global visual-text pretraining of the model. This pretraining is facilitated by utilizing the extensive open-source data available online, with the assistance of MLLMs. Subsequently, the pretrained model is distributed among various clients for local training. Finally, once the locally trained models are transmitted back to the server, a global alignment is carried out under the supervision of MLLMs to further enhance the performance. Experimental evaluations on established benchmarks, show that our framework delivers promising performance in the typical scenarios with data heterogeneity and long-tail distribution across different clients in FL.
Jianyi Zhang, Hao Frank Yang, Ang Li +5
Duke University · Johns Hopkins University · University of Maryland College Park +1
Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic deployments often exhibit dual-axis modality missingness: clients have different modality sets, and individual samples may contain only subsets of the modalities available locally. Existing methods typically address these two axes separately. We propose Flux, a multimodal federated learning framework built around two complementary components. First, modality-aware confidence tempering learns sample-specific confidence for each modality through mask-aware unimodal supervision and fuses the confidence estimates from observed modalities into a sample-adaptive temperature that adjusts predictive sharpness according to evidence quality and completeness. Second, gradient-decoupled private adaptation applies this temperature only to a client-private prediction pathway, while training the shared federated model with a standard, untempered objective. This enables sample-specific, client-local confidence adaptation without allowing confidence-dependent gradients to perturb shared representation learning. Across four multimodal datasets, Flux achieves the highest average macro-F1 on every dataset, outperforming the strongest dataset-specific baseline by 0.8~2.2 points and by 1.6 points on average. Additional analyses demonstrate favorable calibration, temperature sensitivity to both modality missingness and input corruption, and more stable shared optimization under private-only tempering. Our code is available at https://github.com/AdibaOrz/Flux.