Multimodal inputs are inherently heterogeneous, not only across modalities but also in the information pathways required for effective prediction. To address this limitation, we propose InstMoE, an adaptive expert routing framework for multimodal learning. InstMoE dynamically routes each input to specialized unimodal and cross-modal experts, allowing the model to adapt its information pathways to the characteristics of the input. However, routing can be misled when modality-specific variations obscure task-relevant semantics. Such irrelevant variations may distort routing decisions, causing inputs to be assigned to inappropriate experts. We therefore introduce a Contrastive Semantic Alignment module, which encourages semantically similar inputs to share task-relevant representations while suppressing irrelevant modality-specific variations. Experiments on multimodal sentiment analysis benchmarks demonstrate that InstMoE achieves state-of-the-art performance on CMU-MOSEI and CH-SIMS v2 while using substantially fewer parameters than competitive baselines. Further analysis shows that different inputs exhibit distinct expert preferences, demonstrating that InstMoE moves beyond fixed fusion toward adaptive multimodal computation.
Figures & tables
Instance
Language
Audio
Vision
Pathway
1
“I absolutely love this movie.”
Neutral tone
Neutral expression
T
2
“Yeah, that’s great.”
Sarcastic tone
Neutral expression
T+A
3
“I’m fine.”
Neutral tone
Sad expression
T+V
4
“I guess I’m okay.”
Hesitant, low-energy tone
Sad facial expression
T+A+V
Table 1: Illustrative multimodal instances showing that different instances may rely on different compositions of multimodal evidence. T, A, and V denote textual, acoustic, and visual information, respectively.
Figure 1: Overview of InstMoE.
Dataset
Lang.
Total Samples
Split (Train/Valid/Test)
Alignment
CH-SIMS v2
Zh
4,002
2,400 / 800 / 802
Unaligned
CMU-MOSEI
En
22,856
16,326 / 1,871 / 4,659
Aligned
Table 2: Statistics of the primary benchmark datasets.
Model
Venue
MAE ↓
Corr ↑
Acc-2 ↑
F1-Score ↑
Params (M) ↓
Train / ep. (s) ↓
Infer. (ms) ↓
CubeMLP* Sun et al. (2022)
ACM MM 2022
0.334
0.648
71.95
78.53
∼ 6.5
∼ 210
∼ 3.8
CENet † Wang et al. (2022)
TMM 2022
0.310
0.699
79.56
79.63
∼ 6.4
∼ 210
∼ 3.7
ALMT* Zhang et al. (2023)
EMNLP 2023
0.308
0.700
71.86
79.51
∼ 12.4
∼ 310
∼ 5.4
KuDA* Feng et al. (2024)
EMNLP 2024
0.289
0.741
76.21
82.11
∼ 10.6
∼ 280
∼ 5.0
KAN-MCP* Luo et al. (2025)
ACM MM 2025
0.281
0.742
81.60
81.70
∼ 3.95
∼ 140
∼ 2.8
CaReFlow Mai and Han (2026a)
CVPR 2026
0.277
0.745
82.90
82.90
∼ 7
∼ 210
∼ 4.0
Table 3: PA Comparative Analysis of Performance and Computational Efficiency on CH-SIMS v2. Best result per metric in bold . † denotes results from Mao et al. (2022) ; * denotes results from author-provided code. Results marked with ‡ indicate statistical significance ( p<0.05 ) over the previous SOTA (KAN-MCP) via paired t-test.
Model
Venue
MAE ↓
Corr ↑
Acc-2 ↑
F1-Score ↑
Acc-7 ↑
Params (M) ↓
Train / ep. (s) ↓
Infer. (ms) ↓
UniMSE Hu et al. (2022)
EMNLP 2022
0.523
0.773
87.46
87.50
54.39
∼ 12.1
∼ 300
∼ 5.2
HyCon Mai et al. (2023)
TAC 2023
0.590
0.792
86.50
86.40
52.80
∼ 7.6
∼ 200
∼ 3.8
DMD Li et al. (2023)
CVPR 2023
-
-
85.00
84.90
53.70
∼ 9.8
∼ 210
∼ 4.0
MMML Wu et al. (2024a)
NAACL 2024
0.517
0.790
86.49
86.50
54.95
∼ 8.4
∼ 230
∼ 4.1
EMOE Fang et al. (2025)
CVPR 2025
0.536
-
85.30
85.30
54.10
∼ 10.1
∼ 250
∼ 4.7
CaReFlow Mai and Han (2026a)
CVPR 2026
0.504
0.799
87.90
88.00
54.10
∼ 7
∼ 210
∼ 4.0
Table 4: A Comparative Analysis of Performance and Computational Efficiency on the CMU-MOSEI Dataset. Acc-2 and F1-Score are non-zero binary metrics. Best results in bold . Baseline data are primarily retrieved from Wu et al. (2024a) , except for KAN-MCP which is cited from its original paper Luo et al. (2025) . Results marked with ‡ indicate statistical significance ( p<0.05 ) over the second-best F1-Score.
MAE ↓
Corr. ↑
Acc-2 ↑
F1-Score ↑
InstMoE
0.271
0.751
85.85
84.87
w/o MoE Module
0.302 ( ↑ 0.031)
0.721 ( ↓ 0.030)
82.81 ( ↓ 3.04)
81.52 ( ↓ 3.35)
w/o CA module
0.298 ( ↑ 0.027)
0.737 ( ↓ 0.014)
83.94 ( ↓ 1.91)
82.05 ( ↓ 2.82)
w/o Centroid Init
0.284 ( ↑ 0.013)
0.743 ( ↓ 0.008)
84.35 ( ↓ 1.50)
83.13 ( ↓ 1.74)
w/o Aux Loss
0.283 ( ↑ 0.012)
0.747 ( ↓ 0.004)
84.82 ( ↓ 1.03)
83.72 ( ↓ 1.15)
Table 5: Ablation study results on CH-SIMS v2. The values in parentheses indicate the performance degradation compared to the full InstMoE model. “CA” denotes contrastive alignment module.
Figure 2: Average weight allocation per expert across sentiment intensity groups (CH-SIMS v2).
Condition
GA
GV
GT
Original
31.4
34.2
34.4
Audio degraded (80%)
14.7
41.6
43.7
Visual degraded (80%)
42.1
16.3
41.6
Text degraded (80%)
43.5
42.7
13.8
Audio + Visual degraded (80%)
11.2
13.5
75.3
Table 6: Routing behavior under controlled modality degradation. GA , GV , and GT denote the aggregated normalized routing weights assigned to audio-, visual-, and text-related experts, respectively. The reported values are preliminary placeholders.
Dataset
Representation
Speaker ID
Clip Source
Emotion Acc
CH-SIMS v2
w/o CA
72
68
78.4
w/ CA
65 ( ↓ 7)
50 ( ↓ 18)
77.9 ( ↓ 0.5)
MOSEI
w/o CA
90
68
79.6
w/ CA
69 ( ↓ 21)
54 ( ↓ 14)
78.2 ( ↓ 1.4)
Table 7: Representation probing results with and without contrastive alignment (CA).
Dataset
Model
Pitch ↑
Speed ↑
Noise ↑
Δ Acc ↓
CH-SIMS v2
w/o CA
76.4
74.1
70.2
-5.4
w CA
85.1
83.7
82.4
-2.1
MOSEI
w/o CA
74.2
71.5
69.8
-6.7
w/ CA
85.4
86.2
83.8
-1.7
Table 8: Robustness under controlled stylistic perturbations, measured by consistency between original and perturbed samples and the resulting drop in accuracy.
Figure 3: t-SNE visualization of all expert network outputs (CH-SIMS v2), colored by expert type.
Hyperparameter
CMU-MOSEI
CH-SIMS v2
Training Parameters
Initial Learning Rate
6.09×10−5
2.27×10−4
Weight Decay
0.0011
0.0095
Max Epochs
50
50
Main and Alignment Loss Weights
Alignment Loss ( λalign )
0.314
0.315
Table 9: Hyperparameter settings for the CMU-MOSEI and CH-SIMS v2 datasets.