Multimodal inputs are inherently heterogeneous, not only across modalities but also in the information pathways required for effective prediction. To address this limitation, we propose InstMoE, an adaptive expert routing framework for multimodal learning. InstMoE dynamically routes each input to specialized unimodal and cross-modal experts, allowing the model to adapt its information pathways to the characteristics of the input. However, routing can be misled when modality-specific variations obscure task-relevant semantics. Such irrelevant variations may distort routing decisions, causing inputs to be assigned to inappropriate experts. We therefore introduce a Contrastive Semantic Alignment module, which encourages semantically similar inputs to share task-relevant representations while suppressing irrelevant modality-specific variations. Experiments on multimodal sentiment analysis benchmarks demonstrate that InstMoE achieves state-of-the-art performance on CMU-MOSEI and CH-SIMS v2 while using substantially fewer parameters than competitive baselines. Further analysis shows that different inputs exhibit distinct expert preferences, demonstrating that InstMoE moves beyond fixed fusion toward adaptive multimodal computation.
Figures & tables
Instance
Language
Audio
Vision
Pathway
1
“I absolutely love this movie.”
Neutral tone
Neutral expression
T
2
“Yeah, that’s great.”
Sarcastic tone
Neutral expression
T+A
3
“I’m fine.”
Neutral tone
Sad expression
T+V
4
“I guess I’m okay.”
Hesitant, low-energy tone
Sad facial expression
T+A+V
Table 1: Illustrative multimodal instances showing that different instances may rely on different compositions of multimodal evidence. T, A, and V denote textual, acoustic, and visual information, respectively.
Figure 1: Overview of InstMoE.
Dataset
Lang.
Total Samples
Split (Train/Valid/Test)
Alignment
CH-SIMS v2
Zh
4,002
2,400 / 800 / 802
Unaligned
CMU-MOSEI
En
22,856
16,326 / 1,871 / 4,659
Aligned
Table 2: Statistics of the primary benchmark datasets.
Model
Venue
MAE ↓
Corr ↑
Acc-2 ↑
F1-Score ↑
Params (M) ↓
Train / ep. (s) ↓
Infer. (ms) ↓
CubeMLP* Sun et al. (2022)
ACM MM 2022
0.334
0.648
71.95
78.53
∼ 6.5
∼ 210
∼ 3.8
CENet † Wang et al. (2022)
TMM 2022
0.310
0.699
79.56
79.63
∼ 6.4
∼ 210
∼ 3.7
ALMT* Zhang et al. (2023)
EMNLP 2023
0.308
0.700
71.86
79.51
∼ 12.4
∼ 310
∼ 5.4
KuDA* Feng et al. (2024)
EMNLP 2024
0.289
0.741
76.21
82.11
∼ 10.6
∼ 280
∼ 5.0
KAN-MCP* Luo et al. (2025)
ACM MM 2025
0.281
0.742
81.60
81.70
∼ 3.95
∼ 140
∼ 2.8
CaReFlow Mai and Han (2026a)
CVPR 2026
0.277
0.745
82.90
82.90
∼ 7
∼ 210
∼ 4.0
Table 3: PA Comparative Analysis of Performance and Computational Efficiency on CH-SIMS v2. Best result per metric in bold . † denotes results from Mao et al. (2022) ; * denotes results from author-provided code. Results marked with ‡ indicate statistical significance ( p<0.05 ) over the previous SOTA (KAN-MCP) via paired t-test.
Model
Venue
MAE ↓
Corr ↑
Acc-2 ↑
F1-Score ↑
Acc-7 ↑
Params (M) ↓
Train / ep. (s) ↓
Infer. (ms) ↓
UniMSE Hu et al. (2022)
EMNLP 2022
0.523
0.773
87.46
87.50
54.39
∼ 12.1
∼ 300
∼ 5.2
HyCon Mai et al. (2023)
TAC 2023
0.590
0.792
86.50
86.40
52.80
∼ 7.6
∼ 200
∼ 3.8
DMD Li et al. (2023)
CVPR 2023
-
-
85.00
84.90
53.70
∼ 9.8
∼ 210
∼ 4.0
MMML Wu et al. (2024a)
NAACL 2024
0.517
0.790
86.49
86.50
54.95
∼ 8.4
∼ 230
∼ 4.1
EMOE Fang et al. (2025)
CVPR 2025
0.536
-
85.30
85.30
54.10
∼ 10.1
∼ 250
∼ 4.7
CaReFlow Mai and Han (2026a)
CVPR 2026
0.504
0.799
87.90
88.00
54.10
∼ 7
∼ 210
∼ 4.0
Table 4: A Comparative Analysis of Performance and Computational Efficiency on the CMU-MOSEI Dataset. Acc-2 and F1-Score are non-zero binary metrics. Best results in bold . Baseline data are primarily retrieved from Wu et al. (2024a) , except for KAN-MCP which is cited from its original paper Luo et al. (2025) . Results marked with ‡ indicate statistical significance ( p<0.05 ) over the second-best F1-Score.
MAE ↓
Corr. ↑
Acc-2 ↑
F1-Score ↑
InstMoE
0.271
0.751
85.85
84.87
w/o MoE Module
0.302 ( ↑ 0.031)
0.721 ( ↓ 0.030)
82.81 ( ↓ 3.04)
81.52 ( ↓ 3.35)
w/o CA module
0.298 ( ↑ 0.027)
0.737 ( ↓ 0.014)
83.94 ( ↓ 1.91)
82.05 ( ↓ 2.82)
w/o Centroid Init
0.284 ( ↑ 0.013)
0.743 ( ↓ 0.008)
84.35 ( ↓ 1.50)
83.13 ( ↓ 1.74)
w/o Aux Loss
0.283 ( ↑ 0.012)
0.747 ( ↓ 0.004)
84.82 ( ↓ 1.03)
83.72 ( ↓ 1.15)
Table 5: Ablation study results on CH-SIMS v2. The values in parentheses indicate the performance degradation compared to the full InstMoE model. “CA” denotes contrastive alignment module.
Figure 2: Average weight allocation per expert across sentiment intensity groups (CH-SIMS v2).
Condition
GA
GV
GT
Original
31.4
34.2
34.4
Audio degraded (80%)
14.7
41.6
43.7
Visual degraded (80%)
42.1
16.3
41.6
Text degraded (80%)
43.5
42.7
13.8
Audio + Visual degraded (80%)
11.2
13.5
75.3
Table 6: Routing behavior under controlled modality degradation. GA , GV , and GT denote the aggregated normalized routing weights assigned to audio-, visual-, and text-related experts, respectively. The reported values are preliminary placeholders.
Dataset
Representation
Speaker ID
Clip Source
Emotion Acc
CH-SIMS v2
w/o CA
72
68
78.4
w/ CA
65 ( ↓ 7)
50 ( ↓ 18)
77.9 ( ↓ 0.5)
MOSEI
w/o CA
90
68
79.6
w/ CA
69 ( ↓ 21)
54 ( ↓ 14)
78.2 ( ↓ 1.4)
Table 7: Representation probing results with and without contrastive alignment (CA).
Dataset
Model
Pitch ↑
Speed ↑
Noise ↑
Δ Acc ↓
CH-SIMS v2
w/o CA
76.4
74.1
70.2
-5.4
w CA
85.1
83.7
82.4
-2.1
MOSEI
w/o CA
74.2
71.5
69.8
-6.7
w/ CA
85.4
86.2
83.8
-1.7
Table 8: Robustness under controlled stylistic perturbations, measured by consistency between original and perturbed samples and the resulting drop in accuracy.
Figure 3: t-SNE visualization of all expert network outputs (CH-SIMS v2), colored by expert type.
Hyperparameter
CMU-MOSEI
CH-SIMS v2
Training Parameters
Initial Learning Rate
6.09×10−5
2.27×10−4
Weight Decay
0.0011
0.0095
Max Epochs
50
50
Main and Alignment Loss Weights
Alignment Loss ( λalign )
0.314
0.315
Table 9: Hyperparameter settings for the CMU-MOSEI and CH-SIMS v2 datasets.
Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization across modalities and domains despite not being explicitly trained for modularity. Building on this structure, we introduce ExpertLens, a data-free method that identifies domain-specialized experts directly from pretrained model weights by decoding router weights into semantically meaningful vocabulary tokens. We leverage this specialization for efficient multimodal adaptation by selectively fine-tuning experts relevant to a target domain. Across math, medical, and remote sensing tasks, ExpertLens matches or surpasses full fine-tuning while updating only 21.7 - 47.0% of model parameters and achieving a 4.0x average training speedup, and outperforms LoRA in both adaptation performance and training efficiency. These results show that sparsity introduced for efficiency can give rise to semantic modularity that is directly useful for efficient adaptation.
Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert routing remains under-explored. Existing routing strategies are either hand-crafted or modality-agnostic, relying on idealized priors that ignore the layer-dependent modality fusion patterns in MoE-VLMs and provide little guidance for expert specialization. We propose Soft Modality-guided Expert Specialization (SMoES), which consists of dynamic soft modality scores that capture layer-dependent fusion patterns, an expert binning mechanism aligned with expert-parallel deployment, and an inter-bin mutual information regularization that encourages coherent modality specialization. Our method leverages attention-based or Gaussian-statistics modality scores to optimize mutual information regularization. Experiments across four MoE-based VLMs and 16 benchmarks demonstrate improvement on both effectiveness and efficiency: 0.9% and 4.2% average gain on multimodal and language-only tasks, 56.1% reduction in EP communication overhead, and 12.3% throughput improvement under realistic deployment. These results validate that aligning routing with modality-aware expert specialization unlocks MoE-VLM capacity and efficiency.
Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across diverse modalities and tasks. Despite its growing success, a comprehensive and systematic review on the MoE metho addressing multimodal challenges remains lacking. Existing surveys tend to evaluate either multimodal learning or MoE independently from method taxonomy, overlooking the unique interplay between them. This survey fills that gap by answering a central question: \textit{How does MoE effectively resolve multimodal challenges?} We approach this from three key perspectives: (1) \textbf{MoE as an Efficient Multimodal Engine:} enabling scalable multimodal modeling by decoupling computational cost from parameter growth and mitigating modality redundancy through selective expert activation; (2) \textbf{MoE as a Multimodal Representation Learner:} integrating complementary multi-opinion expert knowledge to enrich alignment and interaction representations; and (3) \textbf{MoE as a Multimodal Adapter:} providing a modular and flexible mechanism to model imperfect data scenarios such as modality imbalance and missing modality. Through our extensive literature review, we identify critical research gaps, including interpretable routing, expert communication, modality integration, and lifelong multimodal learning. We position this survey as a foundation for future research toward interpretable and sustainable multimodal Mixture-of-Experts system.
Liangwei Nathan Zheng, Wei Emma Zhang, Olaf Maennel +2