Multimodal sentiment analysis integrates textual, acoustic and visual cues, yet current language-model-based fusion methods typically leave prediction-layer branch allocation implicit. We introduce Branch-Calibrated Multimodal Language Fusion (BC-MLF), which explicitly models prediction-layer branch allocation through a Branch-Calibrated Task Head (BCHead), complemented by Fusion Token Contrastive Learning (FTCL) for sentiment-aware fusion-token regularization. FTCL organizes mean-pooled fusion-token representations according to continuous sentiment affinity, while BCHead combines fusion, text and audiovisual predictions through a lightweight sample-adaptive constrained mixture. Without modifying the fusion backbone, BC-MLF consistently improves the reproduced DeepMLF baseline and achieves the strongest results among the compared methods on CMU-MOSEI and CH-SIMS across classification and regression metrics. The controlled ablations show that sample-adaptive prediction-layer branch allocation consistently outperforms static branch aggregation. Code is available at https://github.com/sunyulin0421/BC-MLF.
Figures & tables
Figure 1: Branch-Calibrated Multimodal Language Fusion (BC-MLF). The unchanged fusion backbone produces fusion-token, text and audiovisual representations. During training, FTCL regularizes fusion-token geometry according to continuous sentiment affinity, whereas BCHead calibrates prediction-layer branch allocation through a lightweight constrained mixture.
MODEL
CMU-MOSEI
CH-SIMS
Acc2 ↑
F1 ↑
MAE ↓
Corr ↑
Acc5 ↑
Acc7 ↑
Acc2 ↑
F1 ↑
MAE ↓
Corr ↑
LF-DNN †
82.78
82.38
0.558
0.731
–
–
76.68
76.48
0.446
0.567
TFN †
82.23
81.47
0.573
0.718
–
–
77.07
76.94
0.437
0.582
MAG-BERT †
84.87
84.85
0.539
0.764
–
–
74.44
71.75
0.492
0.399
MulT †
84.07
83.93
0.564
0.731
53.97
52.56
78.56
78.66
0.453
0.564
MISA †
84.51
84.47
0.549
0.759
53.57
51.96
76.54
76.59
0.447
0.563
Table 1: Comparison with representative state-of-the-art MSA methods on CMU-MOSEI and CH-SIMS. † : results reported in [ 2 ] ; ∗ : results reproduced; ↑/↓ : higher/lower is better. Bold: best result in each column. Results of BC-MLF are averaged over two random seeds.
Method
Acc2 ↑
F1 ↑
MAE ↓
Corr ↑
MOSEI
DeepMLF
86.08
86.10
0.505
0.802
+ FTCL
86.75
86.74
0.501
0.805
+ BCHead
87.40
87.39
0.495
0.807
+ FTCL + Uniform
86.18
86.13
0.518
0.798
+ FTCL + Global
86.15
86.14
0.514
0.799
Table 2: Ablation studies of BC-MLF. The results demonstrate the effectiveness of FTCL and BCHead, and show that sample-adaptive prediction-layer allocation provides additional benefits beyond static branch aggregation.
Figure 2: Pairwise similarity shifts in mean-pooled fusion-token representations between DeepMLF and BC-MLF, shown alongside the continuous label-affinity reference used by FTCL on CH-SIMS.
Dataset
Model
Spearman ↑
DKL(Q∥P)↓
MOSEI
DeepMLF
0.3701 ± 0.0022
0.2410 ± 0.0025
DeepMLF + FTCL
0.3862 ± 0.0043
0.2358 ± 0.0021
CH-SIMS
DeepMLF
0.2610 ± 0.0312
0.1793 ± 0.0035
DeepMLF + FTCL
0.2749 ± 0.0387
0.1757 ± 0.0026
Table 3: Fusion-token alignment with continuous sentiment affinity (mean ± standard deviation over two seeds).
Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled within the same operation. We propose \textbf{SeRIn} (\textbf{Se}gregate, \textbf{R}efine, \textbf{In}tegrate), a multimodal LM fusion scheme that enforces this separation as an architectural prior. Modality-specific representations evolve along isolated pathways, each refined against its respective encoder context, while a dedicated cross-modal pathway accumulates their joint evolution without contaminating unimodal streams. Full cross-modal interaction is deferred to a final prediction step - ablations confirm that structured interactions, not added capacity, drive the gains; gate analysis under visual corruption reveals emergent modality reweighting without explicit supervision. SeRIn achieves state-of-the-art results on CH-SIMS and CMU-MOSEI, improving all metrics on both benchmarks.
Alexios Filippakopoulos, Elias Kallioras, Nikolaos Xiros +2
National Technical University of Athens, Greece · Athena Research Center, Greece · University of Bern, Switzerland +2
Multimodal sentiment analysis relies on language, visual, and acoustic cues, but utterance-level modality quality may vary due to occlusion, background noise, motion blur, or imperfect transcripts, causing conventional fusion to over-trust unreliable modalities. We propose MRUF, a reliability-aware fusion method that combines multi-granularity routing with uncertainty-aware calibration. MRUF summarizes sentiment-relevant representations, performs subspace- and modality-level routing, and supervises modality routing with leave-one-out error increases to estimate utterance-level modality importance. It further predicts modality-wise uncertainty and refines modality gates through inverse-variance reweighting, while modality-invariant contrastive alignment stabilizes the shared representation space. Experiments on CMU-MOSI and CMU-MOSEI under aligned and unaligned settings show consistent improvements over strong baselines, and mechanism analysis verifies that modalities with higher predicted uncertainty receive lower fusion weights.
Haoran Ma, Yinfeng Yu, Liejun Wang
School of Computer Science and Technology, Xinjiang University, Urumqi 830017, China. · Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing. · Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center. +2
Multimodal Sentiment Analysis (MSA) fuses text, acoustic, and visual streams to infer sentiment. Because pre-trained text encoders are far more expressive than their acoustic and visual counterparts, the text modality tends to dominate optimization, suppressing weaker modalities and inducing gradient norm conflicts that destabilize training. To address this, we propose a Conflict-aware Penalty (CP) that detects and penalizes gradient norm conflicts at each training step, and a Statistical Loss (SL) that aligns predicted distribution statistics with empirical input statistics. Crucially, CP prevents dominant modality gradients from interfering with the SL objective, enabling synergistic training within a unified framework incorporating adaptive modality encoding, gated cross-modal fusion, and unimodal auxiliary heads. Experiments on CMU-MOSI demonstrate state-of-the-art performance, with ablation studies confirming the effectiveness of each component.
Jianheng Dai, Jiazhang Liang, Sijie Mai
School of Computer Science, South China Normal University, Guangzhou, Guangdong, China