Prediction-Layer Branch Calibration for Multimodal Sentiment Analysis
Organizations: National University of Defense Technology
Abstract
Multimodal sentiment analysis integrates textual, acoustic and visual cues, yet current language-model-based fusion methods typically leave prediction-layer branch allocation implicit. We introduce Branch-Calibrated Multimodal Language Fusion (BC-MLF), which explicitly models prediction-layer branch allocation through a Branch-Calibrated Task Head (BCHead), complemented by Fusion Token Contrastive Learning (FTCL) for sentiment-aware fusion-token regularization. FTCL organizes mean-pooled fusion-token representations according to continuous sentiment affinity, while BCHead combines fusion, text and audiovisual predictions through a lightweight sample-adaptive constrained mixture. Without modifying the fusion backbone, BC-MLF consistently improves the reproduced DeepMLF baseline and achieves the strongest results among the compared methods on CMU-MOSEI and CH-SIMS across classification and regression metrics. The controlled ablations show that sample-adaptive prediction-layer branch allocation consistently outperforms static branch aggregation. Code is available at https://github.com/sunyulin0421/BC-MLF.
Figures & tables
| MODEL | CMU-MOSEI | CH-SIMS | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Acc2 | F1 | MAE | Corr | Acc5 | Acc7 | Acc2 | F1 | MAE | Corr | ||
| LF-DNN | 82.78 | 82.38 | 0.558 | 0.731 | – | – | 76.68 | 76.48 | 0.446 | 0.567 | |
| TFN | 82.23 | 81.47 | 0.573 | 0.718 | – | – | 77.07 | 76.94 | 0.437 | 0.582 | |
| MAG-BERT | 84.87 | 84.85 | 0.539 | 0.764 | – | – | 74.44 | 71.75 | 0.492 | 0.399 | |
| MulT | 84.07 | 83.93 | 0.564 | 0.731 | 53.97 | 52.56 | 78.56 | 78.66 | 0.453 | 0.564 | |
| MISA | 84.51 | 84.47 | 0.549 | 0.759 | 53.57 | 51.96 | 76.54 | 76.59 | 0.447 | 0.563 | |
| Method | Acc2 | F1 | MAE | Corr |
|---|---|---|---|---|
| MOSEI | ||||
| DeepMLF | 86.08 | 86.10 | 0.505 | 0.802 |
| + FTCL | 86.75 | 86.74 | 0.501 | 0.805 |
| + BCHead | 87.40 | 87.39 | 0.495 | 0.807 |
| + FTCL + Uniform | 86.18 | 86.13 | 0.518 | 0.798 |
| + FTCL + Global | 86.15 | 86.14 | 0.514 | 0.799 |
| Dataset | Model | Spearman | |
|---|---|---|---|
| MOSEI | DeepMLF | 0.3701 0.0022 | 0.2410 0.0025 |
| DeepMLF + FTCL | 0.3862 0.0043 | 0.2358 0.0021 | |
| CH-SIMS | DeepMLF | 0.2610 0.0312 | 0.1793 0.0035 |
| DeepMLF + FTCL | 0.2749 0.0387 | 0.1757 0.0026 |