Organizations: School of Computer Science, South China Normal University, Guangzhou, China, 510631 · School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, China, 510275
Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating information from multiple modalities, such as text, audio, and visual data. However, existing methods often suffer from spurious correlations both within and across modalities, leading models to rely on statistical shortcuts rather than true causal relationships, thereby undermining generalization. To mitigate this issue, we propose a Multi-relational Multimodal Causal Intervention (MMCI) framework, which leverages the backdoor adjustment from causal theory to address the confounding effects of such shortcuts. Specifically, we first model the multimodal inputs as a multi-relational graph to explicitly capture intra- and inter-modal dependencies. Then, we apply an attention mechanism to separately estimate and disentangle the causal features and shortcut features corresponding to these intra- and inter-modal relations. Finally, by approximating backdoor adjustment, we stratify the shortcut features and dynamically combine them with the causal features to encourage MMCI to produce stable predictions under distribution shifts. Extensive experiments on several standard MSA datasets and out-of-distribution (OOD) settings demonstrate that our method effectively suppresses biases and improves performance.
Figures & tables
Fig. 1: The distribution of several words from the training set on the CMU-MOSI dataset [ 2 ] .
Fig. 2: A testing sample from the CMU-MOSI [ 2 ] dataset has a sentiment label of -2.4. The re-implemented ITHP model [ 3 ] makes correct predictions using text inputs but fails on multimodal inputs.
Notation
Definition
Lm , dm
Sequence length and feature dimension of modality m .
y , y^
Ground-truth label and predicted sentiment value.
D , N
Training set and number of training samples.
Z
Value space of the confounding variable Z .
Xm∈RL×d
Aligned feature representation of modality m .
Hm∈RL×d
Encoded representation of modality m .
TABLE I: Summary of notations.
Fig. 3: (a) A causal graph tailored for modality fusion in MSA. (b) The same graph with backdoor adjustment to address confounding effects.
Fig. 4: The illustration of the proposed MMCI consists of three main components, shown from left to right: (1) Multi-relational Graph Construction, (2) Causal and Shortcut Attention Estimation, and (3) Disentanglement and Causal Intervention.
Dataset
CMU-MOSI
CMU-MOSEI
CH-SIMS
Train
1,284
16,326
1,368
Valid
229
1,871
456
Test
686
4,659
457
Batch Size
8
32
16
Epochs
50
15
50
Warm-up
✓
✓
✓
TABLE II: Dataset statistics, data splits, and hyper-parameter settings for the three multimodal datasets.
Method
CMU-MOSI
CMU-MOSEI
Acc2( ↑ )
F1( ↑ )
Acc7( ↑ )
MAE( ↓ )
Corr( ↑ )
Acc2( ↑ )
F1( ↑ )
Acc7( ↑ )
MAE( ↓ )
Corr( ↑ )
Graph-MFN § [ 58 ]
76.8 / 78.0
76.7 / 78.0
34.4
0.966
0.653
80.3 / 83.8
80.8 / 83.7
52.0
0.564
0.727
MTAG § [ 37 ]
80.8 / 83.1
80.4 / 82.8
25.9
0.879
0.706
81.5 / 84.6
81.7 / 84.5
52.6
0.559
0.735
HGraph-CL † [ 36 ]
84.3 / 86.2
84.6 / 86.2
-
0.717
0.799
84.5 / 85.9
84.5 / 85.8
-
0.527
0.769
ALMT [ 25 ]
82.4 / 84.5
82.2 / 84.4
45.9
0.741
0.776
80.7 / 84.6
81.3 / 84.7
52.7
0.543
0.761
TETFN † [ 23 ]
84.1 / 86.1
83.8 / 86.1
-
0.717
0.800
84.3 / 85.2
84.2 / 85.3
-
0.551
0.748
TABLE III: Comparison on the CMU-MOSI and CMU-MOSEI datasets. † denotes results reported in the original papers, § denotes results from [ 57 ] , while the remaining results are obtained from our own experiments. The best results are highlighted in bold, and the second-best results are underlined.
Method
CH-SIMS
Acc-5( ↑ )
Acc-3( ↑ )
Acc-2( ↑ )
F1( ↑ )
MAE( ↓ )
Corr( ↑ )
TFN [ 15 ]
39.3
65.1
78.4
78.6
0.432
0.591
LMF [ 66 ]
40.5
64.7
77.8
77.9
0.441
0.576
MFN [ 67 ]
39.5
65.7
77.9
77.9
0.435
0.582
MulT [ 5 ]
37.9
64.8
78.6
79.7
0.453
0.564
Self-MM [ 68 ]
41.5
65.5
80.0
80.4
0.425
0.595
TABLE IV: Comparison on the CH-SIMS dataset. † indicates that the results are taken from [ 60 ] , while the remaining results are reported in [ 65 ] .
Method
CMU-MOSI (OOD)
Acc7( ↑ )
Acc2*( ↑ )
Acc2( ↑ )
F1*( ↑ )
F1( ↑ )
MulT [ 5 ]
29.8
75.0
76.7
74.8
76.5
MAG-BERT [ 6 ]
39.9
75.6
77.3
75.5
77.3
MISA [ 30 ]
38.0
75.9
77.4
75.8
77.4
Self-MM [ 68 ]
40.2
76.7
78.1
76.7
78.1
CLUE [ 13 ]
41.8
78.8
79.9
78.8
79.9
TABLE V: Comparison on the OOD version of the CMU-MOSI dataset. Results marked with † are from the original papers, and § indicates results from our experiments. Other results are taken from [ 13 ] .
Fig. 5: Cross-dataset evaluation of methods trained on the CMU-MOSI and CMU-MOSEI training sets and tested on the corresponding opposite datasets, including both MOSI → MOSEI and MOSEI → MOSI transfer settings. The results of ITHP and C-MIB are obtained from our own implementations, while the remaining results are adopted from [ 51 ] .
Fig. 6: Parameter sensitivity analysis of disentanglement weight λ and causal intervention weight β . The gray dashed line denotes the best baseline.
Fig. 7: Effect of graph attention layers on CMU-MOSI (left) and CMU-MOSEI (right).
Method
Acc2( ↑ )
F1( ↑ )
Acc7( ↑ )
w/o Intra-modal Relation
83.6 / 85.5
83.6 / 85.5
45.7
w/o Inter-modal Relation
87.2 / 88.7
87.1 / 88.7
46.6
w/o Causal Attention
86.3 / 87.8
86.3 / 87.8
45.8
w/o Shortcut Attention
86.9 / 88.4
86.8 / 88.4
48.0
w/o Dual Attention Assignment
84.9 /86.1
85.0 / 86.2
45.4
w/o Disentanglement
86.6 / 88.4
86.5 / 88.4
47.3
TABLE VI: Ablation experiments on the CMU-MOSI dataset.
Method
Gaussian Noise
Salt-Pepper Noise
ϵ=1
ϵ=5
ϵ=10
ϵ=1
ϵ=5
ϵ=10
Noisy CMU-MOSI
C-MIB [ 8 ]
84.6
78.5
68.4
84.6
78.6
67.3
ITHP [ 3 ]
86.1
79.2
64.9
82.8
76.3
63.1
QMF [ 69 ]
84.1
76.3
69.5
83.1
80.0
63.7
MMCI (Ours)
87.0
81.2
68.2
86.3
73.7
67.5
TABLE VII: Acc2 performance under Gaussian and Salt-Pepper noise on the CMU-MOSI and CMU-MOSEI datasets.
Method
Acc2( ↑ )
F1( ↑ )
MAE( ↓ )
Corr( ↑ )
Self-MM b [ 68 ]
84.0
84.4
0.713
0.798
MMIM b [ 70 ]
84.1
84.0
0.700
0.800
MAG b [ 6 ]
84.2
84.1
0.712
0.796
C-MIB † b [ 8 ]
84.7
84.7
0.717
0.795
Self-MM d [ 68 ]
55.1
53.5
1.440
0.158
MMIM d [ 70 ]
85.8
85.9
0.649
0.829
TABLE VIII: Performance comparison on the CMU-MOSI dataset. Methods based on BERT and DeBERTa are marked with subscripts “b” and “d”, respectively. † indicates results obtained from our experiments, while other results are taken from [ 3 ] .
Modality
ITHP [ 3 ]
MMCI (Ours)
T
V
A
Acc7 ( ↑ )
Acc2 ( ↑ )
Acc7 ( ↑ )
Acc2 ( ↑ )
✓
x
x
42.3
85.3 / 87.0
46.0
86.0 / 87.9
✓
✓
x
43.5
85.4 / 87.5
48.0
86.9 / 88.5
✓
x
✓
46.7
84.8 / 86.7
47.4
86.3 / 88.2
✓
✓
✓
46.3
86.1 / 88.2
47.6
87.4 / 89.3
TABLE IX: Performance Comparison of MMCI and ITHP across various modality combinations on the CMU-MOSI dataset.
Method
Parameters
Training Time
Inference Time
EMOE [ 62 ]
317,295,771
15.99 s
1.11 s
GLoMo [ 32 ]
109,818,887
13.37 s
0.60 s
C-MIB [ 8 ]
190,930,180
16.09 s
0.65 s
ITHP [ 3 ]
184,883,706
21.16 s
0.89 s
MMCI (Ours)
186,461,076
20.56 s
0.71 s
TABLE X: Comparison of the number of parameters, training time, and inference time between MMCI and baselines on the CMU-MOSI dataset. All times are reported per epoch with batch size 8.
Fig. 8: Visualization of causal features (left) and non-causal features (right). HN: Highly Negative; N: Negative; WN: Weakly Negative; NT: Neutral; WP: Weak Positive; P: Positive; HP: Highly Positive.
Fig. 9: Comparison of FLOPs and peak GPU memory consumption. Peak memory is measured on CMU-MOSI with batch size 32.
Fig. 10: Qualitative case studies of MMCI on the CMU-MOSI dataset. For Case 1, (a)–(e) respectively show the predictions of ITHP, text-only ITHP, MMCI, MMCI without shortcut supervision, and MMCI without causal intervention. Cases 2 and 3 illustrate spurious correlations associated with visual attributes, while Cases 4 and 5 illustrate dataset-specific lexical associations. For Cases 4 and 5, the attention assigned by ITHP and MMCI to the selected biased word is normalized across the two models so that the reported ratios sum to one, representing relative inter-model attention rather than raw values.
Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbalance. Existing methods generate features for modality missing from available ones, but differences in expression mechanisms and sentiment dynamics across modalities may cause the generated features to deviate from true distributions and mislead prediction. In addition, unreliable modalities may dominate fusion, resulting in representation shift across modality combinations and unstable sentiment representations. To address these challenges, we propose a two-level reference alignment framework. The framework introduces stable references at the feature representation and sentiment decision levels to improve robustness under modality missing. First-level reference alignment leverages complete-modality samples to constrain representations and align different modality combinations into a shared sentiment space. Second-level reference alignment enforces cross-modal consistency at the decision level by suppressing unreliable modalities through prototype retrieval and voting. As a result, the framework maintains stable and reliable sentiment predictions under diverse missing-modality patterns. Experiments on CMU-MOSI and CMU-MOSEI show consistent improvements across various missing-modality settings. Under full-modality input, the proposed method achieves state-of-the-art performance, with ACC of 86.28% and 85.88%, and F1 of 86.24% and 85.86%.
Chenglizhao Chen, Yuchen Cao, Xinyu Liu +3
Qingdao Institute of Software, College of Computer Science and Technology, China University of Petroleum (East China) · Shandong Key Laboratory of Intelligent Oil & Gas Industrial Software · The Hong Kong University of Science and Technology (Guangzhou)
Multimodal Sentiment Analysis (MSA) fuses text, acoustic, and visual streams to infer sentiment. Because pre-trained text encoders are far more expressive than their acoustic and visual counterparts, the text modality tends to dominate optimization, suppressing weaker modalities and inducing gradient norm conflicts that destabilize training. To address this, we propose a Conflict-aware Penalty (CP) that detects and penalizes gradient norm conflicts at each training step, and a Statistical Loss (SL) that aligns predicted distribution statistics with empirical input statistics. Crucially, CP prevents dominant modality gradients from interfering with the SL objective, enabling synergistic training within a unified framework incorporating adaptive modality encoding, gated cross-modal fusion, and unimodal auxiliary heads. Experiments on CMU-MOSI demonstrate state-of-the-art performance, with ablation studies confirming the effectiveness of each component.
Jianheng Dai, Jiazhang Liang, Sijie Mai
School of Computer Science, South China Normal University, Guangzhou, Guangdong, China
Recent research on Multimodal Sentiment Analysis (MSA) has focused on learning from language, visual, and acoustic modalities with incomplete data to infer human sentiment. Most studies typically compensate for missing information by reconstructing modality features or designing complicated fusion mechanisms. However, these methods still suffer from spurious generation and noisy guidance due to the lack of high-level semantic grounding in partially observed multimodal evidence. To address these issues, we propose SemMSA, a latent semantic-aided framework that constructs rich sentiment-relevant semantics with LLMs, fully integrating with all modalities via anchor-free spectral alignment. It mainly consists of Cross-modal Semantic Refinement (CSR) and Cross-modal Spectral Alignment (CSA). Specifically, CSR first adaptively extracts visual and acoustic representations by corresponding adapters to form a unified multimodal prefix with language in the frozen LLM embedding space. It then iteratively produces continuous discriminative semantic states through a token-efficient latent refinement process without decoding explicit text. Next, CSA simultaneously aligns the refined semantics with all modalities by enhancing the dominant spectral component of their kernel Gram matrix. This captures global nonlinear dependencies among all representations without relying on a predefined anchor modality. In addition, an instance-level spectral separation constraint preserves cross-sample discriminability and mitigates representation collapse. Extensive experiments on SIMS, MOSI, and MOSEI benchmarks demonstrate that SemMSA achieves state-of-the-art performance.
Wenhao Li, Zhibin Wu, Chong Xiao +1
Software School, Shandong University · Shenzhen Loop Area Institute