Organizations: School of Computer Science, South China Normal University, Guangzhou, China, 510631 · School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, China, 510275
Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating information from multiple modalities, such as text, audio, and visual data. However, existing methods often suffer from spurious correlations both within and across modalities, leading models to rely on statistical shortcuts rather than true causal relationships, thereby undermining generalization. To mitigate this issue, we propose a Multi-relational Multimodal Causal Intervention (MMCI) framework, which leverages the backdoor adjustment from causal theory to address the confounding effects of such shortcuts. Specifically, we first model the multimodal inputs as a multi-relational graph to explicitly capture intra- and inter-modal dependencies. Then, we apply an attention mechanism to separately estimate and disentangle the causal features and shortcut features corresponding to these intra- and inter-modal relations. Finally, by approximating backdoor adjustment, we stratify the shortcut features and dynamically combine them with the causal features to encourage MMCI to produce stable predictions under distribution shifts. Extensive experiments on several standard MSA datasets and out-of-distribution (OOD) settings demonstrate that our method effectively suppresses biases and improves performance.
Figures & tables
Fig. 1: The distribution of several words from the training set on the CMU-MOSI dataset [ 2 ] .
Fig. 2: A testing sample from the CMU-MOSI [ 2 ] dataset has a sentiment label of -2.4. The re-implemented ITHP model [ 3 ] makes correct predictions using text inputs but fails on multimodal inputs.
Notation
Definition
Lm , dm
Sequence length and feature dimension of modality m .
y , y^
Ground-truth label and predicted sentiment value.
D , N
Training set and number of training samples.
Z
Value space of the confounding variable Z .
Xm∈RL×d
Aligned feature representation of modality m .
Hm∈RL×d
Encoded representation of modality m .
TABLE I: Summary of notations.
Fig. 3: (a) A causal graph tailored for modality fusion in MSA. (b) The same graph with backdoor adjustment to address confounding effects.
Fig. 4: The illustration of the proposed MMCI consists of three main components, shown from left to right: (1) Multi-relational Graph Construction, (2) Causal and Shortcut Attention Estimation, and (3) Disentanglement and Causal Intervention.
Dataset
CMU-MOSI
CMU-MOSEI
CH-SIMS
Train
1,284
16,326
1,368
Valid
229
1,871
456
Test
686
4,659
457
Batch Size
8
32
16
Epochs
50
15
50
Warm-up
✓
✓
✓
TABLE II: Dataset statistics, data splits, and hyper-parameter settings for the three multimodal datasets.
Method
CMU-MOSI
CMU-MOSEI
Acc2( ↑ )
F1( ↑ )
Acc7( ↑ )
MAE( ↓ )
Corr( ↑ )
Acc2( ↑ )
F1( ↑ )
Acc7( ↑ )
MAE( ↓ )
Corr( ↑ )
Graph-MFN § [ 58 ]
76.8 / 78.0
76.7 / 78.0
34.4
0.966
0.653
80.3 / 83.8
80.8 / 83.7
52.0
0.564
0.727
MTAG § [ 37 ]
80.8 / 83.1
80.4 / 82.8
25.9
0.879
0.706
81.5 / 84.6
81.7 / 84.5
52.6
0.559
0.735
HGraph-CL † [ 36 ]
84.3 / 86.2
84.6 / 86.2
-
0.717
0.799
84.5 / 85.9
84.5 / 85.8
-
0.527
0.769
ALMT [ 25 ]
82.4 / 84.5
82.2 / 84.4
45.9
0.741
0.776
80.7 / 84.6
81.3 / 84.7
52.7
0.543
0.761
TETFN † [ 23 ]
84.1 / 86.1
83.8 / 86.1
-
0.717
0.800
84.3 / 85.2
84.2 / 85.3
-
0.551
0.748
TABLE III: Comparison on the CMU-MOSI and CMU-MOSEI datasets. † denotes results reported in the original papers, § denotes results from [ 57 ] , while the remaining results are obtained from our own experiments. The best results are highlighted in bold, and the second-best results are underlined.
Method
CH-SIMS
Acc-5( ↑ )
Acc-3( ↑ )
Acc-2( ↑ )
F1( ↑ )
MAE( ↓ )
Corr( ↑ )
TFN [ 15 ]
39.3
65.1
78.4
78.6
0.432
0.591
LMF [ 66 ]
40.5
64.7
77.8
77.9
0.441
0.576
MFN [ 67 ]
39.5
65.7
77.9
77.9
0.435
0.582
MulT [ 5 ]
37.9
64.8
78.6
79.7
0.453
0.564
Self-MM [ 68 ]
41.5
65.5
80.0
80.4
0.425
0.595
TABLE IV: Comparison on the CH-SIMS dataset. † indicates that the results are taken from [ 60 ] , while the remaining results are reported in [ 65 ] .
Method
CMU-MOSI (OOD)
Acc7( ↑ )
Acc2*( ↑ )
Acc2( ↑ )
F1*( ↑ )
F1( ↑ )
MulT [ 5 ]
29.8
75.0
76.7
74.8
76.5
MAG-BERT [ 6 ]
39.9
75.6
77.3
75.5
77.3
MISA [ 30 ]
38.0
75.9
77.4
75.8
77.4
Self-MM [ 68 ]
40.2
76.7
78.1
76.7
78.1
CLUE [ 13 ]
41.8
78.8
79.9
78.8
79.9
TABLE V: Comparison on the OOD version of the CMU-MOSI dataset. Results marked with † are from the original papers, and § indicates results from our experiments. Other results are taken from [ 13 ] .
Fig. 5: Cross-dataset evaluation of methods trained on the CMU-MOSI and CMU-MOSEI training sets and tested on the corresponding opposite datasets, including both MOSI → MOSEI and MOSEI → MOSI transfer settings. The results of ITHP and C-MIB are obtained from our own implementations, while the remaining results are adopted from [ 51 ] .
Fig. 6: Parameter sensitivity analysis of disentanglement weight λ and causal intervention weight β . The gray dashed line denotes the best baseline.
Fig. 7: Effect of graph attention layers on CMU-MOSI (left) and CMU-MOSEI (right).
Method
Acc2( ↑ )
F1( ↑ )
Acc7( ↑ )
w/o Intra-modal Relation
83.6 / 85.5
83.6 / 85.5
45.7
w/o Inter-modal Relation
87.2 / 88.7
87.1 / 88.7
46.6
w/o Causal Attention
86.3 / 87.8
86.3 / 87.8
45.8
w/o Shortcut Attention
86.9 / 88.4
86.8 / 88.4
48.0
w/o Dual Attention Assignment
84.9 /86.1
85.0 / 86.2
45.4
w/o Disentanglement
86.6 / 88.4
86.5 / 88.4
47.3
TABLE VI: Ablation experiments on the CMU-MOSI dataset.
Method
Gaussian Noise
Salt-Pepper Noise
ϵ=1
ϵ=5
ϵ=10
ϵ=1
ϵ=5
ϵ=10
Noisy CMU-MOSI
C-MIB [ 8 ]
84.6
78.5
68.4
84.6
78.6
67.3
ITHP [ 3 ]
86.1
79.2
64.9
82.8
76.3
63.1
QMF [ 69 ]
84.1
76.3
69.5
83.1
80.0
63.7
MMCI (Ours)
87.0
81.2
68.2
86.3
73.7
67.5
TABLE VII: Acc2 performance under Gaussian and Salt-Pepper noise on the CMU-MOSI and CMU-MOSEI datasets.
Method
Acc2( ↑ )
F1( ↑ )
MAE( ↓ )
Corr( ↑ )
Self-MM b [ 68 ]
84.0
84.4
0.713
0.798
MMIM b [ 70 ]
84.1
84.0
0.700
0.800
MAG b [ 6 ]
84.2
84.1
0.712
0.796
C-MIB † b [ 8 ]
84.7
84.7
0.717
0.795
Self-MM d [ 68 ]
55.1
53.5
1.440
0.158
MMIM d [ 70 ]
85.8
85.9
0.649
0.829
TABLE VIII: Performance comparison on the CMU-MOSI dataset. Methods based on BERT and DeBERTa are marked with subscripts “b” and “d”, respectively. † indicates results obtained from our experiments, while other results are taken from [ 3 ] .
Modality
ITHP [ 3 ]
MMCI (Ours)
T
V
A
Acc7 ( ↑ )
Acc2 ( ↑ )
Acc7 ( ↑ )
Acc2 ( ↑ )
✓
x
x
42.3
85.3 / 87.0
46.0
86.0 / 87.9
✓
✓
x
43.5
85.4 / 87.5
48.0
86.9 / 88.5
✓
x
✓
46.7
84.8 / 86.7
47.4
86.3 / 88.2
✓
✓
✓
46.3
86.1 / 88.2
47.6
87.4 / 89.3
TABLE IX: Performance Comparison of MMCI and ITHP across various modality combinations on the CMU-MOSI dataset.
Method
Parameters
Training Time
Inference Time
EMOE [ 62 ]
317,295,771
15.99 s
1.11 s
GLoMo [ 32 ]
109,818,887
13.37 s
0.60 s
C-MIB [ 8 ]
190,930,180
16.09 s
0.65 s
ITHP [ 3 ]
184,883,706
21.16 s
0.89 s
MMCI (Ours)
186,461,076
20.56 s
0.71 s
TABLE X: Comparison of the number of parameters, training time, and inference time between MMCI and baselines on the CMU-MOSI dataset. All times are reported per epoch with batch size 8.
Fig. 8: Visualization of causal features (left) and non-causal features (right). HN: Highly Negative; N: Negative; WN: Weakly Negative; NT: Neutral; WP: Weak Positive; P: Positive; HP: Highly Positive.
Fig. 9: Comparison of FLOPs and peak GPU memory consumption. Peak memory is measured on CMU-MOSI with batch size 32.
Fig. 10: Qualitative case studies of MMCI on the CMU-MOSI dataset. For Case 1, (a)–(e) respectively show the predictions of ITHP, text-only ITHP, MMCI, MMCI without shortcut supervision, and MMCI without causal intervention. Cases 2 and 3 illustrate spurious correlations associated with visual attributes, while Cases 4 and 5 illustrate dataset-specific lexical associations. For Cases 4 and 5, the attention assigned by ITHP and MMCI to the selected biased word is normalized across the two models so that the reported ratios sum to one, representing relative inter-model attention rather than raw values.
Qingdao Institute of Software, College of Computer Science and Technology, China University of Petroleum (East China) · Shandong Key Laboratory of Intelligent Oil & Gas Industrial Software · The Hong Kong University of Science and Technology (Guangzhou)