MissMAC-Bench: Building Solid Benchmark for Missing Modality Issue in Robust Multimodal Affective Computing
Authors: Ronghao Lin, Honghao Lu, Ruixing Wu, Aolin Xiong, Qinggong Chu, Qiaolin He, Sijie Mai, Haifeng Hu
Organizations: Sun Yat-Sen University Guangzhou, China · South China Normal University Guangzhou, China · Sun Yat-Sen University Pazhou Laboratory Guangzhou, Guangdong, China
Current Multimodal Affective Computing (MAC) systems heavily rely on the completeness of multiple modalities to accurately understand human's affective state. However, in real-world scenarios, the availability of modality data is often dynamic and uncertain, leading to substantial performance fluctuations due to the distribution shifts and semantic deficiencies of the incomplete multimodal inputs. Known as the missing modality issue, this challenge poses a critical barrier to the robustness and practical deployment of MAC models. To systematically quantify this issue, we introduce \textbf{MissMAC-Bench}, a comprehensive benchmark designed to establish fair and unified evaluation standards from the perspective of cross-modal synergy. Two guiding principles are proposed, including no missing prior during training, and one single model capable of handling both complete and incomplete modality scenarios, thereby ensuring better generalization. Moreover, to bridge the gap between academic research and real-world applications, our benchmark integrates evaluation protocols with both fixed and random missing patterns at the dataset and instance levels. Extensive experiments conducted on 3 widely-used language models across 4 datasets validate the effectiveness of diverse MAC approaches in tackling the missing modality issue. Our benchmark provides a solid foundation for advancing robust MAC and promotes the development of multimedia data mining. Our code is released in https://github.com/RH-Lin/MissMAC-Bench.
Figures & tables
Figure 1: Missing modality issue for robust multimodal affective computing in downstream application.
Figure 2: Illustration of random missing protocol, including dataset-level and instance-level evaluation.
Figure 3: Illustration of C&R dimensions, where μ and σ denotes the mean and variance of the performance during various missing circumstances.
Language Model
BERT-base / sBERT-base
DeBERTaV3-large
Fix
Models
CMU-MOSI
CMU-MOSEI
IEMOCAP
MELD
CMU-MOSI
CMU-MOSEI
IEMOCAP
MELD
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc ↑
wF1 ↑
Acc ↑
wF1 ↑
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc ↑
wF1 ↑
Acc ↑
wF1 ↑
{l}
BBFN τ
44.2
84.5
0.745
0.790
51.9
85.1
0.557
0.749
51.2
52.0
59.6
58.2
-
-
-
-
-
-
-
-
-
-
-
-
MTMD τ
46.7
83.5
0.764
0.782
53.1
85.2
0.544
0.753
58.0
57.5
57.7
58.0
-
-
-
-
-
-
-
-
-
-
-
-
ALMT τ
48.3
85.2
0.691
0.807
53.1
82.2
0.533
0.764
53.6
53.6
59.9
58.1
50.7
89.5
0.576
0.883
51.6
83.0
0.551
0.777
48.8
45.5
59.2
57.2
ITHP τ
43.8
82.6
0.762
0.784
52.1
84.7
0.551
0.756
56.5
56.2
56.3
56.3
49.2
87.6
0.592
0.860
50.3
83.5
0.542
0.786
50.2
49.3
57.2
56.3
Table 1: Performance comparison on fixed missing protocol with input modalities set as u∈{l}/{a,v} .
Language Model
BERT-base / sBERT-base
DeBERTaV3-large
Random
Models
CMU-MOSI
CMU-MOSEI
IEMOCAP
MELD
CMU-MOSI
CMU-MOSEI
IEMOCAP
MELD
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc ↑
wF1 ↑
Acc ↑
wF1 ↑
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc ↑
wF1 ↑
Acc ↑
wF1 ↑
MR=0.7
BBFN τ
25.3
61.3
1.189
0.430
41.4
56.9
0.794
0.492
36.3
33.1
28.2
32.4
-
-
-
-
-
-
-
-
-
-
-
-
MTMD τ
26.3
62.2
1.191
0.459
45.2
67.5
0.735
0.447
38.4
37.7
51.2
43.7
-
-
-
-
-
-
-
-
-
-
-
-
ALMT τ
25.9
61.9
1.169
0.462
32.5
65.1
1.098
0.309
35.5
32.0
50.7
44.2
27.4
55.7
1.176
0.499
34.4
49.2
0.891
0.469
34.7
29.7
48.2
43.5
ITHP τ
24.9
53.4
1.229
0.441
45.0
64.8
0.748
0.425
30.3
30.5
25.9
27.5
26.3
54.6
1.210
0.452
33.9
48.9
0.856
0.503
25.4
24.8
43.9
43.3
Table 2: Performance comparison on random missing protocol of the input modalities with missing rates MR set as 0.7 in dataset-level evaluation and MP set as 1.0 in instance-level evaluation.
Figure 4: Evaluation on C&R Dimension with BERT/sBERT on MSA and MER subtasks at both dataset-level ( MR ) and instance-level ( MP ) random missing protocols.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Incomplete multimodal input with diverse missing degrees and task generalization ability.
MAC Task
Sentiment Analysis
Emotion Recognition
Dataset
MOSI
MOSEI
IEMOCAP
MELD
Train
1,284
16,326
5,228
9,765
Valid
229
1,871
519
1,102
Test
686
4,659
1,622
2,524
Appendix
Table 3: Dataset distribution of MissMAC-Bench.
Language Model
BERT-base / sBERT-base
DeBERTaV3-large
Comp
Models
CMU-MOSI
CMU-MOSEI
IEMOCAP
MELD
CMU-MOSI
CMU-MOSEI
IEMOCAP
MELD
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc ↑
wF1 ↑
Acc ↑
wF1 ↑
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc7 ↑
F1 ↑
MAE ↓
Corr ↑
Acc ↑
wF1 ↑
Acc ↑
wF1 ↑
{l,a,v}
BBFN τ
45.9
84.5
0.750
0.790
53.9
86.0
0.522
0.788
60.2
60.2
59.6
58.1
-
-
-
-
-
-
-
-
-
-
-
-
MTMD τ
47.5
85.3
0.705
0.799
53.7
86.0
0.535
0.760
60.1
59.9
57.7
58.0
-
-
-
-
-
-
-
-
-
-
-
-
ALMT τ
49.0
84.8
0.694
0.806
54.3
84.0
0.519
0.779
63.6
63.5
60.7
59.5
51.5
89.3
0.559
0.884
55.0
86.2
0.494
0.807
67.7
67.7
59.1
58.0
ITHP τ
44.5
83.6
0.758
0.779
52.0
84.7
0.551
0.756
56.5
56.2
56.6
56.6
50.2
88.5
0.576
0.872
54.2
87.5
0.501
0.810
66.3
66.2
60.2
59.5
Appendix
Table 4: Performance comparison with complete input modalities set as u∈{l,a,v} .
Figure 6: Evaluation on C&R Dimension with DeBERTa on MSA and MER subtasks at both dataset-level ( MR ) and instance-level ( MP ) random missing protocols.