Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
Authors: Lucas H. Ueda, João G. T. Lima, Paula D. P. Costa
Organizations: Dept. of Computer Engineering and Automation (DCA), Faculdade de Engenharia El´etrica e de Computac¸˜ao, and are part of the AI Lab., Recod.ai, Institute of Computing, Universidade Estadual de Campinas, UNICAMP
Speech Emotion Recognition (SER) in real-world scenarios remains challenging due to severe class imbalance and the prevalence of spontaneous, natural speech. While recent approaches leverage self-supervised learning (SSL) representations and multimodal fusion of speech and text, most existing methods apply supervision only at the final classification layer, limiting the discriminative power of intermediate representations. In this work, we propose Crab (Contrastive Representation and Multimodal Aligned Bottleneck), a bimodal Cross-Modal Transformer architecture that integrates speech representations from WavLM and textual representations from RoBERTa, together with a novel \textit{Multi Layer Contrastive Supervision} (MLCS) strategy. MLCS injects multi-positive contrastive learning signals at multiple layers of the network, encouraging emotionally discriminative representations throughout the model without introducing additional parameters at inference time. To further address data imbalance, we adopt weighted cross-entropy during training. We evaluate the proposed approach on three benchmark datasets covering different degrees of emotional naturalness: IEMOCAP, MELD, and MSP-Podcast 2.0. Experimental results demonstrate that Crab consistently outperforms strong unimodal and multimodal baselines across all datasets, with particularly large gains under naturalistic and highly imbalanced conditions. These findings highlight the effectiveness of \textit{Multi Layer Contrastive Supervision} as a general and robust strategy for SER. Official implementation can be found in https://github.com/AI-Unicamp/Crab.
Figures & tables
Dataset
Emotion
Count
Proportion
IEMOCAP
Angry
1103
0,20
Happy
1636
0,30
Neutral
1708
0,31
Sad
1084
0,19
MELD
Anger
1109
0,11
Disgust
271
0,03
Table 1: Emotion distribution across the training datasets. Since IEMOCAP does not provide official data partitions, the reported values correspond to the total number of available samples. For MELD, the validation and test splits preserve the same emotion distribution as the training set. MSP-Podcast 2.0 includes two test partitions with similar emotion distributions as training and a third test set that is balanced across emotion categories and is only accessible through an online evaluation platform.
Figure 1: Crab model architecture. The modules highlighted in green in each block have the Contrastive Guidance Leg attached and are optimized using the proposed MPCL contrastive loss.
Model
IEMOCAP
MELD
WAR
UAR
WAR
UAR
WavLM Baseline
0.6616
0.6832
0.3011
0.3046
FocalSer
0.6676
0.6943
0.3751
0.4085
MemoCMT
0.6964
0.7123
0.6490
0.4314
Medusa
0.6605
0.6778
0.4303
0.4548
Crab
0.7348
0.7585
0.6257
0.5565
Table 2: Comparison of model performance on IEMOCAP and MELD datasets.
Figure 2: Confusion matrices for MELD dataset across all evaluated models. Values in red on the principal diagonal represent values under a threshold of 0.2. Note that Crab has higher values on the principal diagonal compared to baselines with no emotion category having values under the 0.2 treshold. In particular, the proposed model is robust for the least represented classes ( Fear , Sadness , and Disgust ).
Figure 3: UAR × WAR performance for different values of α . The dashed horizontal lines indicate the proposed model trained only with CE and the Medusa model, using UAR as a reference.
Model
IEMOCAP
MELD
WAR
UAR
WAR
UAR
Speech
0.6153
0.6518
0.3724
0.3176
Text
0.7024
0.7198
0.6034
0.5381
Speech + Text
0.7348
0.7585
0.6257
0.5565
Table 3: Ablation study of the proposed model on IEMOCAP and MELD datasets.
Training setting
WAR
UAR
CE
0.5797
0.5253
CE + MPCL
0.6050
0.5375
MLS + CE
0.6077
0.5539
MLCS w/ SCL
0.6337
0.5534
MLCS w/o hierarchical lr
0.5904
0.5182
MLCS
0.6257
0.5565
Table 4: Comparison of training objectives for the proposed model on MELD. CE stands for Cross-Entropy and MLS denotes Multi Layer Supervision ( α=2 ).
SSL
Loss
WAR
UAR
wav2vec 2.0
CE
0.6023
0.5195
MLCS
0.6123
0.5440
HuBERT
CE
0.5851
0.5229
MLCS
0.6230
0.5453
WavLM
CE
0.5797
0.5253
MLCS
0.6257
0.5565
Table 5: Performance of Crab using different SSL on MELD.
SSL model
Version
Data
WAR
UAR
wav2vec 2.0
Base
960 hr
0.6115
0.5529
Large
960 hr
0.6123
0.5440
Large LV
60k hr
0.6092
0.5292
HuBERT
Base
960 hr
0.6126
0.5443
Large
60k hr
0.6230
0.5453
WavLM
Base
960 hr
0.6023
0.5470
Table 6: Ablation of different versions of wav2vec 2.0, Hubert and WavLM in Crab model.
Model
Test 1
Test 2
WAR
UAR
WAR
UAR
WavLM Baseline
0.5131
0.4105
0.4830
0.3131
FocalSer
0.2892
0.3360
0.1981
0.2819
MemoCMT
0.6239
0.3285
0.6058
0.2621
Medusa
0.4157
0.3549
0.4226
0.2772
Crab
0.5232
0.4390
0.4861
0.3382
Table 7: Comparison of model performance on Test 1 and Test 2 partitions of MSP-Podcast 2.0.
Model
Ensemble
WAR
Macro-F1
WavLM Baseline
No
0.3556
0.3293
FocalSER
No
0.3272
0.3148
MemoCMT
No
0.3184
0.2691
Medusa
No
0.3391
0.3210
StackingSER (*)
Yes
0.4128
0.4094
MATER (*)
Yes
0.4101
0.4097
Table 8: Performance of different models under naturalistic conditions on the MSP-Podcast 2.0 Test 3 set. Models with (*) have the results extracted from platform and were not reproduced.