Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
Authors: Lucas H. Ueda, João G. T. Lima, Paula D. P. Costa
Organizations: Dept. of Computer Engineering and Automation (DCA), Faculdade de Engenharia El´etrica e de Computac¸˜ao, and are part of the AI Lab., Recod.ai, Institute of Computing, Universidade Estadual de Campinas, UNICAMP
Speech Emotion Recognition (SER) in real-world scenarios remains challenging due to severe class imbalance and the prevalence of spontaneous, natural speech. While recent approaches leverage self-supervised learning (SSL) representations and multimodal fusion of speech and text, most existing methods apply supervision only at the final classification layer, limiting the discriminative power of intermediate representations. In this work, we propose Crab (Contrastive Representation and Multimodal Aligned Bottleneck), a bimodal Cross-Modal Transformer architecture that integrates speech representations from WavLM and textual representations from RoBERTa, together with a novel \textit{Multi Layer Contrastive Supervision} (MLCS) strategy. MLCS injects multi-positive contrastive learning signals at multiple layers of the network, encouraging emotionally discriminative representations throughout the model without introducing additional parameters at inference time. To further address data imbalance, we adopt weighted cross-entropy during training. We evaluate the proposed approach on three benchmark datasets covering different degrees of emotional naturalness: IEMOCAP, MELD, and MSP-Podcast 2.0. Experimental results demonstrate that Crab consistently outperforms strong unimodal and multimodal baselines across all datasets, with particularly large gains under naturalistic and highly imbalanced conditions. These findings highlight the effectiveness of \textit{Multi Layer Contrastive Supervision} as a general and robust strategy for SER. Official implementation can be found in https://github.com/AI-Unicamp/Crab.
Figures & tables
Dataset
Emotion
Count
Proportion
IEMOCAP
Angry
1103
0,20
Happy
1636
0,30
Neutral
1708
0,31
Sad
1084
0,19
MELD
Anger
1109
0,11
Disgust
271
0,03
Table 1: Emotion distribution across the training datasets. Since IEMOCAP does not provide official data partitions, the reported values correspond to the total number of available samples. For MELD, the validation and test splits preserve the same emotion distribution as the training set. MSP-Podcast 2.0 includes two test partitions with similar emotion distributions as training and a third test set that is balanced across emotion categories and is only accessible through an online evaluation platform.
Figure 1: Crab model architecture. The modules highlighted in green in each block have the Contrastive Guidance Leg attached and are optimized using the proposed MPCL contrastive loss.
Model
IEMOCAP
MELD
WAR
UAR
WAR
UAR
WavLM Baseline
0.6616
0.6832
0.3011
0.3046
FocalSer
0.6676
0.6943
0.3751
0.4085
MemoCMT
0.6964
0.7123
0.6490
0.4314
Medusa
0.6605
0.6778
0.4303
0.4548
Crab
0.7348
0.7585
0.6257
0.5565
Table 2: Comparison of model performance on IEMOCAP and MELD datasets.
Figure 2: Confusion matrices for MELD dataset across all evaluated models. Values in red on the principal diagonal represent values under a threshold of 0.2. Note that Crab has higher values on the principal diagonal compared to baselines with no emotion category having values under the 0.2 treshold. In particular, the proposed model is robust for the least represented classes ( Fear , Sadness , and Disgust ).
Figure 3: UAR × WAR performance for different values of α . The dashed horizontal lines indicate the proposed model trained only with CE and the Medusa model, using UAR as a reference.
Model
IEMOCAP
MELD
WAR
UAR
WAR
UAR
Speech
0.6153
0.6518
0.3724
0.3176
Text
0.7024
0.7198
0.6034
0.5381
Speech + Text
0.7348
0.7585
0.6257
0.5565
Table 3: Ablation study of the proposed model on IEMOCAP and MELD datasets.
Training setting
WAR
UAR
CE
0.5797
0.5253
CE + MPCL
0.6050
0.5375
MLS + CE
0.6077
0.5539
MLCS w/ SCL
0.6337
0.5534
MLCS w/o hierarchical lr
0.5904
0.5182
MLCS
0.6257
0.5565
Table 4: Comparison of training objectives for the proposed model on MELD. CE stands for Cross-Entropy and MLS denotes Multi Layer Supervision ( α=2 ).
SSL
Loss
WAR
UAR
wav2vec 2.0
CE
0.6023
0.5195
MLCS
0.6123
0.5440
HuBERT
CE
0.5851
0.5229
MLCS
0.6230
0.5453
WavLM
CE
0.5797
0.5253
MLCS
0.6257
0.5565
Table 5: Performance of Crab using different SSL on MELD.
SSL model
Version
Data
WAR
UAR
wav2vec 2.0
Base
960 hr
0.6115
0.5529
Large
960 hr
0.6123
0.5440
Large LV
60k hr
0.6092
0.5292
HuBERT
Base
960 hr
0.6126
0.5443
Large
60k hr
0.6230
0.5453
WavLM
Base
960 hr
0.6023
0.5470
Table 6: Ablation of different versions of wav2vec 2.0, Hubert and WavLM in Crab model.
Model
Test 1
Test 2
WAR
UAR
WAR
UAR
WavLM Baseline
0.5131
0.4105
0.4830
0.3131
FocalSer
0.2892
0.3360
0.1981
0.2819
MemoCMT
0.6239
0.3285
0.6058
0.2621
Medusa
0.4157
0.3549
0.4226
0.2772
Crab
0.5232
0.4390
0.4861
0.3382
Table 7: Comparison of model performance on Test 1 and Test 2 partitions of MSP-Podcast 2.0.
Model
Ensemble
WAR
Macro-F1
WavLM Baseline
No
0.3556
0.3293
FocalSER
No
0.3272
0.3148
MemoCMT
No
0.3184
0.2691
Medusa
No
0.3391
0.3210
StackingSER (*)
Yes
0.4128
0.4094
MATER (*)
Yes
0.4101
0.4097
Table 8: Performance of different models under naturalistic conditions on the MSP-Podcast 2.0 Test 3 set. Models with (*) have the results extracted from platform and were not reproduced.
Zero-shot cross-lingual speech emotion recognition (SER) remains challenging due to distribution mismatches across languages and the lack of emotion annotations in target language. Under such conditions, models trained solely on source-language data frequently suffer from degraded generalization when evaluated on unseen target languages. To address this limitation, we propose an emotion-discriminative representation learning method that integrates supervised contrastive learning and speaker adversarial learning. The contrastive learning promotes cross-lingual emotion alignment, while speaker adversarial learning suppresses speaker-related cues to encourage speaker-invariant representations. Experimental results under a zero-shot cross-lingual SER setting demonstrate that the proposed method significantly improves SER performance over conventional training strategies.
Jinyi Mi, Ding Ma, Tomoki Toda
Graduate School of Informatics, Nagoya University, Japan · Information Technology Center, Nagoya University, Japan
Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.
Hezhao Zhang, Thomas Hain
School of Computer Science, University of Sheffield Sheffield, United Kingdom
On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges remain: teacher reliability varies across batches, and logit-level distillation ignores inter-sample relational structure. We propose Adaptive Multi-teacher Relational Distillation (AMRD) to address both. A one-class SVM on each teacher's logit similarity matrix assigns per-batch weights favoring more coherent teachers. A relational distillation loss aligns teacher and student similarity matrices, capturing structure that logit matching misses. On IEMOCAP and CREMA-D datasets across four student architectures, AMRD outperforms single-teacher distillation baselines in most settings, and ablations confirm both components yield complementary gains.
Yuqi Li, Yi-Cheng Lin, Xianglong Wang +5
Department of Electrical Engineering, The City College of New York · National Taiwan University · Wyze Inc. +1