Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.
Figures & tables
Predict the emotion expressed in the utterance from the provided information.
Transcript: ⟨ transcript ⟩
Pitch level: ⟨ 5 levels from very low to very high ⟩ .
Pitch variation: ⟨ 5 levels from very low to very high ⟩ .
Volume level: ⟨ 5 levels from very low to very high ⟩ .
Volume variation: ⟨ 5 levels from very low to very high ⟩ .
Speech rate: ⟨ 5 levels from very slow to very fast ⟩ .
Table 1: Prompt template with all three concept groups. Unselected groups are omitted. Volume denotes intensity and Predicted sex the speaker’s gender.
Zero-shot
Fine-tuned
Dataset
Model
T
A
TA
TAP
T
A
TA
TAP
CREMA-D
Qwen2.5
4.26
27.88
5.87
7.13
11.74
43.03
44.21
45.50
Qwen2.5-Omni
4.26
24.24
4.79
4.79
11.28
42.18
44.51
45.57
Llama 3.1
4.26
25.01
16.80
13.80
11.15
41.87
45.10
45.37
IEMOCAP
Qwen2.5
46.11
36.49
51.71
52.15
69.90
45.49
74.37
74.16
Qwen2.5-Omni
48.43
31.41
51.99
51.83
70.26
45.17
74.97
74.58
Table 2: Test Macro-F1 (%) of each combination of concept groups, zero-shot and fine-tuned.
Zero-shot
Fine-tuned
Dataset
Concepts
Audio
Concepts
Audio
CREMA-D
24.24
54.95
45.57
78.67
IEMOCAP
51.99
69.32
74.97
82.66
MELD
33.40
34.91
38.48
41.07
Table 3: Qwen2.5-Omni test Macro-F1 (%): highest-scoring concept combination from Table 2 versus direct audio input.
Figure 1: Changes in Macro-F1 relative to the fine-tuned TAP scores in Table 2 , after removing one acoustic concept while keeping the predictor fixed. Points show three-run means and error bars show ±1 standard deviation. IEMOCAP folds are averaged equally within each run.
Figure 2: Predictions that change when one concept is removed, as a share of all predictions of the first class (%), by the utterance’s level of the removed concept. Bars pool the three models and whiskers span the per-model values.
Explainable and trustworthy speech emotion recognition (SER) remains a challenging task to date, largely due to the scarcity of SER data with reliable speech emotion descriptor (SED) labels, such as prosodic features and speaker traits. This paper presents a confidence score and reinforcement learning (RL) based on-the-fly SED rectification approach for post-training SER systems on automatically annotated SED labels. Experiments on IEMOCAP and MELD suggest that explainable SER systems incorporating the proposed confidence score and RL-based SED rectification approach consistently outperform baselines without data selection or SED rectification. The best performing system, which integrates both components, surpasses the baseline without data selection and SED rectification, achieving SER gains of 2.9% and 3.3% absolute (3.7% and 5.4% relative) on IEMOCAP and MELD benchmarks, respectively.
Youjun Chen, Xurong Xie, Mengzhe Geng +9
The Chinese University of Hong Kong, Hong Kong SAR, China · Institute of Software, Chinese Academy of Sciences, China · National Research Council Canada, Canada +1
SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state through a classification head, producing a label in one forward pass without modifying the backbone. Because this readout starts from the hidden state the model would otherwise decode, it gives a controlled comparison of generative and discriminative inference in an otherwise identical speechLLM. We keep the head a single linear layer, trading little accuracy for interpretability: each emotion becomes one direction in the LLM output token space, revealing associated tokens. On IEMOCAP, across two speechLLM architectures, it improves Macro F1 and removes hallucinations, with largest gains on realistic ASR transcripts. Our analysis reveals that these emotion directions encode indirect associations mirroring biases in web-scale text.
Speech emotion recognition (SER) is commonly formulated as utterance-level classification, although conversational emotion depends on a speaker's usual vocal range and the emotional context established by previous utterances. Speech-language models provide strong pretrained acoustic and semantic representations, and can adapts them to SER labels via finetune, but this mechanism still missing per-dialogue state. We study whether test-time neural memory can supply this missing context while leaving the large audio language models (LALMs) backbone intact. Building on Titans, we introduce a plug-and-play Memory-as-a-Layer (MAL) adapter that writes dialogue history into a small neural memory and reads it back as an audio-token-aligned residual update, avoiding changes to the host model's token positions. Across different audio LLMs and emotion recognition datasets evaluations, our design improves SER performs across different evaluation metrics, supporting test-time memory as a residual contextual mechanism for conversational SER.
Daniel Chen, Qicong Hu, Yang Xiao +2
Department of Computer Science, University of Auckland, Auckland, New Zealand · School of Computer and Information Technology, Melbourne, Australia