Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.
Figures & tables
Predict the emotion expressed in the utterance from the provided information.
Transcript: ⟨ transcript ⟩
Pitch level: ⟨ 5 levels from very low to very high ⟩ .
Pitch variation: ⟨ 5 levels from very low to very high ⟩ .
Volume level: ⟨ 5 levels from very low to very high ⟩ .
Volume variation: ⟨ 5 levels from very low to very high ⟩ .
Speech rate: ⟨ 5 levels from very slow to very fast ⟩ .
Table 1: Prompt template with all three concept groups. Unselected groups are omitted. Volume denotes intensity and Predicted sex the speaker’s gender.
Zero-shot
Fine-tuned
Dataset
Model
T
A
TA
TAP
T
A
TA
TAP
CREMA-D
Qwen2.5
4.26
27.88
5.87
7.13
11.74
43.03
44.21
45.50
Qwen2.5-Omni
4.26
24.24
4.79
4.79
11.28
42.18
44.51
45.57
Llama 3.1
4.26
25.01
16.80
13.80
11.15
41.87
45.10
45.37
IEMOCAP
Qwen2.5
46.11
36.49
51.71
52.15
69.90
45.49
74.37
74.16
Qwen2.5-Omni
48.43
31.41
51.99
51.83
70.26
45.17
74.97
74.58
Table 2: Test Macro-F1 (%) of each combination of concept groups, zero-shot and fine-tuned.
Zero-shot
Fine-tuned
Dataset
Concepts
Audio
Concepts
Audio
CREMA-D
24.24
54.95
45.57
78.67
IEMOCAP
51.99
69.32
74.97
82.66
MELD
33.40
34.91
38.48
41.07
Table 3: Qwen2.5-Omni test Macro-F1 (%): highest-scoring concept combination from Table 2 versus direct audio input.
Figure 1: Changes in Macro-F1 relative to the fine-tuned TAP scores in Table 2 , after removing one acoustic concept while keeping the predictor fixed. Points show three-run means and error bars show ±1 standard deviation. IEMOCAP folds are averaged equally within each run.
Figure 2: Predictions that change when one concept is removed, as a share of all predictions of the first class (%), by the utterance’s level of the removed concept. Bars pool the three models and whiskers span the per-model values.
The Chinese University of Hong Kong, Hong Kong SAR, China · Institute of Software, Chinese Academy of Sciences, China · National Research Council Canada, Canada +1