cs.SDSep 30, 2026

From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models

Authors: Hezhao Zhang, Thomas Hain

Organizations: School of Computer Science, University of Sheffield Sheffield, United Kingdom

Abstract

Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.

Figures & tables

Explore similar work

CardsList
  1. Explainable and Trustworthy Speech Emotion Recognition Using Confidence Score and Reinforcement Learning Rectified Speech Emotion Descriptors

    Jun 12, 2026Youjun Chen, Xurong Xie, Mengzhe Geng +9Emotion RecognitionProsody

  2. Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

    Sep 17, 2026Hasindri Watawana, Sergio Burdisso, Esaú Villatoro-Tello +4Emotion RecognitionSpeech Language Models

  3. Titans-as-a-Layer: Test-Time Memory for Conversational Speech Emotion Recognition

    Jun 7, 2026Daniel Chen, Qicong Hu, Yang Xiao +2Speech Language ModelsLarge Audio Language Models