cs.CLSep 30, 2026

Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue

Authors: Yutong Hu, Jinho Choi

Organizations: Emory University Atlanta, GA

Abstract

Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensional evaluation on IEMOCAP, incorporating acoustic cues as natural language descriptions following the SpeechCueLLM approach. We evaluate six models spanning the LLaMA, GPT, and Qwen families under zero-shot prompting, few-shot prompting, and LoRA fine-tuning. LoRA fine-tuned LLaMA models substantially outperform prompt-engineered GPT models on both tasks despite GPT's larger scale, a gap we attribute to domain adaptation rather than model capacity. Our best model achieves a Valence CCC of 0.7822, a new state-of-the-art on IEMOCAP. Ablation studies confirm that textual audio descriptions meaningfully improve smaller models (+3.5 to 3.6 weighted F1) while contributing little for the largest model, suggesting audio cues are most valuable when linguistic capacity is limited. The performance asymmetry across VAD dimensions closely mirrors the annotator agreement hierarchy in IEMOCAP's own annotations.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

    Sep 17, 2026Hasindri Watawana, Sergio Burdisso, Esaú Villatoro-Tello +4Emotion RecognitionSpeech Language Models

  2. Multimodal Hidden Markov Models for Persistent Emotional State Tracking

    May 13, 2026Anamika Ragu, Aneesh JonelagaddaEmotion RecognitionValence-Arousal Estimation

  3. Quantifying the Affective Gap: A Zero-Shot Evaluation of LLMs on Fine-Grained Emotion Taxonomies

    Jul 1, 2026Lawrence Obiuwevwi, Krzysztof J. Rechowicz, Jessica M. Johnson +3Emotion RecognitionZero-Shot Robustness