Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue
Organizations: Emory University Atlanta, GA
Abstract
Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensional evaluation on IEMOCAP, incorporating acoustic cues as natural language descriptions following the SpeechCueLLM approach. We evaluate six models spanning the LLaMA, GPT, and Qwen families under zero-shot prompting, few-shot prompting, and LoRA fine-tuning. LoRA fine-tuned LLaMA models substantially outperform prompt-engineered GPT models on both tasks despite GPT's larger scale, a gap we attribute to domain adaptation rather than model capacity. Our best model achieves a Valence CCC of 0.7822, a new state-of-the-art on IEMOCAP. Ablation studies confirm that textual audio descriptions meaningfully improve smaller models (+3.5 to 3.6 weighted F1) while contributing little for the largest model, suggesting audio cues are most valuable when linguistic capacity is limited. The performance asymmetry across VAD dimensions closely mirrors the annotator agreement hierarchy in IEMOCAP's own annotations.
Figures & tables
| LLaMA-2-7B | LLaMA-3.1-8B | LLaMA-3.3-70B | Qwen3.5-35B-A3B | GPT 4o mini | GPT 5 mini | SpeechCueLLM Wu et al. (2025b) | |
| Zero-shot | 9.058 | 31.293 | 60.299 | 14.317 | 54.700 | 58.490 | - |
| Few-shot | 25.675 | 38.762 | 58.280 | 21.832 | 56.822 | 59.600 | - |
| LoRA | 73.196 | 71.818 | 72.122 | 68.493 | - | - | 72.021 |
| GPT-4o mini | GPT-5 mini | LoRA Fine-tuned LLaMA Model | |||||
|---|---|---|---|---|---|---|---|
| Metric | Zero-shot | Few-shot | Zero-shot | Few-shot | 2-7B | 3.1-8B | 3.370B |
| Valence | 0.6024 | 0.6311 | 0.6630 | 0.6697 | 0.7672 | 0.7433 | 0.7822 |
| Arousal | 0.3393 | 0.3589 | 0.3926 | 0.3416 | 0.4406 | 0.4778 | 0.4653 |
| Dominance | 0.1458 | 0.1235 | 0.2990 | 0.3046 | 0.4388 | 0.4400 | 0.4413 |
| Overall | 0.3625 | 0.3712 | 0.4515 | 0.4386 | 0.5489 | 0.5537 | 0.5629 |
| Modalities | Valence CCC | Arousal CCC | Dominance CCC | |
| Atmaja et al. Atmaja and Akagi (2021) | Audio, Text | 0.553 | 0.579 | 0.456 |
| Messaoudi et al. Messaoudi et al. (2024) | Audio | 0.236 | 0.571 | 0.408 |
| Awatef et al. Awatef et al. (2025) | Audio, Text | 0.603 | 0.736 | 0.604 |
| The Proposed | Audio, Text | 0.782 | 0.465 | 0.441 |
| Model | Happy | Sad | Neutral | Angry | Excited | Frustrated |
| GPT-4o-mini ZS | 49 | 70 | 53 | 32 | 53 | 60 |
| GPT-4o-mini FS | 46 | 74 | 55 | 40 | 56 | 60 |
| GPT-5-mini ZS | 42 | 70 | 57 | 45 | 63 | 61 |
| GPT-5-mini FS | 43 | 70 | 60 | 53 | 61 | 61 |
| LLaMA-2-7B ZS | 28 | 2 | 4 | 3 | 5 | 8 |
| LLaMA-2-7B FS | 40 | 3 | 5 | 4 | 9 | 7 |
| Happy Excited | Angry Frustrated | Sad Frustrated | ||||
| Model | Hap Exc | Exc Hap | Ang Fru | Fru Ang | Sad Fru | Fru Sad |
| GPT-4o-mini ZS | 7.6 | 29.8 | 75.3 | 3.1 | 22.0 | 3.4 |
| GPT-4o-mini FS | 12.5 | 22.7 | 65.9 | 6.3 | 15.1 | 6.0 |
| GPT-5-mini ZS | 30.6 | 18.4 | 60.0 | 6.3 | 18.8 | 4.5 |
| GPT-5-mini FS | 29.2 | 22.1 | 46.5 | 12.1 | 11.0 | 8.7 |
| LLaMA-2-7B ZS | 1.1 | 81.5 | 10.2 | 3.3 | 2.5 | 10.4 |
| Reliability Measurements | ||
|---|---|---|
| Krippendorff’s alpha | Fleiss’ Kappa | |
| Discrete Emotion Labels | 0.279 | 0.273 |
| Valence Score | 0.680 | 0.316 |
| Arousal Score | 0.304 | 0.085 |
| Dominance Score | 0.271 | 0.014 |
| LLaMA-2-7B | LLaMA-3.1-8B | LLaMA-3.3-70B | |
|---|---|---|---|
| LoRA | 73.196 | 71.818 | 72.122 |
| LoRA without audio description | 69.672 | 68.23 | 72.108 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Modalities | # Emotions | Continuous Measurements | Dimensional Metric | Dialogue |
|---|---|---|---|---|---|
| IEMOCAP | T, A, V | 9 | ✓ | VAD | ✓ |
| MELD Poria et al. (2019) | T, A, V | 7 | × | - | ✓ |
| EmotionLines Chen et al. (2018) | T | 6 | × | - | ✓ |
| DailyDialog Li et al. (2017) | T | 7 | × | - | ✓ |
| CMU-MOSEI Bagher Zadeh et al. (2018) | T, A, V | 6 | ✓ | 0-3 Likert scale | × |
| MEISD Firdaus et al. (2020) | T, A, V | 8 | ✓ | 1-3 Likert scale | ✓ |
| Property | Value |
|---|---|
| Number of sessions | 5 dyadic sessions |
| Total dialogues | 151 |
| Total utterances | 10,086 |
| Avg. utterances per dialogue | 66 |
| Total audio duration | 12 hours |
| Available modalities | Audio, Video, Text, Motion Capture |
| Context Window = 12 | Context Window = 3 | |||||||
|---|---|---|---|---|---|---|---|---|
| GPT-4o mini | GPT-5 mini | GPT-4o mini | GPT-5 mini | |||||
| Dimension | Zero-shot | Few-shot | Zero-shot | Few-shot | Zero-shot | Few-shot | Zero-shot | Few-shot |
| Valence | 0.5795 | 0.6193 | 0.6892 | 0.6861 | 0.6132 | 0.6329 | 0.6806 | 0.6857 |
| Arousal | 0.3787 | 0.3720 | 0.3959 | 0.3460 | 0.3481 | 0.3462 | 0.3710 | 0.3415 |
| Dominance | 0.1487 | 0.1504 | 0.3396 | 0.3424 | 0.1555 | 0.1340 | 0.3343 | 0.3333 |
| Overall | 0.3690 | 0.3806 | 0.4749 | 0.4582 | 0.3723 | 0.3710 | 0.4620 | 0.4535 |
| Context Window = 12 | Context Window = 3 | |||||
|---|---|---|---|---|---|---|
| Dimension | 2-7B | 3.1-8B | 3.3-70B | 2-7B | 3.1-8B | 3.3-70B |
| Valence | 0.6647 | 0.7677 | 0.5368 | 0.6291 | 0.7160 | 0.4626 |
| Arousal | 0.3033 | 0.4411 | 0.3046 | 0.3412 | 0.3968 | 0.2034 |
| Dominance | 0.2662 | 0.4252 | 0.3585 | 0.3166 | 0.3719 | 0.2412 |
| Overall | 0.4114 | 0.5447 | 0.4000 | 0.5489 | 0.4949 | 0.3024 |