Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensional evaluation on IEMOCAP, incorporating acoustic cues as natural language descriptions following the SpeechCueLLM approach. We evaluate six models spanning the LLaMA, GPT, and Qwen families under zero-shot prompting, few-shot prompting, and LoRA fine-tuning. LoRA fine-tuned LLaMA models substantially outperform prompt-engineered GPT models on both tasks despite GPT's larger scale, a gap we attribute to domain adaptation rather than model capacity. Our best model achieves a Valence CCC of 0.7822, a new state-of-the-art on IEMOCAP. Ablation studies confirm that textual audio descriptions meaningfully improve smaller models (+3.5 to 3.6 weighted F1) while contributing little for the largest model, suggesting audio cues are most valuable when linguistic capacity is limited. The performance asymmetry across VAD dimensions closely mirrors the annotator agreement hierarchy in IEMOCAP's own annotations.
Figures & tables
Figure 1: Overall Pipeline
LLaMA-2-7B
LLaMA-3.1-8B
LLaMA-3.3-70B
Qwen3.5-35B-A3B
GPT 4o mini
GPT 5 mini
SpeechCueLLM Wu et al. (2025b)
Zero-shot
9.058
31.293
60.299
14.317
54.700
58.490
-
Few-shot
25.675
38.762
58.280
21.832
56.822
59.600
-
LoRA
73.196
71.818
72.122
68.493
-
-
72.021
Table 1: Discrete ER Performance of All Models (Weighted F1 score)
GPT-4o mini
GPT-5 mini
LoRA Fine-tuned LLaMA Model
Metric
Zero-shot
Few-shot
Zero-shot
Few-shot
2-7B
3.1-8B
3.370B
Valence
0.6024
0.6311
0.6630
0.6697
0.7672
0.7433
0.7822
Arousal
0.3393
0.3589
0.3926
0.3416
0.4406
0.4778
0.4653
Dominance
0.1458
0.1235
0.2990
0.3046
0.4388
0.4400
0.4413
Overall
0.3625
0.3712
0.4515
0.4386
0.5489
0.5537
0.5629
Table 2: VAD Evaluation Performance Comparison Across Different Models
Modalities
Valence CCC
Arousal CCC
Dominance CCC
Atmaja et al. Atmaja and Akagi (2021)
Audio, Text
0.553
0.579
0.456
Messaoudi et al. Messaoudi et al. (2024)
Audio
0.236
0.571
0.408
Awatef et al. Awatef et al. (2025)
Audio, Text
0.603
0.736
0.604
The Proposed
Audio, Text
0.782
0.465
0.441
Table 3: Comparison with Prior Work on VAD Evaluation on IEMOCAP
Model
Happy
Sad
Neutral
Angry
Excited
Frustrated
GPT-4o-mini ZS
49
70
53
32
53
60
GPT-4o-mini FS
46
74
55
40
56
60
GPT-5-mini ZS
42
70
57
45
63
61
GPT-5-mini FS
43
70
60
53
61
61
LLaMA-2-7B ZS
28
2
4
3
5
8
LLaMA-2-7B FS
40
3
5
4
9
7
Table 4: Per-emotion F1 score (%)
Happy ↔ Excited
Angry ↔ Frustrated
Sad ↔ Frustrated
Model
Hap → Exc
Exc → Hap
Ang → Fru
Fru → Ang
Sad → Fru
Fru → Sad
GPT-4o-mini ZS
7.6
29.8
75.3
3.1
22.0
3.4
GPT-4o-mini FS
12.5
22.7
65.9
6.3
15.1
6.0
GPT-5-mini ZS
30.6
18.4
60.0
6.3
18.8
4.5
GPT-5-mini FS
29.2
22.1
46.5
12.1
11.0
8.7
LLaMA-2-7B ZS
1.1
81.5
10.2
3.3
2.5
10.4
Table 5: Key Bidirectional Confusion Rates (%)
Reliability Measurements
Krippendorff’s alpha
Fleiss’ Kappa
Discrete Emotion Labels
0.279
0.273
Valence Score
0.680
0.316
Arousal Score
0.304
0.085
Dominance Score
0.271
0.014
Table 6: Reliability Measurements of IEMOCAP
LLaMA-2-7B
LLaMA-3.1-8B
LLaMA-3.3-70B
LoRA
73.196
71.818
72.122
LoRA without audio description
69.672
68.23
72.108
Table 7: The Impact of Audio Description on Discrete ER Performance of LLaMA Models(Weighted F1 score)
Figure 2: The Impact of Past VAD on GPT Models in Emotion Dimensional Evaluation
Figure 3: The Impact of Past VAD on LoRA Fine-tuned LLaMA Models in Emotion Dimensional Evaluation
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Example of The Proposed Pipeline
Figure 5: Example of the Audio Feature Description Generation
Dataset
Modalities
# Emotions
Continuous Measurements
Dimensional Metric
Dialogue
IEMOCAP
T, A, V
9
✓
VAD
✓
MELD Poria et al. (2019)
T, A, V
7
×
-
✓
EmotionLines Chen et al. (2018)
T
6
×
-
✓
DailyDialog Li et al. (2017)
T
7
×
-
✓
CMU-MOSEI Bagher Zadeh et al. (2018)
T, A, V
6
✓
0-3 Likert scale
×
MEISD Firdaus et al. (2020)
T, A, V
8
✓
1-3 Likert scale
✓
Appendix
Table 8: Existing Emotion Datasets
Property
Value
Number of sessions
5 dyadic sessions
Total dialogues
151
Total utterances
10,086
Avg. utterances per dialogue
≈ 66
Total audio duration
≈ 12 hours
Available modalities
Audio, Video, Text, Motion Capture
Appendix
Table 9: Summary Statistics of the IEMOCAP Dataset
Figure 6: Example of IEMOCAP Dataset
Context Window = 12
Context Window = 3
GPT-4o mini
GPT-5 mini
GPT-4o mini
GPT-5 mini
Dimension
Zero-shot
Few-shot
Zero-shot
Few-shot
Zero-shot
Few-shot
Zero-shot
Few-shot
Valence
0.5795
0.6193
0.6892
0.6861
0.6132
0.6329
0.6806
0.6857
Arousal
0.3787
0.3720
0.3959
0.3460
0.3481
0.3462
0.3710
0.3415
Dominance
0.1487
0.1504
0.3396
0.3424
0.1555
0.1340
0.3343
0.3333
Overall
0.3690
0.3806
0.4749
0.4582
0.3723
0.3710
0.4620
0.4535
Appendix
Table 10: VAD Evaluation Performance Comparison Across OpenAI Models and Prompting Strategies with Past VAD Values
Context Window = 12
Context Window = 3
Dimension
2-7B
3.1-8B
3.3-70B
2-7B
3.1-8B
3.3-70B
Valence
0.6647
0.7677
0.5368
0.6291
0.7160
0.4626
Arousal
0.3033
0.4411
0.3046
0.3412
0.3968
0.2034
Dominance
0.2662
0.4252
0.3585
0.3166
0.3719
0.2412
Overall
0.4114
0.5447
0.4000
0.5489
0.4949
0.3024
Appendix
Table 11: VAD Evaluation Performance Comparison Across LLaMA Models with LoRA Finetuning With Past VAD