While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically proximal emotions based on fine-grained visual evidence, which could be decoupled as two limitations: 1) Insufficient Attribution. The global reasoning paradigm of conventional MLLMs severely dilutes fine-grained emotion cues, where subtle emotional states are usually implicitly encoded, thereby generating emotional misjudgments in complex scenarios. 2) Insufficient Discrimination. Existing methods could only identify regions generally associated with emotions, which fails to distinguish discriminative regions between semantically similar emotions, leading to ambiguous emotion judgements. To overcome these limitations, we present a training-free inference-time optimization framework, named Decoding Affective Nuances (DAN). Specifically, we propose a Hierarchical Emotional Reasoning Chain (HERC) that enhances the insufficient attribution by harmonizing fine-grained scene/object-level cues and performing a soft-gated reasoning. Furthermore, to discriminate between semantically proximal emotions, we design a Contrastive Discriminative Visual Pruning (CDVP), which isolates discriminative visual tokens to reason the final emotion category by computing the absolute discrepancy between the attention distributions of similar emotions. Performances on several benchmarks demonstrate that DAN significantly improves discrimination for affective nuances without consuming additional training resources, especially achieving +10.47% improvements with Qwen3-VL-8B-Instruct on WebEmo25 dataset that contains 25 fine-grained emotion categories.
Figures & tables
Figure 1: Two coupled limitations underlying the difficulty of MLLMs in distinguishing semantically proximal emotions using fine-grained visual evidence. (a) Insufficient Attribution: Global and undifferentiated visual processing dilutes localized affective cues. (b) Insufficient Discrimination: Failure cases exhibit small confidence margins and concentrated confusion among emotions within the same semantic macro-category, indicating difficulty in resolving ambiguous candidates.
Figure 2: The illustration of our proposed DAN framework. The Hierarchical Emotion Reasoning Chain module (HERC) firstly identifies the scene-level and object-level emotion clues, and then performs a soft-gated reasoning for a stable preliminary classification. Then, the Contrastive Discriminative Visual Pruning module (CDVP) is proposed to handle the affectively nuanced samples, which guides MLLM to focus on the most discriminative regions and reselect the final emotion.
Dataset
Emotion6
EmoSet8
WebEmo7
WebEmo25
Abstract8
Average
Qwen2.5-VL-7B-Instruct
Zero-shot
58.33
56.98
47.70
22.45
23.68
41.83
Zero-shot-CoT
59.76
57.19
47.25
22.15
23.25
41.92
SEPM
61.58
57.94
49.10
22.85
26.32
43.56
DAN(Ours)
63.13
59.10
50.90
24.85
29.82
45.56
Qwen3-VL-4B-Instruct
Table 1: Comparison with state-of-the-art on various emotion datasets. The optimal results are denoted by boldface.
HERC
Scene
Object
CDVP
Emotion6
WebEmo25
Abstract8
×
×
×
55.21
21.10
29.38
✓
×
×
54.88
21.25
31.14
×
✓
×
56.40
22.50
28.51
✓
✓
×
57.24
23.15
31.58
×
×
✓
58.08
23.80
30.26
Table 2: Ablation study for key components.
Setting
Emotion6
WebEmo25
Abstract8
No-gating
54.55
20.90
28.51
Hard-gating
58.42
24.05
32.02
Soft-gating
59.09
24.80
33.77
Table 3: Discussion on soft-gated reasoning mechanism.
Setting
Emotion6
WebEmo25
Abstract8
Global Average
58.59
24.55
32.46
First-3 Average
57.58
24.30
30.26
Last Layer
58.42
24.50
31.58
Last-3 Average
59.09
24.80
33.77
Table 4: Discussion for strategies of attention scores.
Setting
Emotion6
WebEmo25
Abstract8
Only HERC
57.24
23.15
31.58
+ S
56.57
22.95
30.70
+ S∗
57.41
23.25
32.01
+ S∗ + Qctr
58.25
24.05
32.89
+ S∗ + Qctr + CDP
59.09
24.80
33.77
Table 5: Discussion for CDVP module.
Dataset
Emotion6
WebEmo25
Abstract8
Random
55.56
23.70
27.63
Query-related
56.57
23.95
29.82
FoE-related
57.91
24.25
31.58
Ours
59.09
24.80
33.77
Table 6: Results of different pruning strategies.
Figure 3: Parameter sensitivity analysis.
Figure 4: Case study of HERC.
Figure 5: Case study of CDVP.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
k
Emotion6
EmoSet
WebEmo7
WebEmo25
1
72.39
72.36
76.70
29.39
2
91.41
94.86
92.10
41.22
3
91.41
94.94
93.00
42.11
Appendix
Table A.1: Result for different truncation threshold k .
Figure A.2: Comparison of discriminative abilities.
Dataset
Emotion6
EmoSet8
Abstract8
SEPM
82.83
90.13
69.74
Scene
87.37
95.17
76.75
Object
86.03
92.91
71.93
HERC(Ours)
89.90
96.44
80.26
Appendix
Table A.2: Results of emotion polarity classification.
Setting
Emotion6
WebEmo25
Abstract8
Entire
58.59
24.60
32.01
Module-based
58.75
24.75
33.33
Alone
59.09
24.80
33.77
Appendix
Table A.3: Discussion for prompt injection strategies.