Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8% on average cross-domain accuracy than EMO-R3.
Figures & tables
Figure 1: Comparison between traditional GRPO and our DSPO. 1) Traditional self-reflection mechanisms struggle to correct the internal biases of MLLMs. Counterfactual visual intervention mitigates hallucinations about visual evidence. 2) Diversity Distribution Reward addresses the issue where the sparse, discrete rewards of GRPO struggle to accommodate sentiment-close answers.
Figure 2: Overview of DSPO framework. The left part presents the construction of the emotional distribution prior and the input prompt. The middle part show two crucial counterfactual visual intervention gating and distribution-aligned emotional diversity reward modules. Finally, DSPO are jointly optimized with the original Format and Accuracy rewards under the GRPO framework.
Methods
Rollout
EmoSet I
Emotion6
WebEmo
Emotion6 I
EmoSet
WebEmo
AI
AO
A
LLaVA-1.5-7B
Zero-shot
-
52.77
48.32
25.56
48.32
52.77
25.56
50.55
38.05
42.22
SFT
-
56.04
54.21
42.39
54.21
56.04
42.39
55.13
48.76
50.88
Qwen2.5-VL-3B-Instruct
Zero-shot
-
51.55
50.00
40.65
50.00
51.55
40.65
50.77
45.71
47.40
SFT
-
77.15
34.51
17.75
69.53
26.45
37.65
73.34
29.09
43.84
Table 1: Comparison with GRPO variants and SOTA methods across in-domain and out-of-domain settings. The best and suboptimal results are highlighted in bold and underline, respectively.
Methods
KL ↓
JS ↓
Hmodel↑
En-MAE↓
SFT
1.940
0.457
0.395
0.397
GRPO
0.702
0.260
0.581
0.298
GRPO+ Ren
0.884
0.295
0.763
0.329
DAPO
1.059
0.328
0.524
0.348
EMO-R3
0.650
0.244
0.623
0.274
DSPO
0.372
0.169
0.696
0.209
Table 2: The results for distribution-based evaluation.
CVIG
DEDR
EmoSet I
Emotion6
WebEmo
A
×
×
75.45
57.91
49.40
60.92
✓
×
76.10
58.85
48.91
61.29
×
✓
75.60
64.25
52.60
64.15
✓
✓
76.65
66.87
54.00
65.84
Table 3: The ablation study for CVIG and DEDR modules.
Top- K
EmoSet I
Emotion6
WebEmo
A
1
72.35
59.20
50.15
60.57
3
75.90
62.74
51.20
63.28
5
76.65
66.87
54.00
65.84
7
74.45
60.81
50.99
62.08
Table 4: The ablation study for the number of candidate emotions.
Setting
EmoSet I
Emotion6
WebEmo
A
w/o DEDR
76.10
58.85
48.91
61.29
Center Distance
76.30
60.01
48.75
61.69
Pairwise Distance
74.74
58.40
47.89
60.34
LOO-Margin Distance
75.09
65.49
52.60
64.39
Combined Distance
76.65
66.87
54.00
65.84
Table 5: The discussion for diversity computation.
Setting
EmoSet I
Emotion6
WebEmo
A
w/o CVIG
75.60
64.25
52.60
64.15
Bounding Box
Random Box
67.80
57.30
46.59
57.23
w/o Area Penalty
71.55
60.89
49.10
60.51
Intervention
Mean Replace
72.11
60.60
49.70
60.80
Table 6: The discussion on CVIG module.
Figure 3: Impact for different reward weights.
Figure 5: Efficiency analysis visualization.
Figure 4: Case study between the most powerful method EMO-R3 and DSPO on the EmoSet dataset.