Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8% on average cross-domain accuracy than EMO-R3.
Figures & tables
Figure 1: Comparison between traditional GRPO and our DSPO. 1) Traditional self-reflection mechanisms struggle to correct the internal biases of MLLMs. Counterfactual visual intervention mitigates hallucinations about visual evidence. 2) Diversity Distribution Reward addresses the issue where the sparse, discrete rewards of GRPO struggle to accommodate sentiment-close answers.
Figure 2: Overview of DSPO framework. The left part presents the construction of the emotional distribution prior and the input prompt. The middle part show two crucial counterfactual visual intervention gating and distribution-aligned emotional diversity reward modules. Finally, DSPO are jointly optimized with the original Format and Accuracy rewards under the GRPO framework.
Methods
Rollout
EmoSet I
Emotion6
WebEmo
Emotion6 I
EmoSet
WebEmo
AI
AO
A
LLaVA-1.5-7B
Zero-shot
-
52.77
48.32
25.56
48.32
52.77
25.56
50.55
38.05
42.22
SFT
-
56.04
54.21
42.39
54.21
56.04
42.39
55.13
48.76
50.88
Qwen2.5-VL-3B-Instruct
Zero-shot
-
51.55
50.00
40.65
50.00
51.55
40.65
50.77
45.71
47.40
SFT
-
77.15
34.51
17.75
69.53
26.45
37.65
73.34
29.09
43.84
Table 1: Comparison with GRPO variants and SOTA methods across in-domain and out-of-domain settings. The best and suboptimal results are highlighted in bold and underline, respectively.
Methods
KL ↓
JS ↓
Hmodel↑
En-MAE↓
SFT
1.940
0.457
0.395
0.397
GRPO
0.702
0.260
0.581
0.298
GRPO+ Ren
0.884
0.295
0.763
0.329
DAPO
1.059
0.328
0.524
0.348
EMO-R3
0.650
0.244
0.623
0.274
DSPO
0.372
0.169
0.696
0.209
Table 2: The results for distribution-based evaluation.
CVIG
DEDR
EmoSet I
Emotion6
WebEmo
A
×
×
75.45
57.91
49.40
60.92
✓
×
76.10
58.85
48.91
61.29
×
✓
75.60
64.25
52.60
64.15
✓
✓
76.65
66.87
54.00
65.84
Table 3: The ablation study for CVIG and DEDR modules.
Top- K
EmoSet I
Emotion6
WebEmo
A
1
72.35
59.20
50.15
60.57
3
75.90
62.74
51.20
63.28
5
76.65
66.87
54.00
65.84
7
74.45
60.81
50.99
62.08
Table 4: The ablation study for the number of candidate emotions.
Setting
EmoSet I
Emotion6
WebEmo
A
w/o DEDR
76.10
58.85
48.91
61.29
Center Distance
76.30
60.01
48.75
61.69
Pairwise Distance
74.74
58.40
47.89
60.34
LOO-Margin Distance
75.09
65.49
52.60
64.39
Combined Distance
76.65
66.87
54.00
65.84
Table 5: The discussion for diversity computation.
Setting
EmoSet I
Emotion6
WebEmo
A
w/o CVIG
75.60
64.25
52.60
64.15
Bounding Box
Random Box
67.80
57.30
46.59
57.23
w/o Area Penalty
71.55
60.89
49.10
60.51
Intervention
Mean Replace
72.11
60.60
49.70
60.80
Table 6: The discussion on CVIG module.
Figure 3: Impact for different reward weights.
Figure 5: Efficiency analysis visualization.
Figure 4: Case study between the most powerful method EMO-R3 and DSPO on the EmoSet dataset.
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of human emotions. Existing approaches based on supervised fine-tuning often suffer from limited generalization and poor interpretability, while reinforcement learning methods such as Group Relative Policy Optimization fail to align with the intrinsic characteristics of emotional cognition. To address these challenges, we propose Reflective Reinforcement Learning for Emotional Reasoning (EMO-R3), a framework designed to enhance the emotional reasoning ability of MLLMs. Specifically, we introduce Structured Emotional Thinking to guide the model to perform step-by-step emotional reasoning in a structured and interpretable manner, and design a Reflective Emotional Reward that enables the model to re-evaluate its reasoning based on visual-text consistency and emotional coherence. Extensive experiments demonstrate that EMO-R3 significantly improves both the interpretability and emotional intelligence of MLLMs, achieving superior performance across multiple visual emotional understanding benchmarks.
Yiyang Fang, Wenke Huang, Pei Fu +5
School of Computer Science, Wuhan University. · MiLM Plus, Xiaomi Inc.
While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically proximal emotions based on fine-grained visual evidence, which could be decoupled as two limitations: 1) Insufficient Attribution. The global reasoning paradigm of conventional MLLMs severely dilutes fine-grained emotion cues, where subtle emotional states are usually implicitly encoded, thereby generating emotional misjudgments in complex scenarios. 2) Insufficient Discrimination. Existing methods could only identify regions generally associated with emotions, which fails to distinguish discriminative regions between semantically similar emotions, leading to ambiguous emotion judgements. To overcome these limitations, we present a training-free inference-time optimization framework, named Decoding Affective Nuances (DAN). Specifically, we propose a Hierarchical Emotional Reasoning Chain (HERC) that enhances the insufficient attribution by harmonizing fine-grained scene/object-level cues and performing a soft-gated reasoning. Furthermore, to discriminate between semantically proximal emotions, we design a Contrastive Discriminative Visual Pruning (CDVP), which isolates discriminative visual tokens to reason the final emotion category by computing the absolute discrepancy between the attention distributions of similar emotions. Performances on several benchmarks demonstrate that DAN significantly improves discrimination for affective nuances without consuming additional training resources, especially achieving +10.47% improvements with Qwen3-VL-8B-Instruct on WebEmo25 dataset that contains 25 fine-grained emotion categories.
Cheng Ye, Weidong Chen, Zhaobo Qi +2
University of Science and Technology of China, Hefei · Harbin Institute of Technology, Weihai
We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit unfaithful behavior, often hallucinating modality-specific statements from other modalities. Building on these insights, we propose OPPO (Omni-Perception Policy Optimization), a reinforcement learning framework that explicitly optimizes multimodal perception. First, an Omni-Perception Reward decomposes ground-truth reasoning into fine-grained visual, acoustic, and emotion cues and rewards trajectories that semantically recover these cues. Second, an Omni-Perception Loss compares the policy under full and unimodally masked inputs, applying a KL penalty only to modality-specific evidence tokens to suppress cross-modal hallucination. We further introduce MEP-Bench, a diagnostic benchmark that quantifies utilization and faithfulness. Experiments show that OPPO achieves state-of-the-art performance on MER-UniBench and MME-Emotion, while substantially improving utilization and faithfulness scores on MEP-Bench, highlighting the importance of sufficient and faithful omni perception for multimodal emotion reasoning.
Zhiyuan Han, Beier Zhu, Wenwen Tong +6
University of Science and Technology of China · SenseTime · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center +1