cs.CVSep 29, 2026

Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning

Authors: Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao

Organizations: University of Science and Technology of China, Hefei · Harbin Institute of Technology, Weihai

Abstract

While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically proximal emotions based on fine-grained visual evidence, which could be decoupled as two limitations: 1) Insufficient Attribution. The global reasoning paradigm of conventional MLLMs severely dilutes fine-grained emotion cues, where subtle emotional states are usually implicitly encoded, thereby generating emotional misjudgments in complex scenarios. 2) Insufficient Discrimination. Existing methods could only identify regions generally associated with emotions, which fails to distinguish discriminative regions between semantically similar emotions, leading to ambiguous emotion judgements. To overcome these limitations, we present a training-free inference-time optimization framework, named Decoding Affective Nuances (DAN). Specifically, we propose a Hierarchical Emotional Reasoning Chain (HERC) that enhances the insufficient attribution by harmonizing fine-grained scene/object-level cues and performing a soft-gated reasoning. Furthermore, to discriminate between semantically proximal emotions, we design a Contrastive Discriminative Visual Pruning (CDVP), which isolates discriminative visual tokens to reason the final emotion category by computing the absolute discrepancy between the attention distributions of similar emotions. Performances on several benchmarks demonstrate that DAN significantly improves discrimination for affective nuances without consuming additional training resources, especially achieving +10.47% improvements with Qwen3-VL-8B-Instruct on WebEmo25 dataset that contains 25 fine-grained emotion categories.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

    Feb 27, 2026Yiyang Fang, Wenke Huang, Pei Fu +5Multimodal Large Language ModelsEmotion Recognition

  2. DSPO: Diversity-aware Subjective Policy Optimization for Robust Emotional Reasoning

    Sep 29, 2026Cheng Ye, Weidong Chen, Bingyan Xu +1Emotion RecognitionFrictive Policy Optimization

  3. NTDH: Complex Reasoning for Comprehensive Affective Analysis

    Aug 5, 2026Tianlei Zhu, Zhiwei Liu, Yuyan Wang +2Affective ComputingSentiment