Organizations: South China University of Technology, Guangzhou, China · Shenzhen MSU-BIT University, Shenzhen, China · Tsinghua University, Beijing, China
Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible reaction can convey different emotions depending on events beyond the input: tears, for example, may indicate grief or joy. We formulate Event-Grounded Emotion Recognition (EGER), where emotion recognition requires recovering the affect-determining event. We construct EGER-Bench, comprising 10,052 videos and 10,734 images across 11 emotions, two source domains, and four visual settings. A study with six annotators shows that event context raises human recognition accuracy from 33.96% to 72.08%, confirming that visual evidence alone is often insufficient. Semantic relevance alone does not solve EGER: a plausible event may imply the wrong emotion if its identity, focal-person role, relationship, or outcome is misinterpreted. We therefore propose AffectReveal, a tuning-free framework that first constructs and independently verifies evidence-grounded alternatives over these affect-critical factors. It then cross-checks the recovered event against face-masked in-media facts through bidirectional atomic evidence support, while retaining the original unmasked input for final prediction. Across three downstream models and four input settings, AffectReveal yields average UAR gains of 5.26--10.53 points. For three fine-tunable models, it also enables untuned models to outperform their fine-tuned visual-only counterparts in all 12 accuracy comparisons, without updating downstream parameters.
Figures & tables
Figure 1: Appearance-based versus event-grounded emotion recognition. The visible reaction suggests sadness, whereas the recovered event context supports happiness. AffectReveal recovers and verifies the affect-determining event before prediction.
Figure 2: Overview of EGER-Bench, covering 11 emotions from news/interview and television-series sources. The lower panel summarizes dataset statistics, four visual settings, and class distributions across the two domains.
Evidence Condition
UAR
Acc.
WAF
Visual Input
27.80
33.96
33.89
Visual + Event Context
65.80
72.08
72.66
Table 1: Human performance with and without reference event context on 480 instances.
Figure 3: Overview of AffectReveal. The framework recovers external event context by retrieving a primary event and verifying evidence-anchored alternatives, then cross-checks the recovered event against non-facial in-media facts before emotion prediction.
Method
Affect-Critical Factor Correctness (%) ↑
Overall Event Recovery (%)
Event Identity
Focal Role
Relationship
Event Outcome
Match ↑
Partial ↑
Mismatch ↓
Primary Retrieval
16.35
36.67
23.33
12.25
13.25
21.55
65.20
Generic Multi-Search
15.44
38.90
25.53
14.32
12.96
25.92
61.12
Affect-Critical Event Recovery
16.52
40.64
31.06
15.03
13.95
28.91
57.14
Table 4: Event recovery quality under VID-V. Generic Multi-Search and Affect-Critical Event Recovery use the same maximum retrieval budget.
Figure 4: Qualitative comparisons between Primary Retrieval and AffectReveal. Although Primary Retrieval identifies the correct focal person, it retrieves a topically related but affectively incorrect event. AffectReveal corrects the event identity and outcome in the Olympic example, and the event identity and interpersonal relationship in the television-series example, leading to reference-aligned emotion predictions. Green text highlights the corrected event details.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Acc.
UAR
WAF
Qwen3.5-Omni-Plus
78.63
69.66
79.20
Qwen3.5-Omni-Plus+AffectReveal
81.51
69.60
81.61
Emotion-LLaMA
42.81
32.71
36.60
Emotion-LLaMA+AffectReveal
50.96
38.44
48.18
AffectGPT
48.44
20.57
49.47
AffectGPT+AffectReveal
53.84
23.69
56.73
Appendix
Table 6: Vision-only transfer results on MER2023.
Method
New/Interview (57%)
TV Series (43%)
All
UAR
Acc.
WAF
UAR
Acc.
WAF
UAR
Acc.
WAF
Primary Retrieval
28.87
32.40
27.23
23.64
23.92
23.67
28.09
28.74
25.09
AffectReveal
29.45
32.69
27.99
25.24
26.00
27.45
29.25
29.80
27.11
Appendix
Table 7: Emotion recognition across source domains under VID-V using Qwen2.5-Omni-7B. News/Interview and TV Series account for 57% and 43% of the evaluation samples, respectively.
Video emotion analysis is typically framed as a static classification problem, treating each clip as an independent labeled unit. However, such a formulation overlooks a key psychological fact: emotions change as a result of cumulative reactions to consecutive causal events. To bridge this gap, we introduce Dynamic Affective Reasoning, the first large-scale benchmark for viewer-centric affect transitions and causal reasoning over consecutive video events. DAR contains 15,087 videos and 36,908 event-aligned affective segments annotated with 27 emotion categories. Unlike existing video-based emotion datasets, DAR presents a new viewer-centric perspective on fine-grained emotional expressions and transitions, and provides dense, temporally grounded, and causally explicit reasoning chains. Based on DAR, we formally define three challenging tasks: affective segmentation, fine-grained emotion classification, and affective reasoning. Complementing this benchmark, we propose DAR-R1, a two-stage framework that combines supervised fine-tuning with Group Relative Policy Optimization. Experiments across 10+ MLLMs show that DAR-R1 sets a new state-of-the-art for dynamic affective reasoning, in terms of both emotional localization and affective reasoning. Project page: https://github.com/Zhang-Zhiyan/DAR.
Zhiyan Zhang, Peipei Song, Jinpeng Hu +3
University of Science and Technology of China, Hefei, China · Hefei University of Technology, Hefei, China
Visual emotion understanding requires models not only to recognize emotional states, but also to why they arise and perform higher-level cognitive reasoning. However, existing benchmarks mainly focus on emotion recognition, offering limited support for grounded understanding and response-oriented analysis. To address this gap, we introduce \textbf{InsightVQA}, a large-scale dataset for hierarchical visual question answering on emotion understanding and cognitive reasoning. Building from 351K images collected from six public sources, we apply a rigorous multi-stage filtering pipeline to curate 138K high-confidence images. Each image is annotated at three hierarchical levels: perception QA for emotion and valence recognition, grounded understanding QA constructed from visual trigger extraction through constraint-guided generation, and cognition QA centered on response intent prediction and sequential insight reasoning. In total, InsightVQA contains 725K QA pairs. We further present \textbf{InsightVQA-Bench}, a high-quality evaluation benchmark comprising 30K samples for fine-grained evaluation. To support evaluation, we introduce \textbf{InsightNet}, an emotion-tuned baseline for MLLMs. Results demonstrate that InsightVQA poses significant challenges for grounded emotion understanding and reasoning.
Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows angry. We introduce CHIARO, a 1,000 human-annotated sentence benchmark for contrastive emotion inference grounded in appraisal theory. Each scene describes one causal trigger eliciting a positive emotion in one person and a negative emotion in the other, drawn from a ten-class taxonomy. We benchmark seven frontier LLMs and four off-the-shelf emotion classifiers. The strongest LLM reaches 67.3 macro-F1, well below human agreement, while existing emotion classifiers score near chance. Beyond evaluation, CHIARO also serves as a training signal. When combined with an existing emotion corpus, the resulting downstream classifier improves on CHIARO itself and on six of ten external emotion benchmarks, which positions our dataset as a complementary signal for emotion recognition.