cs.CVSep 29, 2026

AffectReveal: Event-Grounded Emotion Recognition Beyond Visual Appearances

Authors: Yihao Qian, Runhao Zeng, Sicheng Zhao, Feng Liang, Hongmin Cai, Mingkui Tan

Organizations: South China University of Technology, Guangzhou, China · Shenzhen MSU-BIT University, Shenzhen, China · Tsinghua University, Beijing, China

Abstract

Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible reaction can convey different emotions depending on events beyond the input: tears, for example, may indicate grief or joy. We formulate Event-Grounded Emotion Recognition (EGER), where emotion recognition requires recovering the affect-determining event. We construct EGER-Bench, comprising 10,052 videos and 10,734 images across 11 emotions, two source domains, and four visual settings. A study with six annotators shows that event context raises human recognition accuracy from 33.96% to 72.08%, confirming that visual evidence alone is often insufficient. Semantic relevance alone does not solve EGER: a plausible event may imply the wrong emotion if its identity, focal-person role, relationship, or outcome is misinterpreted. We therefore propose AffectReveal, a tuning-free framework that first constructs and independently verifies evidence-grounded alternatives over these affect-critical factors. It then cross-checks the recovered event against face-masked in-media facts through bidirectional atomic evidence support, while retaining the original unmasked input for final prediction. Across three downstream models and four input settings, AffectReveal yields average UAR gains of 5.26--10.53 points. For three fine-tunable models, it also enables untuned models to outperform their fine-tuned visual-only counterparts in all 12 accuracy comparisons, without updating downstream parameters.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset

    Jul 11, 2026Zhiyan Zhang, Peipei Song, Jinpeng Hu +3Emotion RecognitionAffective Computing

  2. InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark

    Jun 1, 2026Shiyu Wang, Ziyu Liu, Chaoyi Yu +6Emotion RecognitionHuman-Annotated Benchmark

  3. Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory

    Sep 3, 2026Divyesh Bommana, Mohammad Saim, Tianyu JiangEmotion RecognitionEmotion