cs.CVSep 30, 2026

LongEmo: Towards Emotion Understanding and Reasoning in Long Videos

Authors: Shuo Zhang, Yifan Zhou, Han Wang, Jinsong Zhang, Jingyu Li, Hongbing Li, Zhejun Zhang, Chengyi Zhao, +10 more

Organizations: BUPT · SJTU · THU · HIT · USTC · BNU · CUFE · PKU · UCAS

Abstract

While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion understanding and reasoning in long videos. It assesses progressive capabilities scaling from continuous scene interactions to complex episodic developments. Furthermore, we propose LongEmo, a novel memory-augmented agentic framework designed to tackle the immense challenges of long-range affective reasoning. LongEmo processes continuous video streams to construct an Event Memory Graph, explicitly modeling long-range dependencies and capturing emotional dynamics across discrete events. Given a question, the agent retrieves a query-relevant event stream from the graph, iteratively integrating multimodal memories and relational dependencies to deduce the final answer. Extensive evaluations of 17 representative methods reveal that they struggle significantly with emotion understanding and reasoning in long videos. In contrast, LongEmo achieves state-of-the-art performance, demonstrating the efficacy of its event-centric memory architecture.

Figures & tables

Explore similar work

CardsList
  1. AffectSeek: Agentic Affective Understanding in Long Videos under Vague User Queries

    May 7, 2026Zhen Zhang, Yuhang Yang, Yunxiang Jiang +5Long VideosImpact Analysis

  2. Agentic Very Long Video Understanding

    Jan 26, 2026Aniket Rege, Arka Sadhu, Yuliang Li +5Video AgentVideo Understanding