cs.CVSep 30, 2026

GLARE: Generating Listening Heads with Appropriate Reactions

Authors: Zikai Liao, Yumin Suh, Yi Ouyang, Yi-Lun Lee, Yi-Hsuan Tsai, Zhaozheng Yin

Organizations: Department of Computer Science, Stony Brook University · Atmanity Inc.

Abstract

While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents

    Apr 19, 2026Yu Zhang, Kaiyuan Shen, Yang LiAvatarsTurn-Taking

  2. Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation

    May 31, 2026Zhicheng Zhang, Lei Wang, Yu Zhang +1Speech-To-Text AlignmentSeed-Tts-Eval Benchmark

  3. EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold

    Sep 28, 2026Junjie Chen, Fei Wang, Kun Li +5AvatarsAudio-Video Generation