cs.AISep 28, 2026

Beyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction Generation

Authors: Hanzhong Zhang, Jindong Wang, Siyang Song

Organizations: Department of Computer Science, University of Exeter, Exeter, UK · Department of Data Science, William & Mary, Williamsburg, VA, USA

Abstract

Automatic human-like facial reaction generation (FRG) is essential for building intelligent systems that can engage in human-computer interaction (HCI). While diverse and context-appropriate facial reactions can reflect latent appraisal and affective processes in human interaction, most existing FRG methods rely on end-to-end architectures that directly map speaker behaviours to listener expressions without an explicit intermediate internal state. We reformulate FRG as generation mediated by a structured internal-state process and propose the \textbf{Intentional Agent}, which shifts FRG from direct stimulus-response mapping to stimulus-grounded generation through explicit intermediate states. To represent temporal internal-state evolution, we propose an internal dynamics model that integrates emotional drives with an iterative Inner Thought Flow (ITF) within a structured intermediate state used for subsequent generation. This state can continue to update during conversational silences. Furthermore, to bridge abstract internal states with physiological actions, we formulate FRG as a downstream affective mapping from this latent thought flow to facial expressions. Experiments on the REACT 2025 dataset show an FRDist of 72.39 and an FRDiv of 0.5057; perceptual plausibility is evaluated separately through blinded human ratings. A blinded human evaluation of 96 reactions found no significant difference in mean score between Full and ground truth (5.5275.527 vs.\ 5.1955.195, pHolm=.076p_{\mathrm{Holm}}=.076), while Full significantly outperformed Event-Triggered and Heuristic-Only (both pHolm<.001p_{\mathrm{Holm}}<.001). The Reaction Quality Scorer (RQS) correlated strongly with human judgements (Pearson r=.855r=.855; Spearman ρ=.821ρ=.821, both p<.05p<.05), supporting its use as an automatic metric. These results underscore the immense potential of endogenous dynamics in building highly autonomous, human-like agents.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ResDiffFRG: Residual Diffusion for Multiple Appropriate Facial Reaction Generation

    Sep 27, 2026Shizhe Liu, Jiayan Gu, Xiangyu Kong +1Facial Expression RecognitionResidual Conditional Diffusion Model

  2. REACT 2026: The Fourth Multiple Appropriate Facial Reaction Generation Challenge: Personalised MAFRG and Appropriate EEG Reaction Prediction

    Jun 6, 2026Siyang Song, Micol Spitale, Zijian Wu +11Facial Expression RecognitionPersonality

  3. MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations

    Jun 26, 2026Hejia Chen, Haoxian Zhang, Xu He +4Facial Expression RecognitionHuman Motion Generation