Beyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction Generation
Organizations: Department of Computer Science, University of Exeter, Exeter, UK · Department of Data Science, William & Mary, Williamsburg, VA, USA
Abstract
Automatic human-like facial reaction generation (FRG) is essential for building intelligent systems that can engage in human-computer interaction (HCI). While diverse and context-appropriate facial reactions can reflect latent appraisal and affective processes in human interaction, most existing FRG methods rely on end-to-end architectures that directly map speaker behaviours to listener expressions without an explicit intermediate internal state. We reformulate FRG as generation mediated by a structured internal-state process and propose the \textbf{Intentional Agent}, which shifts FRG from direct stimulus-response mapping to stimulus-grounded generation through explicit intermediate states. To represent temporal internal-state evolution, we propose an internal dynamics model that integrates emotional drives with an iterative Inner Thought Flow (ITF) within a structured intermediate state used for subsequent generation. This state can continue to update during conversational silences. Furthermore, to bridge abstract internal states with physiological actions, we formulate FRG as a downstream affective mapping from this latent thought flow to facial expressions. Experiments on the REACT 2025 dataset show an FRDist of 72.39 and an FRDiv of 0.5057; perceptual plausibility is evaluated separately through blinded human ratings. A blinded human evaluation of 96 reactions found no significant difference in mean score between Full and ground truth ( vs.\ , ), while Full significantly outperformed Event-Triggered and Heuristic-Only (both ). The Reaction Quality Scorer (RQS) correlated strongly with human judgements (Pearson ; Spearman , both ), supporting its use as an automatic metric. These results underscore the immense potential of endogenous dynamics in building highly autonomous, human-like agents.
Figures & tables
| Method | Appropriateness | Diversity | Synchrony | ||
| FRCorr ( ) | FRDist ( ) | FRDiv ( ) | FRVar ( ) | FRSyn ( ) | |
| GT ( Song et al., 2025 ) | 10.00 | 0.00 | 0.1876 | 0.0669 | 48.66 |
| B_Random ( Song et al., 2025 ) | 0.03 | 474.68 | 0.3342 | 0.1671 | 46.64 |
| B_Mime ( Song et al., 2025 ) | 0.52 | 206.02 | 0.0000 | 0.0766 | 43.70 |
| B_MeanFr ( Song et al., 2025 ) | 0.00 | 205.65 | 0.0000 | 0.0000 | 49.00 |
| Trans-VAE ( Song et al., 2025 ) | 0.30 | 181.72 | 0.0076 | 0.0083 | 49.00 |
| Model | Affect-Action Coherence | Pragmatic Utility | Temporal Intentionality | Role Compliance | Affective Appropriateness |
| GT | 3.00 | 1.93 | 2.43 | 3.07 | 2.64 |
| Qwen3.1-8B | 3.23 | 3.23 | 2.85 | 3.62 | 3.92 |
| Llama3.1-8B | 3.21 | 3.07 | 3.36 | 3.43 | 3.79 |
| GPT-4o | 2.71 | 2.43 | 3.50 | 3.36 | 3.43 |
| DeepSeek-V3 | 3.29 | 2.36 | 3.79 | 3.86 | 3.79 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Equation | Theoretical Basis | Functional Design Rationale |
| Eq. 1 | Self-sustaining dynamics in Appendix A | Separates internal and input-driven changes, making silence updates testable. |
| Eq. 2 | ITCMA ( Zhang et al., 2025 ) | Separates action choice, purpose, feasibility, and expected result. |
| Eqs. 3–5 | Predictive coding ( Rao and Ballard, 1999 ) | Gives an explicit prediction, comparison, and correction cycle. |
| Eq. 6 | Appraisal dynamics and EMA ( Marsella and Gratch, 2009 ) | Retains emotional continuity while allowing appraisal and error to change emotion. |
| Eq. 7 | VAD ( Russell and Mehrabian, 1977 ; Mehrabian, 1996 ) and embodied intentionality | Separates affective meaning from character-specific facial realisation. |
| Dimension | Pearson | Spearman | Direction Acc. | vs. |
| Valence | .512 | .484 | ||
| Arousal | .236 | .223 | ||
| Dominance | .151 | .148 |
| Condition | VaG | VeG | DG |
| Full | .04399 | .01369 | .09013 |
| Learned-Only | .04561 | .01477 | .09655 |
| Heuristic-Only | .05639 | .02124 | .12368 |
| Condition | Mean pairwise difference |
| Noise-Only | .01944 |
| State-Only | .00337 |
| Both | .02257 |
| Measure | State-Only | Noise-Only |
| Out-of-range AU proportion | .0097 | .0457 |
| Boundary saturation rate | .1788 | .2153 |
| Change speed | .0190 | .0277 |
| Trajectory jerk | .0126 | .0193 |
| Largest frame-to-frame change | .1776 | .2032 |
| Low-frequency energy |
| Method | AU-MAE | VaG | VeG |
| REACT-trained Full | .19081 | .02123 | .02825 |
| Seamless-adapted Full | .13973 | .01104 | .01294 |
| Method | FDD | SID |
| L2L ( Ng et al., 2022 ) | 25.95 | – |
| ARTalk ( Chu et al., 2025 ) | 30.62 | – |
| DualTalk ( Peng et al., 2025 ) | 43.58 | .690 |
| GDPO-Listener ( Jin et al., 2026 ) | 18.85 | 1.440 |
| Ours | 15.26 | .841 |
| Category | Mean | Median | Std | Samples | GT Score | Relative Ratio (%) |
| Overall | ||||||
| Session 0 | ||||||
| Session 1 | ||||||
| Session 2 | ||||||
| Session 3 | ||||||
| Session 4 |
| Metric | Human Pearson | Human Spearman |
| FRCorr | -.100 | -.040 |
| FRDist | .066 | .078 |
| FRSyn | .000 | .000 |
| VaG | .512 | .491 |
| VeG | -.057 | -.001 |
| DG | .189 | .136 |
| Predictors | Pearson | Spearman | MAE |
| Six automatic metrics | .480 | .432 | 1.461 |
| RQS | .847 | .716 | .851 |
| Six metrics + RQS | .879 | .786 | .745 |