cs.LGSep 27, 2026

ResDiffFRG: Residual Diffusion for Multiple Appropriate Facial Reaction Generation

Authors: Shizhe Liu, Jiayan Gu, Xiangyu Kong, Siyang Song

Organizations: Department of Computer Science, University of Oxford, Oxford, United Kingdom · School of Artificial Intelligence and Big Data, Hefei University, Hefei, China · Department of Computer Science, University of Exeter, Exeter, United Kingdom

Abstract

In dyadic human speaker-listener conversations, the listener's facial reactions allows the speaker to accurately perceive the listener's emotional states. Since human facial reactions are non-deterministic, the ability to generate multiple appropriate human-like facial reactions is crucial for realistic human-agent interactions. Although diffusion models are naturally suited to such one-to-many generation, existing diffusion-based Multiple Appropriate Facial Reaction Generation (MAFRG) methods attempt to denoise random Gaussian initialisations directly into multiple appropriate facial reactions (AFRs). These random initialisations are usually not well-aligned with the target listener facial reaction, which requires complex denoising trajectories from these initialisations, and subsequently creates substantial opportunities for deviations away from the range of trajectories leading to appropriate AFRs. Given the inherent mimicry between the human listener's and speaker's facial behaviours, we address the above denoising trajectory issue by leveraging this strong prior. Specifically, we propose ResDiffFRG, a novel diffusion-based MAFRG framework that explicitly anchors the diffusion process to the speaker behaviour by defining its diffusion target as the residual between the speaker anchor and an AFR. The denoiser only needs to model the comparatively small, reaction-specific residual needed to transform this anchor into an AFR, rather than reconstructing the complete reaction from an unstructured state. Extensive experiments show that ResDiffFRG achieves large improvements in correlation-based appropriateness over existing methods. Our denoising trajectory analysis showed that even at the start of the denoising trajectory, ResDiffFRG already achieves a higher facial-reaction correlation score than the Gaussian Diffusion baseline does after completing 60% of its denoising trajectory.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. REACT 2026: The Fourth Multiple Appropriate Facial Reaction Generation Challenge: Personalised MAFRG and Appropriate EEG Reaction Prediction

    Jun 6, 2026Siyang Song, Micol Spitale, Zijian Wu +11Facial Expression RecognitionPersonality

  2. Beyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction Generation

    Sep 28, 2026Hanzhong Zhang, Jindong Wang, Siyang SongFacial Expression RecognitionAgentic Framework

  3. IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

    May 28, 2026Hao Wu, Xiangyang Luo, Hao Wang +3Precise Lip SynchronizationModel Fine-Tuning