cs.CVOct 8, 2026

No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

Authors: Yu Han, Dejan Markovic, Alexander Richard, Wojciech Zielonka, Akshay Venkatesh, Cheng-hsin Wuu, Michael Zollhoefer

Organizations: Meta Reality Labs

Abstract

Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents. And it must run online: each frame emitted from audio observed up to the current time, at interactive rates. Recent progress is dominated by diffusion models, which need many network evaluations per sample and are therefore a poor fit for streaming. We argue the cost is unnecessary in this domain. Audio-conditioned facial motion occupies a comparatively low-dimensional manifold, a regime where a single-pass GAN suffices. The obstacle is not capacity but stochastic structure. We show that a causal, time-invariant generator driven by i.i.d. noise cannot suppress its output spectrum over a band without collapsing its per-step innovation. We proposed FaceGAN, which dissolved the limitation by shaping the noise pathway acausally. Because the driving noise is synthetic, its future can be sampled now, so the audio-to-expression path stays causal, and the model supports fully causal operation. FaceGAN emits expression and head pose in a single forward pass per frame and matches or outperforms state-of-art approaches in generation quality. Being feed-forward with bounded attention windows, it generates indefinitely without drift.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens

    May 29, 2026Qingcheng Zhao, Yifang Pan, Karan SinghTalking Head GenerationHuman Avatar Animation

  2. Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval

    Apr 26, 2026Xuangeng Chu, Yu Han, Wei Mao +1Audio-Driven Facial Animation

  3. Decoupled Self-Forcing Distillation for Streaming Talking Head Generation

    Sep 9, 2026Yanru An, Ruiyan Wang, Wenwu Wei +7Audio-Video GenerationDiffusion Model Distillation