Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents. And it must run online: each frame emitted from audio observed up to the current time, at interactive rates. Recent progress is dominated by diffusion models, which need many network evaluations per sample and are therefore a poor fit for streaming. We argue the cost is unnecessary in this domain. Audio-conditioned facial motion occupies a comparatively low-dimensional manifold, a regime where a single-pass GAN suffices. The obstacle is not capacity but stochastic structure. We show that a causal, time-invariant generator driven by i.i.d. noise cannot suppress its output spectrum over a band without collapsing its per-step innovation. We proposed FaceGAN, which dissolved the limitation by shaping the noise pathway acausally. Because the driving noise is synthetic, its future can be sampled now, so the audio-to-expression path stays causal, and the model supports fully causal operation. FaceGAN emits expression and head pose in a single forward pass per frame and matches or outperforms state-of-art approaches in generation quality. Being feed-forward with bounded attention windows, it generates indefinitely without drift.
Figures & tables
Figure 1: Streaming audio-to-expression with latency-free acausal noise shaping. Top: audio streams through our method called FaceGAN , which emits every frame in a single forward pass at 25 Hz with 40 ms of audio lookahead; the filmstrip shows rendered output. Bottom: the same causal generator driven by causal noise (left) produces jittery, muted motion, while our acausally shaped noise (right) yields smooth, expressive motion. Bottom traces are illustrative schematics.
Method
Arch.
# steps
Causal
Live-system
DiffPoseTalk Sun et al. (2024)
Diffusion
500
∼ (chunk-level)
✗
ARTalk Chu et al. (2025)
LLM
5
∼ (chunk-level)
✗
MemoryTalker Kim et al. (2025)
Transformer
1
✗
✗
Fallingwater Chu et al. (2026)
LLM
176
∼ (chunk-level)
✗
FaceGAN (ours)
GAN
1
✓ (frame-level)
✓
Table 1: Inference characteristics of the compared methods. Only our method is both single-pass and frame-level streaming; see Appendix D for how # steps is counted. Live-system means deployable for live animation: frame-level streaming with low latency, not throughput alone; see Appendix E for the measured benchmark.
Figure 2: System overview. Speech is encoded by a frozen Mimi encoder while an independent noise pathway samples past and future noise (costing no latency, as the noise is drawn rather than observed) and shapes it with two bidirectional layers into a temporally coherent motion prior. A six-layer causal conditioning stack fuses the audio features with the prior and emits 128-D expression codes plus head pose at 25 Hz in a single forward pass, with no iterative sampling.
Figure 3: Lip articulation on three utterances, with the target phoneme highlighted in each row label. Our mouth shapes track the ground-truth articulation faithfully with single-step inference, unlike the multi-step or non-causal baselines.
Method
FED ↓
FPD ↓
Similarity ↑
Sync Score ↑
Expr Var →1
Pose Var →1
Fallingwater
12.4824 †
1.6070
0.0976
0.6964
0.9803
1.1863 †
DiffPoseTalk
11.3546
1.4720 †
0.1016
0.5768 †
0.7591 †
0.5354
MemoryTalker
20.7047
3.7693
0.1328 †
0.5346
0.2171
0.0832
ARTalk
14.5405
1.2727
0.2911
0.2342
1.3998
1.0539
FaceGAN (ours)
12.1483
1.3335
0.1542
0.8767
0.9647
0.9921
Table 2: Quantitative comparison against Fallingwater Chu et al. (2026) , DiffPoseTalk Sun et al. (2024) , MemoryTalker Kim et al. (2025) and ARTalk Chu et al. (2025) . Expr/Pose Var are within-clip variance ratios against ground truth (how much the face and head move over time during an utterance), so best is closest to 1.0 . bold gold = best, underline silver = second-best, bronze† = third-best per column (ranking direction per arrow). Our method matches the multi-step or non-causal methods, achieving the best Sync Score and within-clip pose variance, while being the second best on every other metric, and operating in real time with single-step inference.
Comparison
Lip sync
Naturalness
vs. Fallingwater Chu et al. (2026)
45.3±7.1
55.3±7.1
vs. DiffPoseTalk Sun et al. (2024)
52.1±7.1
52.1±7.1
vs. ARTalk Chu et al. (2025)
45.6±9.1
61.1±9.0
vs. MemoryTalker Kim et al. (2025)
67.1±10.6
68.4±10.5
all baselines
50.5±4.1
57.1±4.1
Table 3: User study: percentage of comparisons preferring our method, with 95% confidence intervals, over 19 participants and 684 comparisons. 50% means the two are perceptually indistinguishable; bold marks intervals that exclude 50% .
causal noise
radius
att.
Sync ↑
FED ↓
Expr Div
Pose Div
Expr Var
✓
512
512
0.8654
-2%
11.59
-1%
0.89
-3%
0.84
+4%
0.92
-8%
✓
129
129
0.8905
+1%
11.60
-1%
0.79
-14%
0.64
-22%
0.93
-7%
✓
64
64
0.9074
+3%
12.61
+7%
0.68
-25%
0.52
-36%
0.92
-9%
2
5
0.8616
-2%
11.59
-1%
0.82
-11%
0.81
-0%
1.03
+3%
32
65
0.8902
+1%
11.64
-1%
0.91
-0%
0.91
+12%
1.02
+1%
ours
64
129
0.8810
± 0.0065
11.76
± 0.52
0.91
± 0.03
0.82
± 0.05
1.01
± 0.02
Table 4: Noise stack, causal versus acausal; causal rows and ours are means of three runs. “att.” is frames seen per layer: Rn for a causal stack, 2Rn+1 for an acausal one. Deltas against ours, whose ± is the run-to-run sd over its three runs; grey is inside that band ( <2σ ). For Expr Div and Pose Div only collapse is marked, since higher diversity is not in itself better.
Sync ↑
FED ↓
FPD ↓
Expr Div
Pose Div
Ours ( L=1 )
0.8810
± 0.0065
11.76
± 0.52
1.62
± 0.32
0.91
± 0.03
0.82
± 0.05
Lookahead
L=0
0.8698
-1%
13.08
+11%
2.38
+46%
1.03
+13%
0.99
+21%
L=2
0.8737
-1%
11.42
-3%
1.99
+22%
0.90
-1%
0.85
+4%
L=4
0.8721
-1%
11.82
+1%
2.24
+38%
0.90
-2%
0.81
-0%
Discriminators
no Du
0.6403
-27%
11.98
+2%
4.44
+173%
1.21
+33%
0.08
-90%
no Dc
0.0234
-97%
15.16
+29%
1.87
+15%
1.21
+33%
0.86
+5%
Table 5: Ablations on the 64-clip test set. Deltas are percentages against ours (mean of three runs), whose ± is the run-to-run sd over three runs; grey marks differences inside that band ( <2σ ). For Expr Div and Pose Div only collapse is marked, since higher diversity is not in itself better.
concurrent streams
1
256
1024
2048
4096
forward / latency (ms)
3.3 / 43.3
3.6 / 43.6
6.5 / 46.4
11.4 / 51.4
21.7 / 61.7
× real time
12.1
11.0
6.2
3.5
1.8
Table 6: Latency and throughput of the generator on one H200. Latency is the forward pass plus the 40 ms audio lookahead; real-time factor is against the 40 ms frame budget at 25 Hz.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
style
FED ↓
FPD ↓
Similarity ↑
Sync Score ↑
Expr Var
Pose Var
Fallingwater
✓
12.3718 †
1.7186
0.7052
0.6977
1.1058
1.2905 †
DiffPoseTalk
✓
9.8356
1.3321
0.5683
0.5100
0.6428 †
0.4661
MemoryTalker
✓
16.2776
3.9006
0.1486
0.5722 †
0.2651
0.0227
ARTalk
✓
12.4887
1.3644 †
0.5223 †
0.3121
1.4379
0.8492
FaceGAN (ours)
12.1483
1.3335
0.1542
0.8767
0.9647
0.9921
Appendix
Table 7: Style-conditioned comparison against Fallingwater Chu et al. (2026) , DiffPoseTalk Sun et al. (2024) , MemoryTalker Kim et al. (2025) and ARTalk Chu et al. (2025) ; ✓marks style conditioning. Expr/Pose Var are within-clip variance ratios against ground truth (how much the face and head move over time during an utterance), so best is closest to 1.0 . bold gold = best, underline silver = second-best, bronze† = third-best per column (ranking direction per arrow).
Figure 4: Head pose variation across three identities.
Method
Arch.
# steps
Causal
Live-system ready
Step counting
DiffPoseTalk
Diffusion
500
∼ (chunk-level)
✗
Full 500-level DDPM ancestral schedule per 100-frame chunk; one denoising pass per level, no step-skipping
ARTalk
LLM
5
∼ (chunk-level)
✗
Coarse-to-fine AR schedule over 5 token scales ([1, 5, 25, 50, 100]); one pass per scale
MemoryTalker
Transformer
1
✗
✗
Single deterministic forward pass over the full sequence
Fallingwater
LLM
176
∼ (chunk-level)
✗
Per-token AR over 176 tokens across 4 scales ([1, 25, 50, 100]), with style-bank cross-attention per step
FaceGAN (ours)
GAN
1
✓ (frame-level)
✓
One generator pass per frame; causal attention with 40 ms lookahead
Appendix
Table 8: How # steps is counted. For each method, # steps is the number of sequential forward evaluations of the core generative network needed to finalize one generation unit, independent of clip length. DiffPoseTalk runs the full 500-level DDPM ancestral schedule per 100-frame chunk: one denoising pass per level, with no DDIM or step-skipping. ARTalk generates each 100-frame chunk through a coarse-to-fine autoregressive schedule over 5 token scales (patch sizes [1, 5, 25, 50, 100]), invoking the network once per scale. Fallingwater decodes its 4-scale pyramid (patch sizes [1, 25, 50, 100]) token by token, invoking the network 176 times per chunk with style-bank cross-attention at each step. MemoryTalker produces the full sequence in a single deterministic forward pass (audio encoder → memory retrieval → decoder). FaceGAN emits one frame per forward pass of the generator.
Method
Decoding
Compute
Latency ↓
MemoryTalker
one-shot, non-causal
0.02 s
full utterance
DiffPoseTalk
4 s chunk, 500-step diffusion
11.2 s
5.56 s
Fallingwater
4 s chunk, 176-step AR
9.0 s
5.29 s
ARTalk
4 s chunk, 5-step AR
0.3 s
4.20 s
Ours
per-frame
2.4 s
83 ms
Appendix
Table 9: Compute is total wall-clock time to generate 30 s of motion (motion generation only, H200, fp32, batch 1); values below 30 s are faster than real time. Latency is the delay before the first frame can be emitted: for the chunked methods, 4.00 s of audio buffering + per-chunk compute + 0.16 s of centred Savitzky–Golay look-ahead; for ours, one 40 ms frame + 40 ms acausal look-ahead + 3.2 ms compute.
Recent advances in Audio-LLMs like GPT-4o have ushered in an era of conversational interaction with language models. Conversational avatars however, still seem robotic in facial expression and conversational flow, in part due to sequential stages of speech recognition, text generation, turn-based text response, speech synthesis, and audio driven facial animation. Based on our insight that audio-tokens produced by current Audio-LLMs carry sufficient information to reconstruct a plausible facial performance, we present TokTalk, a system that directly outputs expressive facial animation in real-time from streaming audio-tokens. We construct a novel audio-token to 3D facial motion dataset, on which TokTalk is trained using a Chunk-based Conditional Flow Matching model. A lightweight adaptation strategy allows our trained model to seamlessly connect to any token-based Audio-LLM at minimal computational overhead. Our chunk-based processing further enables parametric trade-off between latency and facial quality, shown through ablation studies. We further show that the real-time performance of TokTalk is comparable in latency to prior art solutions, and significantly favorable (via a perceptual study) in terms of quality, expressivity and control of the 3D facial performance. We showcase TokTalk's flexibility using a chatbot Avatar, a voice-driven user Avatar, and an animation Director's interface, as diverse audio-visual face applications.
Audio-driven facial animation is essential for immersive digital interaction, yet existing frameworks fail to reconcile real-time streaming with high-fidelity personalization. Current methods often rely on latency-inducing audio look-ahead, or require high user compliance to pre-encode static embeddings that fails to capture dynamic idiosyncrasies. We present an end-to-end causal framework for personalizing causal facial motion generation via dynamic multi-modal style retrieval, enabling ultra-low latency while uniquely leveraging unstructured style references. We introduce two key innovations: (1) a temporal hierarchical motion representation that captures global temporal context and high-frequency details while maintaining decoding causality, and (2) a multi-modal style retriever that jointly queries audio and motion to dynamically extract stylistic priors without breaking causality. This mechanism allows for scalable personalization with total flexibility regarding the number and contents of templates. By integrating these components into a causal autoregressive architecture, our method significantly outperforms state-of-the-art approaches in lip-sync accuracy, identity consistency, and perceived realism, supported by extensive quantitative evaluations and user studies.
Xuangeng Chu, Yu Han, Wei Mao +1
The University of Tokyo, Codec Avatars Lab, Meta, USA · Codec Avatars Lab, Meta, USA · Codec Avatars Lab, Meta, Pittsburgh, PA, USA
Streaming talking-head generation produces each frame as its driving audio arrives, yet fidelity and efficiency have so far pulled in opposite directions: end-to-end methods condition a video diffusion model on audio directly and achieve high quality but only at large scale, while cheaper two-stage methods generate an intermediate motion representation and trail in fidelity. We argue the cost of the former lies in the target of fusion: the video latent is dominated by identity, appearance and background, none of which audio bears on, so coupling audio to every pixel blurs detail and wastes capacity. We instead fuse conditions in a low-dimensional identity-disentangled motion space, routing audio and motion captions by their temporal granularity, and generate motion latents with a small causal autoregressive transformer that a pretrained diffusion renderer turns into video. Conditions thus control video transitively, and high fidelity no longer requires a large backbone. Streaming this decomposition needs both models to be causal, and the exposure-bias problem could be solved by self-forcing given a bidirectional teacher. But there is no such teacher in motion space. Our decoupled self-forcing distillation resolves both models under one frozen teacher: conditioned on motion, it distills the renderer into a block-causal student; unconditionally, it scores rendered rollouts against real videos, supervising motion by the video it produces. This lifts the fidelity ceiling from the motion generator onto the stronger renderer. The two models run as parallel causal streams, reaching 15.4 FPS at 1.3 s latency with no quality degradation.