EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold
Authors: Junjie Chen, Fei Wang, Kun Li, Yiqi Nie, Xun Yang, Yanbin Hao, Linfeng Zhang, Meng Wang
Organizations: Hefei University of Technology · EPIC Lab, Shanghai Jiao Tong University · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center · United Arab Emirates University · Anhui University · University of Science and Technology of China · SAI, Shanghai Jiao Tong University
Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAvatar, a causal generator that uses test-time training to adapt to user face video and dyadic audio during interaction. Its dyadic context prediction objective provides a self-supervised learning signal from audiovisual context without target motion labels at test time. Persistent fast weights accumulate these updates within each conversation to guide motion generation, while transient jaw adaptation responds to current audiovisual context. Predicted speech activity controls how persistent adaptation guides motion. We also introduce InterHead-Bench, a unified 455.95-hour benchmark built from single-view and dual-view conversation videos. Experiments show improved conversational motion statistics over strong baselines. On the hardest out-of-distribution split, generation improves as conversations unfold, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
Figures & tables
Figure 2: EvolvingAvatar learns from the ongoing conversation at deployment without motion labels. Dyadic context prediction drives two complementary forms of adaptation: persistent fast weights capture conversational regularities, while transient jaw adaptation tracks current articulation. Regional latent routing combines their predictions through a shared causal decoder.
Figure 3: dialog3d-factory processes dual-view recordings with separate audio and single-view recordings with mixed audio. A shared workflow extracts face video, participant speech, FLAME motion, transcripts, speaking and interaction states, and turn statistics on a common timeline.
Method
Venue
Context
Parameter Size
U-Video
U-FLAME
U-Audio
A-Audio
Total (M)
Trainable (M)
DiffPoseTalk ( Sun et al., 2024 )
TOG ’24
✗
✗
✗
✓
129.32
110.55
ARTalk ( Chu et al., 2025 )
SIGGRAPH Asia ’25
✗
✗
✗
✓
382.87
37.89
UniLS ( Chu et al., 2026 )
CVPR ’26
✗
✗
✓
✓
422.88
27.38
DualTalk ( Peng et al., 2025 )
CVPR ’25
✗
✓
✓
✓
647.27
638.85
EvolvingAvatar (Ours)
This work
✓
✗
✓
✓
159.57
54.82
Table 1: Generation-time conditioning context and parameter sizes of the compared methods. U and A denote User and Avatar, respectively. Parameter counts are reported in millions (M).
Figure 4: InterHead-Bench composition by split, including scale and the distributions of interaction duration and conversational turns.
Method
FD ↓
P-FD ↓
MSE ↓
rPCC ↓
SID ↑
PDD ↓
JDD ↓
EXP
JAW (×103)
POSE (×102)
EXP
JAW (×103)
POSE (×102)
EXP (×101)
JAW (×103)
POSE (×102)
EXP (×102)
JAW (×101)
POSE (×101)
EXP
JAW
POSE
In-Distribution Test Set (ID)
Speaking
DiffPoseTalk ( Sun et al., 2024 )
48.45
22.62
24.60
50.30
22.93
25.47
6.01
10.01
11.61
20.11
1.95
3.31
1.84
1.36
1.35
9.38
2.88
ARTalk ( Chu et al., 2025 )
21.11
3.66
18.89
22.12
3.75
19.52
2.77
1.70
8.68
18.83
1.30
3.29
2.18
1.69
1.38
5.67
0.95
UniLS ( Chu et al., 2026 )
15.99
2.36
12.31
17.05
2.45
12.76
2.32
1.30
5.36
14.73
1.07
3.37
2.84
2.20
1.13
7.36
0.84
Table 2: Parameter-space comparison. Speaking and listening results. Green bold and blue underlined values mark the best and second best in each split and state. Lower is better except SID.
ID
OOD
OOD-Hard
Method
Speaking
Listening
Speaking
Listening
Speaking
Listening
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
DiffPoseTalk ( Sun et al., 2024 )
39.26
18.78
4.10
40.74
19.97
4.31
38.37
18.10
4.02
40.82
18.60
4.11
37.57
17.94
3.97
44.10
18.28
4.01
ARTalk ( Chu et al., 2025 )
32.97
9.33
2.21
35.90
11.77
2.75
30.82
8.18
1.98
30.92
9.85
2.32
20.91
10.79
2.49
19.84
10.10
2.34
UniLS ( Chu et al., 2026 )
29.21
7.62
1.79
31.11
9.92
2.21
27.37
7.29
1.71
28.19
8.91
1.97
21.84
10.28
2.27
21.78
10.82
2.37
DualTalk ( Peng et al., 2025 )
68.51
8.12
1.84
61.57
9.26
2.03
64.35
7.29
1.65
55.63
8.03
1.75
42.83
7.65
1.67
38.31
7.64
1.67
Table 3: Mesh-space comparison. Speaking/listening results. LVE and MHD are in millimeters. Green bold and blue underline denote the best and second best per split and state. Lower is better.
Table 7Figure 8
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Restrained neck and mouth motion across activities. EvolvingAvatar preserves the reference’s slightly raised chin and small speaking openings, with a more restrained mouth shape in the intervening listening columns. UniLS opens the mouth too widely at several of these selected instants, while DiffPoseTalk and ARTalk introduce marked downward or sideways neck poses.
Figure 8: Smiling while listening. EvolvingAvatar preserves the reference’s smile and moderate mouth opening in columns 4–6, across speaking and listening. UniLS and DualTalk show weaker openings there, while DiffPoseTalk exaggerates mouth opening in later sampled listening frames.
Figure 9: From listening to speaking. EvolvingAvatar combines the tilted listening posture in columns 1–5 with the larger speaking mouth openings in columns 7–10. UniLS changes the mouth opening less consistently with Avatar GT, while DualTalk largely retains a neutral expression.
Figure 10: Coordinating changes in expression and neck pose. EvolvingAvatar follows the reference’s slight downward tilt during listening in columns 5–8 and sideways lean in the final speaking frames. UniLS exaggerates the downward tilt during listening, while DualTalk remains largely frontal and DiffPoseTalk adds a broad grin where the reference expression is restrained.
Figure 11: Moderate mouth openings and neck motion. EvolvingAvatar captures the speaking openings and tilt in columns 2–4, then retains a listening smile. UniLS exaggerates those speaking openings and closes the eyes in columns 6–9, while DiffPoseTalk adds a larger sideways lean.
Figure 12: Matching neck orientation and articulation together. EvolvingAvatar retains the reference’s near-upright orientation and modest changes in mouth opening across the selected speaking and listening instants. UniLS exaggerates the upward tilt or mouth opening at several instants, while ARTalk introduces a pronounced downward tilt that is absent from the reference.
Figure 13: An expressive listener. EvolvingAvatar follows the reference’s open-mouth smile and upward, sideways orientation during both activities. UniLS largely closes the mouth in the first five listening frames, while ARTalk tilts the face downward and DualTalk retains a more frontal pose.
Figure 14: From varied articulation to a listening smile. EvolvingAvatar captures the speaking openings in columns 1–6 and closed-mouth listening smile in columns 7–10. UniLS exaggerates upward tilt and early openings, while DiffPoseTalk bows and turns the face during listening.
Figure 15: Changing articulation within each activity. EvolvingAvatar follows the reference’s alternation between smaller and larger mouth openings while retaining a moderate neck pose across the sampled instants. UniLS raises the chin too far and opens the mouth in several near-closed reference frames, while DiffPoseTalk over-opens the mouth across most of the displayed columns.
Figure 16: Subtle articulation with an upright neck pose. EvolvingAvatar preserves the near-upright neck pose and the mild increase in mouth opening in the later columns. UniLS and ARTalk lower the face, while DiffPoseTalk replaces the small final openings with larger ones.
Figure 17: A smile across speaking and listening. EvolvingAvatar combines a near-frontal pose with larger speaking openings and a softer listening smile, retaining visible expression across both activities. UniLS and DualTalk understate the reference speaking expression in the first five columns, while DiffPoseTalk produces excessive mouth opening in the subsequent listening frames.
Figure 18: Expressive listening with an upright pose. EvolvingAvatar captures the listening smile in columns 2–4 and stays near the reference orientation during speaking. UniLS and ARTalk lower the face during speaking, while DualTalk shows less of the reference variation in mouth shape.
Figure 19: A frontal smile with varying mouth opening. EvolvingAvatar preserves the reference’s upright pose and smile, including the larger openings in columns 8–9. UniLS and ARTalk lower the face, while DiffPoseTalk adds a pronounced sideways turn and wider grin in the final columns.
Figure 20: Expression changes with restrained neck motion. EvolvingAvatar follows the reference’s smaller openings in columns 6–7 and broader speaking smile in columns 8–10. UniLS and ARTalk bow the head, while DualTalk shows less change in mouth shape than the reference.
Figure 21: Reducing articulation during listening. EvolvingAvatar remains nearly upright while following the smaller listening openings in columns 7–10. UniLS and ARTalk lower the face, while DualTalk retains a parted mouth and DiffPoseTalk produces a wider final opening than Avatar GT.
Method
FD ↓
P-FD ↓
MSE ↓
rPCC ↓
SID ↑
PDD ↓
JDD ↓
EXP
JAW (×103)
POSE (×102)
EXP
JAW (×103)
POSE (×102)
EXP (×101)
JAW (×103)
POSE (×102)
EXP (×102)
JAW (×101)
POSE (×101)
EXP
JAW
POSE
In-Distribution Test Set (ID)
DiffPoseTalk ( Sun et al., 2024 )
45.68
25.60
23.10
47.40
25.90
23.99
5.83
11.35
11.38
18.94
1.90
2.90
1.85
1.45
1.48
9.87
3.04
ARTalk ( Chu et al., 2025 )
22.99
3.94
21.71
23.91
4.03
22.32
2.93
1.86
9.60
21.72
1.57
3.18
1.69
1.55
1.17
5.93
1.26
UniLS ( Chu et al., 2026 )
17.19
3.05
12.95
18.20
3.18
13.37
2.45
1.79
5.60
15.86
1.48
3.16
2.31
1.90
1.04
7.52
1.03
DualTalk ( Peng et al., 2025 )
17.22
4.63
11.52
17.35
4.68
11.60
1.79
1.81
4.05
23.17
1.81
3.66
0.32
0.57
0.28
11.34
2.09
Appendix
Table 6: Parameter-space comparison. Results over all valid frames. Green bold and blue underline indicate the best and second-best values within each split. Lower is better except for SID.
Method
ID
OOD
OOD-Hard
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
DiffPoseTalk ( Sun et al., 2024 )
40.03
19.77
4.28
40.37
18.52
4.09
44.83
18.31
4.02
ARTalk ( Chu et al., 2025 )
35.22
11.17
2.62
30.51
9.35
2.22
19.51
10.09
2.34
UniLS ( Chu et al., 2026 )
30.39
9.33
2.10
27.64
8.33
1.87
21.65
10.65
2.34
DualTalk ( Peng et al., 2025 )
66.70
8.87
1.96
61.46
7.71
1.71
40.10
7.60
1.66
EvolvingAvatar (Ours)
37.91
9.83
2.18
31.53
8.74
1.92
17.87
8.71
1.89
Appendix
Table 7: Mesh-space comparison. All-frame results. LVE and MHD are measured in millimeters. Green bold and blue underlined values mark the best and second best in each split. Lower is better.
Method
FD ↓
P-FD ↓
MSE ↓
rPCC ↓
SID ↑
PDD ↓
JDD ↓
EXP
JAW (×103)
NECK (×102)
EXP
JAW (×103)
NECK (×102)
EXP (×101)
JAW (×103)
NECK (×102)
EXP (×102)
JAW (×101)
NECK (×101)
EXP
JAW
NECK
In-Distribution Test Set (ID)
Speaking
UniLS ( Chu et al., 2026 )
15.99
2.36
12.31
17.05
2.45
12.76
2.32
1.30
5.36
14.73
1.07
3.37
2.84
2.20
1.13
7.36
0.84
DualTalk ( Peng et al., 2025 )
16.80
4.17
12.05
16.97
4.23
12.16
1.74
1.59
4.21
20.69
1.31
4.13
0.31
0.49
0.25
11.37
2.08
Ours (audio)
13.71
2.21
9.34
14.90
2.35
9.83
2.15
1.69
4.70
15.38
1.22
3.29
2.47
2.13
1.42
6.86
0.95
Appendix
Table 8: Context signals in parameter space. Speaking/listening results. Green bold and blue underlining denote the best and second best per split and state. Lower is better except SID.
ID
OOD
OOD-Hard
Method
Speaking
Listening
Speaking
Listening
Speaking
Listening
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
UniLS ( Chu et al., 2026 )
29.21
7.62
1.79
31.11
9.92
2.21
27.37
7.29
1.71
28.19
8.91
1.97
21.84
10.28
2.27
21.78
10.82
2.37
DualTalk ( Peng et al., 2025 )
68.51
8.12
1.84
61.57
9.26
2.03
64.35
7.29
1.65
55.63
8.03
1.75
42.83
7.65
1.67
38.31
7.64
1.67
Ours (audio)
36.64
9.15
2.05
35.49
9.91
2.17
32.36
8.45
1.89
28.49
8.83
1.93
18.90
9.06
1.97
18.35
9.24
2.00
Ours (FLAME)
36.96
9.16
2.06
35.49
9.90
2.18
33.09
8.34
1.87
29.22
8.73
1.91
18.74
8.73
1.90
18.14
8.77
1.90
Appendix
Table 9: Context signals in mesh space. Speaking/listening results with LVE/MHD in millimeters. Best and second best per split and state use green bold and blue underlining. Lower is better.
Table 12: Extraction settings. Models and parameters for extracting aligned dyadic supervision.
Figure 22: User study interface. Left: instructions for the blind A/B comparison. Right: a trial showing the User video and two anonymous avatar outputs with shared conversation audio and activity cues. Annotators choose Avatar A or Avatar B separately for lip synchronization, motion naturalness, turn-taking coherence, audiovisual responsiveness, and overall conversational realism.
We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking behavior, and high rendered visual quality simultaneously. Our framework couples the first Rectified-Flow Diffusion Transformer (DiT) for this task with a differentiable renderer, enabling diverse, high-fidelity generation in as few as four sampling steps. Prior listening-speaking methods rely on dual-stream audio, introducing an interlocutor look-ahead dependency incompatible with causal user--LLM interaction. We instead adopt a single-stream interface with explicit per-frame listening-speaking state conditioning and a Streaming Audio Scheduler, suppressing spurious mouth motion during listening while enabling seamless turn-taking. A two-stage training scheme of coefficient-space pretraining and joint image-domain refinement further closes the gap between motion-level supervision and rendered quality. Extensive experiments demonstrate state-of-the-art visual quality and motion fidelity in both speaking and listening scenarios.
Yu Zhang, Kaiyuan Shen, Yang Li
School of Computer Science and Technology, East China Normal University, Shanghai, China · Garabido Shanghai Technology Co., Ltd., Shanghai, China
This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with previous talking-avatar methods that mainly rely on audio or reference-image conditioning, we introduce additional SMPL-X-derived 3D guidance, enabling global control over body pose and head motion. Specifically, we render the driving SMPL-X sequence into dense mesh frames and encode them with a lightweight 3D convolutional encoder, whose outputs are injected into the latent tokens to provide overall motion control. Furthermore, we extend Wan2.2-TI2V-5B with additional audio and face cross-attention modules to enable fine-grained expression control and preserve the input identity, respectively. In addition, we implement an audio-to-SMPL-X model to predict an SMPL-X sequence conditioned on the reference image and input audio, allowing MegaAvatar to support audio-driven inference without user-provided SMPL-X frames. Experiments show that MegaAvatar achieves high-quality talking avatar generation with controllable body and head motion, speech-synchronized facial expressions, and consistent identity preservation. MegaAvatar also supports inference with flexible resolutions and video lengths. Codes, dataset, models will be avaliable in https://github.com/Jeoyal/MegaAvatar
Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agents and digital humans. However, collecting such data is time-consuming, expensive, and ethically sensitive. To address this, we propose CHAT, a new dyadic interactive audio-visual dialogue generation (DIADG) framework that generates diverse, paired, and mutually responsive speech-face dialogue clips from a single textual prompt. CHAT unifies large language models and talking face models with interactive audio and facial behaviour refinement modules, enabling the generation of aligned dyadic dialogue clips with diverse contents and facial identities. Experiments show that CHAT outperforms existing related methods designed for similar tasks under both objective and subjective evaluations. Moreover, our synthesised CHAT-AVD-50k dataset serves as effective pre-training data for downstream interactive head generation, consistently improving PerFRDiff and ReactDiff on REACT 2024. CHAT offers a scalable alternative to the costly and ethically sensitive collection of real dyadic interaction data.
Junhao Song, Lluis Guasch, Xilin He +8
Department of Computing, Imperial College London, UK · Department of Earth Science & Engineering, Imperial College London, UK · Mohamed bin Zayed University of Artificial Intelligence, UAE +8