EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold
Authors: Junjie Chen, Fei Wang, Kun Li, Yiqi Nie, Xun Yang, Yanbin Hao, Linfeng Zhang, Meng Wang
Organizations: Hefei University of Technology · EPIC Lab, Shanghai Jiao Tong University · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center · United Arab Emirates University · Anhui University · University of Science and Technology of China · SAI, Shanghai Jiao Tong University
Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAvatar, a causal generator that uses test-time training to adapt to user face video and dyadic audio during interaction. Its dyadic context prediction objective provides a self-supervised learning signal from audiovisual context without target motion labels at test time. Persistent fast weights accumulate these updates within each conversation to guide motion generation, while transient jaw adaptation responds to current audiovisual context. Predicted speech activity controls how persistent adaptation guides motion. We also introduce InterHead-Bench, a unified 455.95-hour benchmark built from single-view and dual-view conversation videos. Experiments show improved conversational motion statistics over strong baselines. On the hardest out-of-distribution split, generation improves as conversations unfold, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
Figures & tables
Figure 2: EvolvingAvatar learns from the ongoing conversation at deployment without motion labels. Dyadic context prediction drives two complementary forms of adaptation: persistent fast weights capture conversational regularities, while transient jaw adaptation tracks current articulation. Regional latent routing combines their predictions through a shared causal decoder.
Figure 3: dialog3d-factory processes dual-view recordings with separate audio and single-view recordings with mixed audio. A shared workflow extracts face video, participant speech, FLAME motion, transcripts, speaking and interaction states, and turn statistics on a common timeline.
Method
Venue
Context
Parameter Size
U-Video
U-FLAME
U-Audio
A-Audio
Total (M)
Trainable (M)
DiffPoseTalk ( Sun et al., 2024 )
TOG ’24
✗
✗
✗
✓
129.32
110.55
ARTalk ( Chu et al., 2025 )
SIGGRAPH Asia ’25
✗
✗
✗
✓
382.87
37.89
UniLS ( Chu et al., 2026 )
CVPR ’26
✗
✗
✓
✓
422.88
27.38
DualTalk ( Peng et al., 2025 )
CVPR ’25
✗
✓
✓
✓
647.27
638.85
EvolvingAvatar (Ours)
This work
✓
✗
✓
✓
159.57
54.82
Table 1: Generation-time conditioning context and parameter sizes of the compared methods. U and A denote User and Avatar, respectively. Parameter counts are reported in millions (M).
Figure 4: InterHead-Bench composition by split, including scale and the distributions of interaction duration and conversational turns.
Method
FD ↓
P-FD ↓
MSE ↓
rPCC ↓
SID ↑
PDD ↓
JDD ↓
EXP
JAW (×103)
POSE (×102)
EXP
JAW (×103)
POSE (×102)
EXP (×101)
JAW (×103)
POSE (×102)
EXP (×102)
JAW (×101)
POSE (×101)
EXP
JAW
POSE
In-Distribution Test Set (ID)
Speaking
DiffPoseTalk ( Sun et al., 2024 )
48.45
22.62
24.60
50.30
22.93
25.47
6.01
10.01
11.61
20.11
1.95
3.31
1.84
1.36
1.35
9.38
2.88
ARTalk ( Chu et al., 2025 )
21.11
3.66
18.89
22.12
3.75
19.52
2.77
1.70
8.68
18.83
1.30
3.29
2.18
1.69
1.38
5.67
0.95
UniLS ( Chu et al., 2026 )
15.99
2.36
12.31
17.05
2.45
12.76
2.32
1.30
5.36
14.73
1.07
3.37
2.84
2.20
1.13
7.36
0.84
Table 2: Parameter-space comparison. Speaking and listening results. Green bold and blue underlined values mark the best and second best in each split and state. Lower is better except SID.
ID
OOD
OOD-Hard
Method
Speaking
Listening
Speaking
Listening
Speaking
Listening
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
DiffPoseTalk ( Sun et al., 2024 )
39.26
18.78
4.10
40.74
19.97
4.31
38.37
18.10
4.02
40.82
18.60
4.11
37.57
17.94
3.97
44.10
18.28
4.01
ARTalk ( Chu et al., 2025 )
32.97
9.33
2.21
35.90
11.77
2.75
30.82
8.18
1.98
30.92
9.85
2.32
20.91
10.79
2.49
19.84
10.10
2.34
UniLS ( Chu et al., 2026 )
29.21
7.62
1.79
31.11
9.92
2.21
27.37
7.29
1.71
28.19
8.91
1.97
21.84
10.28
2.27
21.78
10.82
2.37
DualTalk ( Peng et al., 2025 )
68.51
8.12
1.84
61.57
9.26
2.03
64.35
7.29
1.65
55.63
8.03
1.75
42.83
7.65
1.67
38.31
7.64
1.67
Table 3: Mesh-space comparison. Speaking/listening results. LVE and MHD are in millimeters. Green bold and blue underline denote the best and second best per split and state. Lower is better.
Table 7Figure 8
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Restrained neck and mouth motion across activities. EvolvingAvatar preserves the reference’s slightly raised chin and small speaking openings, with a more restrained mouth shape in the intervening listening columns. UniLS opens the mouth too widely at several of these selected instants, while DiffPoseTalk and ARTalk introduce marked downward or sideways neck poses.
Figure 8: Smiling while listening. EvolvingAvatar preserves the reference’s smile and moderate mouth opening in columns 4–6, across speaking and listening. UniLS and DualTalk show weaker openings there, while DiffPoseTalk exaggerates mouth opening in later sampled listening frames.
Figure 9: From listening to speaking. EvolvingAvatar combines the tilted listening posture in columns 1–5 with the larger speaking mouth openings in columns 7–10. UniLS changes the mouth opening less consistently with Avatar GT, while DualTalk largely retains a neutral expression.
Figure 10: Coordinating changes in expression and neck pose. EvolvingAvatar follows the reference’s slight downward tilt during listening in columns 5–8 and sideways lean in the final speaking frames. UniLS exaggerates the downward tilt during listening, while DualTalk remains largely frontal and DiffPoseTalk adds a broad grin where the reference expression is restrained.
Figure 11: Moderate mouth openings and neck motion. EvolvingAvatar captures the speaking openings and tilt in columns 2–4, then retains a listening smile. UniLS exaggerates those speaking openings and closes the eyes in columns 6–9, while DiffPoseTalk adds a larger sideways lean.
Figure 12: Matching neck orientation and articulation together. EvolvingAvatar retains the reference’s near-upright orientation and modest changes in mouth opening across the selected speaking and listening instants. UniLS exaggerates the upward tilt or mouth opening at several instants, while ARTalk introduces a pronounced downward tilt that is absent from the reference.
Figure 13: An expressive listener. EvolvingAvatar follows the reference’s open-mouth smile and upward, sideways orientation during both activities. UniLS largely closes the mouth in the first five listening frames, while ARTalk tilts the face downward and DualTalk retains a more frontal pose.
Figure 14: From varied articulation to a listening smile. EvolvingAvatar captures the speaking openings in columns 1–6 and closed-mouth listening smile in columns 7–10. UniLS exaggerates upward tilt and early openings, while DiffPoseTalk bows and turns the face during listening.
Figure 15: Changing articulation within each activity. EvolvingAvatar follows the reference’s alternation between smaller and larger mouth openings while retaining a moderate neck pose across the sampled instants. UniLS raises the chin too far and opens the mouth in several near-closed reference frames, while DiffPoseTalk over-opens the mouth across most of the displayed columns.
Figure 16: Subtle articulation with an upright neck pose. EvolvingAvatar preserves the near-upright neck pose and the mild increase in mouth opening in the later columns. UniLS and ARTalk lower the face, while DiffPoseTalk replaces the small final openings with larger ones.
Figure 17: A smile across speaking and listening. EvolvingAvatar combines a near-frontal pose with larger speaking openings and a softer listening smile, retaining visible expression across both activities. UniLS and DualTalk understate the reference speaking expression in the first five columns, while DiffPoseTalk produces excessive mouth opening in the subsequent listening frames.
Figure 18: Expressive listening with an upright pose. EvolvingAvatar captures the listening smile in columns 2–4 and stays near the reference orientation during speaking. UniLS and ARTalk lower the face during speaking, while DualTalk shows less of the reference variation in mouth shape.
Figure 19: A frontal smile with varying mouth opening. EvolvingAvatar preserves the reference’s upright pose and smile, including the larger openings in columns 8–9. UniLS and ARTalk lower the face, while DiffPoseTalk adds a pronounced sideways turn and wider grin in the final columns.
Figure 20: Expression changes with restrained neck motion. EvolvingAvatar follows the reference’s smaller openings in columns 6–7 and broader speaking smile in columns 8–10. UniLS and ARTalk bow the head, while DualTalk shows less change in mouth shape than the reference.
Figure 21: Reducing articulation during listening. EvolvingAvatar remains nearly upright while following the smaller listening openings in columns 7–10. UniLS and ARTalk lower the face, while DualTalk retains a parted mouth and DiffPoseTalk produces a wider final opening than Avatar GT.
Method
FD ↓
P-FD ↓
MSE ↓
rPCC ↓
SID ↑
PDD ↓
JDD ↓
EXP
JAW (×103)
POSE (×102)
EXP
JAW (×103)
POSE (×102)
EXP (×101)
JAW (×103)
POSE (×102)
EXP (×102)
JAW (×101)
POSE (×101)
EXP
JAW
POSE
In-Distribution Test Set (ID)
DiffPoseTalk ( Sun et al., 2024 )
45.68
25.60
23.10
47.40
25.90
23.99
5.83
11.35
11.38
18.94
1.90
2.90
1.85
1.45
1.48
9.87
3.04
ARTalk ( Chu et al., 2025 )
22.99
3.94
21.71
23.91
4.03
22.32
2.93
1.86
9.60
21.72
1.57
3.18
1.69
1.55
1.17
5.93
1.26
UniLS ( Chu et al., 2026 )
17.19
3.05
12.95
18.20
3.18
13.37
2.45
1.79
5.60
15.86
1.48
3.16
2.31
1.90
1.04
7.52
1.03
DualTalk ( Peng et al., 2025 )
17.22
4.63
11.52
17.35
4.68
11.60
1.79
1.81
4.05
23.17
1.81
3.66
0.32
0.57
0.28
11.34
2.09
Appendix
Table 6: Parameter-space comparison. Results over all valid frames. Green bold and blue underline indicate the best and second-best values within each split. Lower is better except for SID.
Method
ID
OOD
OOD-Hard
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
DiffPoseTalk ( Sun et al., 2024 )
40.03
19.77
4.28
40.37
18.52
4.09
44.83
18.31
4.02
ARTalk ( Chu et al., 2025 )
35.22
11.17
2.62
30.51
9.35
2.22
19.51
10.09
2.34
UniLS ( Chu et al., 2026 )
30.39
9.33
2.10
27.64
8.33
1.87
21.65
10.65
2.34
DualTalk ( Peng et al., 2025 )
66.70
8.87
1.96
61.46
7.71
1.71
40.10
7.60
1.66
EvolvingAvatar (Ours)
37.91
9.83
2.18
31.53
8.74
1.92
17.87
8.71
1.89
Appendix
Table 7: Mesh-space comparison. All-frame results. LVE and MHD are measured in millimeters. Green bold and blue underlined values mark the best and second best in each split. Lower is better.
Method
FD ↓
P-FD ↓
MSE ↓
rPCC ↓
SID ↑
PDD ↓
JDD ↓
EXP
JAW (×103)
NECK (×102)
EXP
JAW (×103)
NECK (×102)
EXP (×101)
JAW (×103)
NECK (×102)
EXP (×102)
JAW (×101)
NECK (×101)
EXP
JAW
NECK
In-Distribution Test Set (ID)
Speaking
UniLS ( Chu et al., 2026 )
15.99
2.36
12.31
17.05
2.45
12.76
2.32
1.30
5.36
14.73
1.07
3.37
2.84
2.20
1.13
7.36
0.84
DualTalk ( Peng et al., 2025 )
16.80
4.17
12.05
16.97
4.23
12.16
1.74
1.59
4.21
20.69
1.31
4.13
0.31
0.49
0.25
11.37
2.08
Ours (audio)
13.71
2.21
9.34
14.90
2.35
9.83
2.15
1.69
4.70
15.38
1.22
3.29
2.47
2.13
1.42
6.86
0.95
Appendix
Table 8: Context signals in parameter space. Speaking/listening results. Green bold and blue underlining denote the best and second best per split and state. Lower is better except SID.
ID
OOD
OOD-Hard
Method
Speaking
Listening
Speaking
Listening
Speaking
Listening
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
FDD ↓
LVE ↓
MHD ↓
UniLS ( Chu et al., 2026 )
29.21
7.62
1.79
31.11
9.92
2.21
27.37
7.29
1.71
28.19
8.91
1.97
21.84
10.28
2.27
21.78
10.82
2.37
DualTalk ( Peng et al., 2025 )
68.51
8.12
1.84
61.57
9.26
2.03
64.35
7.29
1.65
55.63
8.03
1.75
42.83
7.65
1.67
38.31
7.64
1.67
Ours (audio)
36.64
9.15
2.05
35.49
9.91
2.17
32.36
8.45
1.89
28.49
8.83
1.93
18.90
9.06
1.97
18.35
9.24
2.00
Ours (FLAME)
36.96
9.16
2.06
35.49
9.90
2.18
33.09
8.34
1.87
29.22
8.73
1.91
18.74
8.73
1.90
18.14
8.77
1.90
Appendix
Table 9: Context signals in mesh space. Speaking/listening results with LVE/MHD in millimeters. Best and second best per split and state use green bold and blue underlining. Lower is better.
Table 12: Extraction settings. Models and parameters for extracting aligned dyadic supervision.
Figure 22: User study interface. Left: instructions for the blind A/B comparison. Right: a trial showing the User video and two anonymous avatar outputs with shared conversation audio and activity cues. Annotators choose Avatar A or Avatar B separately for lip synchronization, motion naturalness, turn-taking coherence, audiovisual responsiveness, and overall conversational realism.
Department of Computing, Imperial College London, UK · Department of Earth Science & Engineering, Imperial College London, UK · Mohamed bin Zayed University of Artificial Intelligence, UAE +8