Organizations: MAIS, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · CAIR, HKISI, Chinese Academy of Sciences · SCSE, FIE, M.U.S.T.
In daily life, each person exhibits unique speaking habits, leading to subtle yet consistent lip-shape variations even when pronouncing the same word. Although recent talking head generation methods have achieved impressive visual fidelity and lip synchronization, they largely overlook user-specific customization, especially the motion patterns that characterize individual speaking habits. These habits are difficult to model and capture, as their motion patterns are highly fine-grained and often similar across individuals. As a result, many approaches produce overly uniform facial motions and fail to capture diverse, person-specific articulation patterns. To address this, we propose TalkLikeYou, an efficient framework that imitates how a target person speaks in talking head generation. Our method models habit in motion-space and achieves real-time performance through Flow Matching with only one sampling step during inference. We further adopt a two-stage imitation learning strategy to capture subtle distinctions between habits, allowing users to specify a target habit through either a preset style from the dataset or a reference video. In addition, we introduce a new metric PLAD that projects mouth motions onto representative articulation axes to evaluate imitation accuracy and generation diversity. Extensive experiments demonstrate that TalkLikeYou generates high-quality talking heads in real-time and significantly improves speaking habit imitation compared with prior methods. The code is available at: https://github.com/BQ-Wang0511/TalkLikeYou
Figures & tables
Figure 1: Overview of TalkLikeYou. Previous methods often generate uniform motions with limited control over speaking habits. We propose TalkLikeYou for efficient habit imitation in talking head generation, together with a new metric PLAD to evaluate habit similarity and diversity.
Figure 2: Framework of TalkLikeYou. Our method first generates habit-aware motion using the Flow Matching Motion Generator with optional habit sources through one-step sampling, and then renders them into a video. We further adopt a two-stage habit imitation learning strategy for more generalized and stable habit imitation. In the unified habit space, ⋆ denotes the one-hot habit code and ∙ denotes reference-video samples. During the first training stage, we use an alignment loss between reference video features and their corresponding one-hot habit codes.
Figure 3: Comparison Results. We compare our method with advanced baselines under three selected speaking habits, showing better visual quality and lip synchronization. Please zoom in for details. Video results are provided in the supplementary material.
Method
HDTF
CelebV-HQ
Sync-C ↑
Sync-D ↓
FID ↓
NIQE↓
FVD↓
CSIM↑
Sync-C ↑
Sync-D ↓
FID ↓
NIQE↓
FVD↓
CSIM↑
SadTalk Zhang et al. (2023)
7.66
7.32
87.60
45.72
534.14
0.85
6.36
8.36
109.11
40.82
829.09
0.80
Hallo2 Cui et al. (2025)
7.33
7.97
25.41
13.02
272.57
0.83
6.32
8.73
63.95
13.13
688.03
0.75
Sonic Ji et al. (2025)
8.62
6.70
23.42
12.91
195.21
0.84
7.68
7.35
66.37
12.88
704.12
0.76
Float Ki et al. (2025)
7.04
8.16
84.70
12.39
382.93
0.82
6.44
8.64
137.23
11.89
854.89
0.70
Ditto Li et al. (2025)
5.23
9.94
21.61
12.62
319.03
0.87
5.80
9.19
62.67
12.13
743.16
0.82
Table 1: Quantitative comparisons with state-of-the-art methods.
Figure 5Figure 6
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: (a) Comparison of generation tendency under DDPM and Flow Matching. (b) PLAD histograms of different methods. (c) t-SNE visualization of the unified habit space, where stars denote preset one-hot habit features and dots denote reference-based habit features.
Methods
Lip-sync
Image
Temporal
Habit
Overall
Accuracy
Quality
Consistency
Imitation
Realistic
Ditto Li et al. (2025)
2.2
2.4
2.7
—-
2.5
Hallo2 Cui et al. (2025)
2.2
2.8
3.3
—-
3.0
Sonic Ji et al. (2025)
4.0
2.4
3.5
—-
3.6
SadTalker Zhang et al. (2023)
1.2
1.9
2.4
—-
1.9
StyleTalk Ma et al. (2023a)
1.3
1.9
1.6
1.5
1.7
Appendix
Table 3: User Study. The rating is on scale of 1-5 where higher rank demonstrate better results.
Figure 8: More results on habit imitation.
Figure 9: More comparison on Lip Synchronization with other baselines.
Figure 10: More visualization results on stylize faces.
Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer from error accumulation and identity drift. To address this drawback, we propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos. At the heart of our approach lies a single-step bridge distillation scheme. On the one hand, departing from the conventional noise-to-data paradigm, we introduce a data-to-data transport formulation based on a Brownian bridge. Anchored by a persistent reference, this strategy effectively mitigates identity drift and enhances long-term temporal stability. On the other hand, to enable smooth knowledge transfer from a pre-trained diffusion teacher to the student bridge model, we explore a heterogeneous distillation framework with an SNR-aligned time transformation Φ(τ), which bridges the functional discrepancy between the two models. Moreover, we propose an audio-driven classifier-free guidance mechanism to maintain fine-grained lip synchronization under extreme step reduction. Extensive experiments demonstrate that our method achieves high-fidelity and temporally consistent video generation with only 1 step at up to 200 FPS, significantly outperforming existing approaches in both efficiency and stability. Project Page: https://zhangrongxiang.github.io/leaptalk-page/
Rongxiang Zhang, Songhua Liu
School of Artificial Intelligence, Shanghai Jiao Tong University · 2Harbin Institute of Technology
High-quality, real-time talking head synthesis remains a fundamental challenge in computer vision. Existing reconstruction- and rendering-based methods typically rely on identity-specific models, limiting cross-identity generalization. To address this issue, we propose SDTalk, a one-shot 3D Gaussian Splatting (3DGS)-based framework that generalizes to unseen identities without personalized training or fine-tuning. Our framework comprises two modules with a two-stage training strategy. In the first stage, we incorporate structured facial priors into the reconstruction module and separately predict 3DGS parameters for visible and occluded regions, enabling complete head reconstruction from a single image. In the second stage, we introduce a dual-branch motion field to model coarse and fine facial dynamics, improving detail fidelity and lip synchronization. Experiments demonstrate that SDTalk surpasses existing methods in both visual quality and inference efficiency.
Peng Jia, Zhen Xiao, Jia Li +3
Hefei University of Technology, Hefei, China · University of Science and Technology of China, Hefei, China
We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking behavior, and high rendered visual quality simultaneously. Our framework couples the first Rectified-Flow Diffusion Transformer (DiT) for this task with a differentiable renderer, enabling diverse, high-fidelity generation in as few as four sampling steps. Prior listening-speaking methods rely on dual-stream audio, introducing an interlocutor look-ahead dependency incompatible with causal user--LLM interaction. We instead adopt a single-stream interface with explicit per-frame listening-speaking state conditioning and a Streaming Audio Scheduler, suppressing spurious mouth motion during listening while enabling seamless turn-taking. A two-stage training scheme of coefficient-space pretraining and joint image-domain refinement further closes the gap between motion-level supervision and rendered quality. Extensive experiments demonstrate state-of-the-art visual quality and motion fidelity in both speaking and listening scenarios.
Yu Zhang, Kaiyuan Shen, Yang Li
School of Computer Science and Technology, East China Normal University, Shanghai, China · Garabido Shanghai Technology Co., Ltd., Shanghai, China