Talk Like You: Imitating How You Speak in Real-Time Talking Head Generation
Organizations: MAIS, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · CAIR, HKISI, Chinese Academy of Sciences · SCSE, FIE, M.U.S.T.
Abstract
In daily life, each person exhibits unique speaking habits, leading to subtle yet consistent lip-shape variations even when pronouncing the same word. Although recent talking head generation methods have achieved impressive visual fidelity and lip synchronization, they largely overlook user-specific customization, especially the motion patterns that characterize individual speaking habits. These habits are difficult to model and capture, as their motion patterns are highly fine-grained and often similar across individuals. As a result, many approaches produce overly uniform facial motions and fail to capture diverse, person-specific articulation patterns. To address this, we propose TalkLikeYou, an efficient framework that imitates how a target person speaks in talking head generation. Our method models habit in motion-space and achieves real-time performance through Flow Matching with only one sampling step during inference. We further adopt a two-stage imitation learning strategy to capture subtle distinctions between habits, allowing users to specify a target habit through either a preset style from the dataset or a reference video. In addition, we introduce a new metric PLAD that projects mouth motions onto representative articulation axes to evaluate imitation accuracy and generation diversity. Extensive experiments demonstrate that TalkLikeYou generates high-quality talking heads in real-time and significantly improves speaking habit imitation compared with prior methods. The code is available at: https://github.com/BQ-Wang0511/TalkLikeYou
Figures & tables
| Method | HDTF | CelebV-HQ | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Sync-C | Sync-D | FID | Sync-C | Sync-D | FID | |||||||
| SadTalk Zhang et al. (2023) | 7.66 | 7.32 | 87.60 | 45.72 | 534.14 | 0.85 | 6.36 | 8.36 | 109.11 | 40.82 | 829.09 | 0.80 |
| Hallo2 Cui et al. (2025) | 7.33 | 7.97 | 25.41 | 13.02 | 272.57 | 0.83 | 6.32 | 8.73 | 63.95 | 13.13 | 688.03 | 0.75 |
| Sonic Ji et al. (2025) | 8.62 | 6.70 | 23.42 | 12.91 | 195.21 | 0.84 | 7.68 | 7.35 | 66.37 | 12.88 | 704.12 | 0.76 |
| Float Ki et al. (2025) | 7.04 | 8.16 | 84.70 | 12.39 | 382.93 | 0.82 | 6.44 | 8.64 | 137.23 | 11.89 | 854.89 | 0.70 |
| Ditto Li et al. (2025) | 5.23 | 9.94 | 21.61 | 12.62 | 319.03 | 0.87 | 5.80 | 9.19 | 62.67 | 12.13 | 743.16 | 0.82 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Methods | Lip-sync | Image | Temporal | Habit | Overall |
|---|---|---|---|---|---|
| Accuracy | Quality | Consistency | Imitation | Realistic | |
| Ditto Li et al. (2025) | 2.2 | 2.4 | 2.7 | —- | 2.5 |
| Hallo2 Cui et al. (2025) | 2.2 | 2.8 | 3.3 | —- | 3.0 |
| Sonic Ji et al. (2025) | 4.0 | 2.4 | 3.5 | —- | 3.6 |
| SadTalker Zhang et al. (2023) | 1.2 | 1.9 | 2.4 | —- | 1.9 |
| StyleTalk Ma et al. (2023a) | 1.3 | 1.9 | 1.6 | 1.5 | 1.7 |