cs.CVOct 5, 2026

Talk Like You: Imitating How You Speak in Real-Time Talking Head Generation

Authors: Baiqin Wang, Zhixing Ding, Jijie Li, Jiankuo Zhao, Zhen Lei, Xiangyu Zhu

Organizations: MAIS, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · CAIR, HKISI, Chinese Academy of Sciences · SCSE, FIE, M.U.S.T.

Abstract

In daily life, each person exhibits unique speaking habits, leading to subtle yet consistent lip-shape variations even when pronouncing the same word. Although recent talking head generation methods have achieved impressive visual fidelity and lip synchronization, they largely overlook user-specific customization, especially the motion patterns that characterize individual speaking habits. These habits are difficult to model and capture, as their motion patterns are highly fine-grained and often similar across individuals. As a result, many approaches produce overly uniform facial motions and fail to capture diverse, person-specific articulation patterns. To address this, we propose TalkLikeYou, an efficient framework that imitates how a target person speaks in talking head generation. Our method models habit in motion-space and achieves real-time performance through Flow Matching with only one sampling step during inference. We further adopt a two-stage imitation learning strategy to capture subtle distinctions between habits, allowing users to specify a target habit through either a preset style from the dataset or a reference video. In addition, we introduce a new metric PLAD that projects mouth motions onto representative articulation axes to evaluate imitation accuracy and generation diversity. Extensive experiments demonstrate that TalkLikeYou generates high-quality talking heads in real-time and significantly improves speaking habit imitation compared with prior methods. The code is available at: https://github.com/BQ-Wang0511/TalkLikeYou

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

    Jul 29, 2026Rongxiang Zhang, Songhua LiuPrecise Lip SynchronizationTalk

  2. SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis

    May 11, 2026Peng Jia, Zhen Xiao, Jia Li +3Precise Lip SynchronizationTalk

  3. EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents

    Apr 19, 2026Yu Zhang, Kaiyuan Shen, Yang LiAvatarsTurn-Taking