cs.CVJul 30, 2026

TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment

Authors: MD Wahiduzzaman KhanMingshan JiaXiaolin ZhangEn YuKaska Musial-Gabrys

Organizations: University of Technology Sydney, Australia · Shandong University of Science and Technology, China

Abstract

Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors during speech and expressive motions. We introduce the first framework for cross-identity tongue dynamics transfer in face reenactment. We propose a foundation-model-assisted bootstrapping pipeline that produces a dedicated tongue segmentation model for in-the-wild reenactment without curated annotations. We further introduce a spatially constrained latent masked diffusion model for realistic tongue synthesis, with adaptive mask dilation for seamless mouth boundary transitions. Extensive experiments demonstrate improvements of more than two times over all baselines on every tongue-specific metric. We additionally propose a VLM-based evaluation protocol that replicates expert annotation at scale, confirming perceptual superiority across all ablation variants.

Explore similar work

CardsList
  1. EmbedTalk: Talking Head Synthesis using Gaussian Embeddings

    Mar 8, 2026Arpita Saggar, Jonathan C. Darling, Duygu Sarikaya +1Facial AnimationTalk