cs.CVSep 29, 2026

Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation

Authors: Bangxun Tang

Organizations: University of California, Irvine

Abstract

We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rather than a generic one. Existing lip-sync systems follow the audio closely and keep the face recognizable, yet the mouth they render is an average mouth: the shape and texture of the lips, the arrangement of the teeth, and how much of them shows as the mouth opens are not that person's. The problem persists because nothing in current training or evaluation asks for the person's own mouth: perceptual losses accept any plausible mouth, face identity is carried mostly by the skin around it, and the released inference code of inpainting systems uses the unmasked target frame as the reference, which hides the gap. To address this, RGOR conditions every generated frame on frames from separate enrollment recordings of the same person and on HD patches of the mouth that bypass the VAE, and trains the generator against a paired judge that compares each rendered mouth with the person's reference and learns to reject a realistic mouth of someone else. We further build an evaluation protocol and use it to compare open-source and commercial lip-sync systems on held-out identities. Experiments show that RGOR achieves the best or second-best result on most metrics, and preserves the person's own lip and dental detail while keeping synchronization and the rest of the face intact.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HighSync: High-Quality Lip Synchronization via Latent Diffusion Models

    May 16, 2026Saeid Firouzi Daghigh, Majid Iranpour Mobarakeh, Mostafa Alavi +1Precise Lip SynchronizationSpeech-To-Text Alignment

  2. SyncEdit: Rethinking Lip Synchronization as Editing with Audio-Driven Diffusion Models

    Mar 10, 2026Lixiang Lin, Siyuan Jin, Jinshan ZhangPrecise Lip SynchronizationAudio-Visual Consistency

  3. Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

    Sep 30, 2026Rui Liu, Bhavin Jawade, Haoqi Li +4Precise Lip SynchronizationSpeech-To-Text Alignment