cs.CVOct 5, 2026

UltraDub: Towards Authentic Dubbing by Unifying Visually-Steered Flow Learning and Trajectory Guidance

Authors: Gaoxiang Cong, Liang Li, Jianwei Wen, Zhedong Zhang, Zheng-Jun Zha, Qingming Huang

Organizations: Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Hangzhou Dianzi University · ByteDance · University of Science and Technology of China

Abstract

Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can improve linguistic accuracy at the expense of lip synchronization. In this paper, we propose UltraDub, a Unifying Visually-Steered Flow learning and trajectory Guidance Dubbing framework that leverages vision in two ways: as continuous motion for multimodal context aggregation, and as structural rhythm for trajectory rectification. Specifically, we introduce the Motion-guided Dual-context Retrieving (MDR) module, which continually recalibrates linguistic and speaker-style retrieval through shared lip-motion query residuals, utilizing independent time-conditioned gates to regulate their contributions. Furthermore, we propose Rhythm-anchored Trajectory Guidance (RTG), a training-free mechanism that evaluates hierarchical multimodal corrections at a visual-only predictive midpoint, safely strengthening semantic conditioning while better preserving temporal alignment. Finally, we construct DiverseDub, a multi-scenario benchmark to evaluate video dubbing in the wild. Extensive experiments demonstrate that UltraDub achieves state-of-the-art performance across four datasets.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

    Nov 23, 2025Kaidi Wang, Yi He, Wenhao Guan +9Audio-Visual ConsistencySpeech Synthesis

  2. Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

    Sep 30, 2026Rui Liu, Bhavin Jawade, Haoqi Li +4Precise Lip SynchronizationSpeech-To-Text Alignment

  3. Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing

    Sep 22, 2026Alejandro Pérez-González-de-Martos, Florian Lux, Angelina Elizarova +3Precise Lip SynchronizationSpeech Synthesis