eess.ASNov 23, 2025

SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

Authors: Kaidi WangYi HeWenhao GuanWeijie WuPeijie ChenHongwu DingXiong ZhangDi Wu+4 more

Organizations: School of Informatics, Xiamen University, China · MiLM Plus, Xiaomi Inc., China · School of Electronic Science and Engineering, Xiamen University, China · WeNet Open Source Community

Abstract

Automatic video dubbing aims to generate high-fidelity speech that is temporally aligned with visual content. However, existing methods still suffer from limited speech naturalness, insufficient audio-visual synchronization, and poor scalability beyond monolingual settings. To address these challenges, we propose SyncVoice, a simple and effective dubbing framework that lightly integrates a Text-Visual Fusion Module into a pretrained text-to-speech (TTS) system. This module aligns visual features with linguistic representations, enabling temporally synchronized speech synthesis without complex architectural redesign. Experiments on the LRS3 dataset show that SyncVoice achieves state-of-the-art performance in zero-shot dubbing. Further training on a large-scale bilingual audio-visual dataset improves vocal fidelity while preserving synchronization, yielding a single unified model for both Chinese and English dubbing.

Explore similar work

CardsList
  1. Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing

    Sep 22, 2026Alejandro Pérez-González-de-Martos, Florian Lux, Angelina Elizarova +3