Organizations: Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Hangzhou Dianzi University · ByteDance · University of Science and Technology of China
Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can improve linguistic accuracy at the expense of lip synchronization. In this paper, we propose UltraDub, a Unifying Visually-Steered Flow learning and trajectory Guidance Dubbing framework that leverages vision in two ways: as continuous motion for multimodal context aggregation, and as structural rhythm for trajectory rectification. Specifically, we introduce the Motion-guided Dual-context Retrieving (MDR) module, which continually recalibrates linguistic and speaker-style retrieval through shared lip-motion query residuals, utilizing independent time-conditioned gates to regulate their contributions. Furthermore, we propose Rhythm-anchored Trajectory Guidance (RTG), a training-free mechanism that evaluates hierarchical multimodal corrections at a visual-only predictive midpoint, safely strengthening semantic conditioning while better preserving temporal alignment. Finally, we construct DiverseDub, a multi-scenario benchmark to evaluate video dubbing in the wild. Extensive experiments demonstrate that UltraDub achieves state-of-the-art performance across four datasets.
Figures & tables
Figure 1: (a) Illustration of the V2C task. (b) Conventional CFG relies on linear extrapolation from fully conditioned velocity ( vc ) via scaling factors β∗ , making it prone to off-manifold trajectory drift. (c) For the multi-conditional V2C task, existing guidance easily triggers imbalance and degradation.
Figure 2: Overview of UltraDub, a U nifying Visually Steered Flow l earning and tra jectory Guidance Dub bing framework. The visually steered flow learning phase consists of an input projection, the core Motion-guided Dual-context Retrieving (MDR) layers, and a final modulation to predict the target velocity field vθ . With the velocity field established, Rhythm-anchored Trajectory Guidance (RTG) acts as a training-free inference mechanism. It constructs a visual-only midpoint as a rhythm anchor to evaluate multimodal corrections, continuously rectifying the base transport velocity.
Figure 3: Statistical overview of DiverseDub . (a) Left: distribution of clips over the eight dubbing categories, with a representative frame for each category. (b) Middle top: per-domain distributions of clip-level DNSMOS quality. (c) Middle bottom: per-domain distributions of SyncNet audiovisual offsets. (d) Right: relative frequency of word types among the nouns and adjectives of each domain, normalized by the corpus average.
Method
WER (%) ↓
SPKSIM (%) ↑
EmoSIM (%) ↑
LSE-C ↑
LSE-D ↓
GT
1.07
69.92
100.00
6.64
7.83
ProDubber
13.86
25.17
69.29
2.40
11.87
VoiceCraftDub
84.68
41.26
72.90
4.89
9.60
AlignDiT
20.61
49.78
80.09
6.33
8.15
InstructDubber
8.38
23.87
67.72
1.85
12.47
CoSyncDiT
12.01
56.32
81.72
6.88
7.86
Table 1: Results on the DiverseDub benchmark. Arrows indicate the preferred direction; the best result among dubbing methods is shown in bold.
Method
WER (%) ↓
SPKSIM (%) ↑
EmoSIM (%) ↑
LSE-C ↑
LSE-D ↓
GT
4.15
66.13
100.00
6.88
7.39
ProDubber
9.71
22.32
65.97
2.44
11.94
VoiceCraftDub
41.19
34.65
72.89
5.46
8.83
AlignDiT
25.86
46.25
74.89
6.29
8.01
InstructDubber
6.47
24.82
72.43
2.43
11.96
CoSyncDiT
12.52
49.05
78.93
6.86
7.56
Table 2: Results on the CelebV-Dub benchmark. Arrows indicate the preferred direction; the best result among dubbing methods is shown in bold.
Table 3: Comparison with dubbing methods on the LRS3 benchmark. Best results among dubbing methods are shown in bold.
Figure 5: Preserving audiovisual alignment under increasing guidance strength on the LRS3 test set. AVSync is a recently introduced metric that provides a more reliable assessment of audiovisual synchronization Choi et al. (2025) .
Methods
WER (%) ↓
SPKSIM (%) ↑
LSE-C ↑
LSE-D ↓
GT
13.14
63.69
7.14
6.84
ProDubber
44.29
34.91
3.71
11.08
VoiceCraftDub
64.47
33.37
5.13
8.50
AlignDiT
19.47
55.06
7.17
6.84
InstructDubber
27.67
36.30
3.33
11.32
CoSyncDiT
19.64
53.45
7.21
6.86
Table 4: Comparison with dubbing methods on the GRID benchmark. EmoSIM is not evaluated because GRID consists primarily of neutral emotion.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Mel-spectrogram comparisons between ground-truth speech and speech synthesized by different models on DiverseDub, LRS3, and GRID. The red and white rectangles highlight audiovisual alignment intervals and local phonetic details, respectively.
Figure 7: More category-wise comparisons on DiverseDub. Error bars indicate 95% confidence intervals. Further details and qualitative examples are provided in the supplementary demo.
Figure 8: Preserving audiovisual alignment under increasing guidance strength on the CelebV-Dub test set. AVSync is a recently introduced metric that provides a more reliable assessment of audiovisual synchronization Choi et al. (2025) .
Methods
WER (%) ↓
SPKSIM (%) ↑
LSE-C ↑
LSE-D ↓
GT
2.33
56.18
7.80
6.85
SA+FFN (base)
1.52
47.12
3.17
11.25
+MDR (Single Linguistic Memory)
1.43
46.53
8.17
6.59
+MDR (Single Speaker-Style Memory)
1.58
51.92
8.13
6.64
+MDR (Dual-context Memories)
1.39
51.49
8.18
6.56
RTG + MDR (Dual-context Memories)
1.37
52.05
8.23
6.50
Appendix
Table 5: Component ablation study of UltraDub on LRS3-Cross. Starting from a self-attention and feed-forward network (SA+FFN) backbone, we evaluate different MDR memory configurations and the contribution of RTG. Arrows indicate the preferred metric directions, and bold denotes the best result. Each architectural variant is trained independently under the same training configuration, while RTG is applied only at inference time.
Method
WER (%) ↓
SPKSIM (%) ↑
EmoSIM (%) ↑
LSE-C ↑
LSE-D ↓
AVSync ↑
GT
1.07
69.92
100.00
6.64
7.83
1.000
ProDubber
13.86
25.17
69.29
2.40
11.87
0.175
VoiceCraftDub
84.68
41.26
72.90
4.89
9.60
0.248
AlignDiT
20.61
49.78
80.09
6.33
8.15
0.344
InstructDubber
8.38
23.87
67.72
1.85
12.47
0.129
CoSyncDiT
12.01
56.32
81.72
6.88
7.86
0.365
Appendix
Table 6: Results on the DiverseDub benchmark. Arrows indicate the preferred direction; the best result among dubbing methods is shown in bold. AVSync is the mean cosine similarity between AV-HuBERT features of generated and GT audiovisual pairs; higher is better.
Methods
WER (%) ↓
SPKSIM (%) ↑
LSE-C ↑
LSE-D ↓
AVSync ↑
GT
2.33
56.18
7.80
6.85
1.000
ProDubber
3.41
24.27
4.27
10.12
0.220
VoiceCraftDub
14.85
35.17
6.49
8.08
0.568
AlignDiT
1.40
51.46
7.91
6.78
0.751
InstructDubber
2.39
24.16
3.29
11.12
0.172
CoSyncDiT
2.09
47.53
7.83
6.88
0.731
Appendix
Table 7: Comparison with dubbing methods on the LRS3-Cross benchmark. Best results among dubbing methods are shown in bold. Note that EmoSIM is not evaluated here, as LRS3 predominantly features neutral TED talks with limited emotional variation. AVSync is the mean cosine similarity between AV-HuBERT features of generated and GT audiovisual pairs; higher is better.
Figure 9: Instructions and an example item for the human listening test of overall speech naturalness.
Method
Naturalness (MOS-N) ↑
Similarity (MOS-S) ↑
Ground Truth
4.15 ± 0.12
4.01 ± 0.13
ProDubber
3.10 ± 0.17
2.55 ± 0.13
VoiceCraftDub
3.31 ± 0.09
3.15 ± 0.15
InstructDubber
3.36 ± 0.12
2.60 ± 0.14
AlignDiT
3.82 ± 0.14
3.99 ± 0.13
CoSyncDiT
3.86 ± 0.12
3.98 ± 0.15
Appendix
Table 8: Subjective evaluation results on LRS3-Cross for naturalness (MOS-N) and speaker similarity (MOS-S).
Automatic video dubbing aims to generate high-fidelity speech that is temporally aligned with visual content. However, existing methods still suffer from limited speech naturalness, insufficient audio-visual synchronization, and poor scalability beyond monolingual settings. To address these challenges, we propose SyncVoice, a simple and effective dubbing framework that lightly integrates a Text-Visual Fusion Module into a pretrained text-to-speech (TTS) system. This module aligns visual features with linguistic representations, enabling temporally synchronized speech synthesis without complex architectural redesign. Experiments on the LRS3 dataset show that SyncVoice achieves state-of-the-art performance in zero-shot dubbing. Further training on a large-scale bilingual audio-visual dataset improves vocal fidelity while preserving synchronization, yielding a single unified model for both Chinese and English dubbing.
Kaidi Wang, Yi He, Wenhao Guan +9
School of Informatics, Xiamen University, China · MiLM Plus, Xiaomi Inc., China · School of Electronic Science and Engineering, Xiamen University, China +1
Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly suited to this setting: even when fine-tuned to recover spoken content from lip motion, they remain largely insensitive to temporal errors. We introduce Align Then Reason (ATR), a multilingual lip-sync judge that first establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, then reasons over this alignment to make the final judgment. An alignment scorer provides the LLM with both local evidence for each phonetic unit and a calibrated global alignment score, enabling it to reason jointly about content and timing. On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively. The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively. They also transfer across datasets to three unseen MuAViC languages. Furthermore, we evaluate on two downstream tasks built from real dubbing lines. On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.
Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works address this problem by conditioning the speech synthesis process on lip movements extracted from the video signal. In this work, we condition the speech generation on a binary voice-activity signal, which has a lightweight representation and can be produced in multiple ways. We show that the model follows the voice-activity signal with high accuracy while maintaining natural prosody and semantically appropriate pause placement within sentences, as demonstrated through extensive objective and subjective evaluations. By randomly masking this condition during training, we make the feature entirely optional during inference, allowing editors to enforce or relax lip-sync constraints when desired.