Organizations: Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Hangzhou Dianzi University · ByteDance · University of Science and Technology of China
Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can improve linguistic accuracy at the expense of lip synchronization. In this paper, we propose UltraDub, a Unifying Visually-Steered Flow learning and trajectory Guidance Dubbing framework that leverages vision in two ways: as continuous motion for multimodal context aggregation, and as structural rhythm for trajectory rectification. Specifically, we introduce the Motion-guided Dual-context Retrieving (MDR) module, which continually recalibrates linguistic and speaker-style retrieval through shared lip-motion query residuals, utilizing independent time-conditioned gates to regulate their contributions. Furthermore, we propose Rhythm-anchored Trajectory Guidance (RTG), a training-free mechanism that evaluates hierarchical multimodal corrections at a visual-only predictive midpoint, safely strengthening semantic conditioning while better preserving temporal alignment. Finally, we construct DiverseDub, a multi-scenario benchmark to evaluate video dubbing in the wild. Extensive experiments demonstrate that UltraDub achieves state-of-the-art performance across four datasets.
Figures & tables
Figure 1: (a) Illustration of the V2C task. (b) Conventional CFG relies on linear extrapolation from fully conditioned velocity ( vc ) via scaling factors β∗ , making it prone to off-manifold trajectory drift. (c) For the multi-conditional V2C task, existing guidance easily triggers imbalance and degradation.
Figure 2: Overview of UltraDub, a U nifying Visually Steered Flow l earning and tra jectory Guidance Dub bing framework. The visually steered flow learning phase consists of an input projection, the core Motion-guided Dual-context Retrieving (MDR) layers, and a final modulation to predict the target velocity field vθ . With the velocity field established, Rhythm-anchored Trajectory Guidance (RTG) acts as a training-free inference mechanism. It constructs a visual-only midpoint as a rhythm anchor to evaluate multimodal corrections, continuously rectifying the base transport velocity.
Figure 3: Statistical overview of DiverseDub . (a) Left: distribution of clips over the eight dubbing categories, with a representative frame for each category. (b) Middle top: per-domain distributions of clip-level DNSMOS quality. (c) Middle bottom: per-domain distributions of SyncNet audiovisual offsets. (d) Right: relative frequency of word types among the nouns and adjectives of each domain, normalized by the corpus average.
Method
WER (%) ↓
SPKSIM (%) ↑
EmoSIM (%) ↑
LSE-C ↑
LSE-D ↓
GT
1.07
69.92
100.00
6.64
7.83
ProDubber
13.86
25.17
69.29
2.40
11.87
VoiceCraftDub
84.68
41.26
72.90
4.89
9.60
AlignDiT
20.61
49.78
80.09
6.33
8.15
InstructDubber
8.38
23.87
67.72
1.85
12.47
CoSyncDiT
12.01
56.32
81.72
6.88
7.86
Table 1: Results on the DiverseDub benchmark. Arrows indicate the preferred direction; the best result among dubbing methods is shown in bold.
Method
WER (%) ↓
SPKSIM (%) ↑
EmoSIM (%) ↑
LSE-C ↑
LSE-D ↓
GT
4.15
66.13
100.00
6.88
7.39
ProDubber
9.71
22.32
65.97
2.44
11.94
VoiceCraftDub
41.19
34.65
72.89
5.46
8.83
AlignDiT
25.86
46.25
74.89
6.29
8.01
InstructDubber
6.47
24.82
72.43
2.43
11.96
CoSyncDiT
12.52
49.05
78.93
6.86
7.56
Table 2: Results on the CelebV-Dub benchmark. Arrows indicate the preferred direction; the best result among dubbing methods is shown in bold.
Table 3: Comparison with dubbing methods on the LRS3 benchmark. Best results among dubbing methods are shown in bold.
Figure 5: Preserving audiovisual alignment under increasing guidance strength on the LRS3 test set. AVSync is a recently introduced metric that provides a more reliable assessment of audiovisual synchronization Choi et al. (2025) .
Methods
WER (%) ↓
SPKSIM (%) ↑
LSE-C ↑
LSE-D ↓
GT
13.14
63.69
7.14
6.84
ProDubber
44.29
34.91
3.71
11.08
VoiceCraftDub
64.47
33.37
5.13
8.50
AlignDiT
19.47
55.06
7.17
6.84
InstructDubber
27.67
36.30
3.33
11.32
CoSyncDiT
19.64
53.45
7.21
6.86
Table 4: Comparison with dubbing methods on the GRID benchmark. EmoSIM is not evaluated because GRID consists primarily of neutral emotion.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Mel-spectrogram comparisons between ground-truth speech and speech synthesized by different models on DiverseDub, LRS3, and GRID. The red and white rectangles highlight audiovisual alignment intervals and local phonetic details, respectively.
Figure 7: More category-wise comparisons on DiverseDub. Error bars indicate 95% confidence intervals. Further details and qualitative examples are provided in the supplementary demo.
Figure 8: Preserving audiovisual alignment under increasing guidance strength on the CelebV-Dub test set. AVSync is a recently introduced metric that provides a more reliable assessment of audiovisual synchronization Choi et al. (2025) .
Methods
WER (%) ↓
SPKSIM (%) ↑
LSE-C ↑
LSE-D ↓
GT
2.33
56.18
7.80
6.85
SA+FFN (base)
1.52
47.12
3.17
11.25
+MDR (Single Linguistic Memory)
1.43
46.53
8.17
6.59
+MDR (Single Speaker-Style Memory)
1.58
51.92
8.13
6.64
+MDR (Dual-context Memories)
1.39
51.49
8.18
6.56
RTG + MDR (Dual-context Memories)
1.37
52.05
8.23
6.50
Appendix
Table 5: Component ablation study of UltraDub on LRS3-Cross. Starting from a self-attention and feed-forward network (SA+FFN) backbone, we evaluate different MDR memory configurations and the contribution of RTG. Arrows indicate the preferred metric directions, and bold denotes the best result. Each architectural variant is trained independently under the same training configuration, while RTG is applied only at inference time.
Method
WER (%) ↓
SPKSIM (%) ↑
EmoSIM (%) ↑
LSE-C ↑
LSE-D ↓
AVSync ↑
GT
1.07
69.92
100.00
6.64
7.83
1.000
ProDubber
13.86
25.17
69.29
2.40
11.87
0.175
VoiceCraftDub
84.68
41.26
72.90
4.89
9.60
0.248
AlignDiT
20.61
49.78
80.09
6.33
8.15
0.344
InstructDubber
8.38
23.87
67.72
1.85
12.47
0.129
CoSyncDiT
12.01
56.32
81.72
6.88
7.86
0.365
Appendix
Table 6: Results on the DiverseDub benchmark. Arrows indicate the preferred direction; the best result among dubbing methods is shown in bold. AVSync is the mean cosine similarity between AV-HuBERT features of generated and GT audiovisual pairs; higher is better.
Methods
WER (%) ↓
SPKSIM (%) ↑
LSE-C ↑
LSE-D ↓
AVSync ↑
GT
2.33
56.18
7.80
6.85
1.000
ProDubber
3.41
24.27
4.27
10.12
0.220
VoiceCraftDub
14.85
35.17
6.49
8.08
0.568
AlignDiT
1.40
51.46
7.91
6.78
0.751
InstructDubber
2.39
24.16
3.29
11.12
0.172
CoSyncDiT
2.09
47.53
7.83
6.88
0.731
Appendix
Table 7: Comparison with dubbing methods on the LRS3-Cross benchmark. Best results among dubbing methods are shown in bold. Note that EmoSIM is not evaluated here, as LRS3 predominantly features neutral TED talks with limited emotional variation. AVSync is the mean cosine similarity between AV-HuBERT features of generated and GT audiovisual pairs; higher is better.
Figure 9: Instructions and an example item for the human listening test of overall speech naturalness.
Method
Naturalness (MOS-N) ↑
Similarity (MOS-S) ↑
Ground Truth
4.15 ± 0.12
4.01 ± 0.13
ProDubber
3.10 ± 0.17
2.55 ± 0.13
VoiceCraftDub
3.31 ± 0.09
3.15 ± 0.15
InstructDubber
3.36 ± 0.12
2.60 ± 0.14
AlignDiT
3.82 ± 0.14
3.99 ± 0.13
CoSyncDiT
3.86 ± 0.12
3.98 ± 0.15
Appendix
Table 8: Subjective evaluation results on LRS3-Cross for naturalness (MOS-N) and speaker similarity (MOS-S).
School of Informatics, Xiamen University, China · MiLM Plus, Xiaomi Inc., China · School of Electronic Science and Engineering, Xiamen University, China +1