Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly suited to this setting: even when fine-tuned to recover spoken content from lip motion, they remain largely insensitive to temporal errors. We introduce Align Then Reason (ATR), a multilingual lip-sync judge that first establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, then reasons over this alignment to make the final judgment. An alignment scorer provides the LLM with both local evidence for each phonetic unit and a calibrated global alignment score, enabling it to reason jointly about content and timing. On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively. The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively. They also transfer across datasets to three unseen MuAViC languages. Furthermore, we evaluate on two downstream tasks built from real dubbing lines. On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.
Figures & tables
Figure 1: Overview of Align Then Reason ( atr ). Stage 1 (Align): the silent mouth video is encoded into frame embeddings h1:T and the candidate line into phonetic-unit embeddings u1:N . Their similarity Mt,n is scored by CTC whose labels are the positions of the units in the line, so only a monotonic path through M (highlighted) scores well, and reordering, shifting, or freezing the frames breaks it. Trained with content and temporal negatives, the scorer emits a calibrated global score a(V,c) and per-unit soft tokens e1:Nsoft . Stage 2 (Reason): the soft tokens are prepended to a prompt containing a(V,c) and the line, and an LLM answers whether the mouth matches the line; the difference of its Yes and No logits is the lip-sync score.
Content
Temporal
Method
Mismatch
Shuffle
Dub
Reverse
Shift
Freeze
Swap
Mean
Auto-AVSR
0.602
0.599
0.507
0.670
0.582
0.609
0.620
0.598
AV-HuBERT
0.500
0.578
0.520
0.513
0.506
0.501
0.508
0.518
LLaMA-AVSR
0.472
0.793
0.484
0.531
0.513
0.500
0.521
0.545
VSP-LLM
0.481
0.764
0.508
0.524
0.507
0.511
0.517
0.545
Qwen3.5-2B Base
0.515
0.448
0.570
0.503
0.499
0.503
0.501
0.506
Table 1: Performance on the seven-language benchmark measured by pooled AUC ( ↑ ). Existing speech recognition and vision-language baselines struggle on temporal corruptions, staying near chance ( 0.50 ). In contrast, atr achieves balanced, high accuracy across all seven evaluation axes and multiple LLM backbones, reaching a mean AUC of 0.920 with Qwen3.5-9B, a 50.8% relative improvement over the best baseline ( 0.610 ).
Content
Temporal
Method
Mismatch
Shuffle
Dub
Reverse
Shift
Freeze
Swap
Mean
Auto-AVSR
0.624
0.751
0.509
0.956
0.892
0.710
0.897
0.763
AV-HuBERT
0.496
0.639
0.519
0.524
0.541
0.514
0.519
0.536
LLaMA-AVSR
0.471
0.874
0.496
0.607
0.596
0.463
0.612
0.588
VSP-LLM
0.480
0.900
0.534
0.597
0.534
0.518
0.585
0.593
Qwen3.5-9B Base
0.553
0.456
0.595
0.522
0.499
0.552
0.499
0.525
Table 2: Paired accuracy ( ↑ ) on the seven-language benchmark. Baseline speech recognizers and frontier models degrade significantly under temporal corruptions, demonstrating insensitivity to fine-grained timing. atr consistently achieves superior accuracy across all content and temporal axes regardless of the underlying LLM backbone.
Method
Mismatch
Shuffle
Reverse
Shift
Freeze
Swap
Mean
Auto-AVSR
0.546
0.552
0.635
0.555
0.601
0.580
0.578
Qwen3.5-9B Base
0.580
0.498
0.501
0.501
0.502
0.501
0.514
atr -9B
0.661
0.930
0.790
0.763
0.572
0.787
0.751
atr -9B (recalibrated)
0.707
0.946
0.927
0.827
0.937
0.905
0.875
Table 3: Generalization on MuAViC across three unseen languages (German, Arabic, and Russian). No model weights are updated. Out-of-the-box atr -9B achieves 0.751 mean AUC. Adjusting the two normalization constants ( μs,σs ) for recalibration lifts the mean AUC to 0.875 , outperforming Auto-AVSR by +51.4% relatively.
Content
Temporal
Method
Mismatch
Shuffle
Dub
Reverse
Shift
Freeze
Swap
Mean
atr -9B
0.786
0.974
0.862
0.964
0.931
0.970
0.952
0.920
w/o phonetic supervision
0.751
0.944
0.793
0.961
0.921
0.965
0.949
0.898
w/o LLM reasoner
0.695
0.757
0.701
0.970
0.935
0.971
0.962
0.856
w/o soft tokens
0.683
0.932
0.748
0.953
0.905
0.942
0.940
0.872
w/o calibrated scalar
0.671
0.927
0.724
0.525
0.534
0.805
0.517
0.672
Table 4: Ablations on core components of atr (Macro AUC ↑ ). Phonetic supervision and the LLM reasoner are critical for content understanding, while soft tokens provide local alignment evidence. The calibrated scalar is essential for temporal sensitivity; removing it causes the largest degradation (reducing mean AUC to 0.672). The full model achieves the best overall performance.
Method
Top-1 ↑
MRR ↑
Chance
0.250
0.521
Isochrony (length fit)
0.237
0.494
Auto-AVSR
0.315
0.558
AV-HuBERT
0.244
0.510
LLaMA-AVSR
0.316
0.575
VSP-LLM
0.309
0.572
Table 5: Dub-line reranking. Each clip is paired with its professional dub line and three meaning-preserving, length-matched paraphrases; the model must rank the professional line first. The alignment scorer achieves 0.479 Top-1 accuracy, outperforming the strongest lip-reading baseline ( 0.316 ) by 52% relative. MRR denotes mean reciprocal rank.
In-domain (5-way)
MuAViC (2-way)
Method
Per-clip ↑
Exact Block ↑
Per-clip ↑
Chance
0.200
0.008
0.500
Isochrony (length fit)
0.656
0.397
0.500
Auto-AVSR
0.634
0.383
0.598
AV-HuBERT
0.365
0.081
0.548
LLaMA-AVSR
0.396
0.134
0.571
Table 6: Script-to-clip assignment. In-domain, we form five-clip blocks from the same title and language and match their five shuffled professional lines one-to-one; Per-clip measures individual assignments, while Exact Block requires all five to be correct. atr -9B achieves 0.746 per-clip accuracy, improving over Auto-AVSR by 17.7% . On MuAViC, where the available evaluation naturally forms two-way clip–line pairs, atr reaches 0.704 versus 0.598 for Auto-AVSR.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Content
Temporal
Method
Mismatch
Shuffle
Dub
Reverse
Shift
Freeze
Swap
Auto-AVSR
0.784
0.817
0.946
0.896
0.639
0.714
0.825
AV-HuBERT
0.544
0.864
0.964
0.584
0.512
0.527
0.545
LLaMA-AVSR
0.587
0.930
0.505
0.599
0.531
0.521
0.546
VSP-LLM
0.562
0.927
0.477
0.598
0.533
0.549
0.575
Qwen3.5-9B Base
0.483
0.535
0.737
0.502
0.502
0.506
0.498
Appendix
Table 7: English-only evaluation measured by pooled AUC ( ↑ ). Recognizer baselines were trained primarily on English data. While Auto-AVSR achieves strong content performance on English (winning mismatch ), atr -9B outperforms all baselines across the remaining six content and temporal axes.
Mismatch
Reverse
Freeze
Language
Auto-AVSR
Scorer
Auto-AVSR
Scorer
Auto-AVSR
Scorer
en
0.78
0.63
0.90
0.98
0.71
0.96
es
0.60
0.72
0.65
0.96
0.65
0.97
fr
0.57
0.67
0.64
0.97
0.64
0.97
ja
0.50
0.73
0.57
0.97
0.46
0.99
ko
0.50
0.78
0.60
0.98
0.54
0.97
Appendix
Table 8: Per-language AUC comparison between Auto-AVSR and our alignment scorer on three representative axes. Auto-AVSR degrades to chance level on non-Latin scripts (Japanese and Korean) due to tokenization constraints and shows weak temporal sensitivity on non-English languages. Our alignment scorer maintains consistent effectiveness across all seven languages, remaining near ceiling ( ≥0.96 AUC) on temporal corruptions.
Model Backbone
Seven-Language Macro
MuAViC (Zero-Shot)
MuAViC (Recalibrated)
atr (Qwen3.5-2B)
0.884±0.018
0.746±0.074
0.815±0.041
atr (Qwen3.5-4B)
0.895±0.023
0.748±0.070
0.802±0.073
atr (Qwen3.5-9B)
0.910±0.015
0.778±0.060
0.852±0.033
atr (LLaMA-3.1-8B)
0.895±0.012
0.711±0.018
0.832±0.011
atr (Mistral-7B)
0.900±0.020
0.777±0.085
0.821±0.045
Appendix
Table 9: Multi-seed evaluation reporting mean AUC and standard deviation across three random training seeds. Performance gains remain stable across diverse LLM reasoner backbones on the 7-language benchmark, out-of-domain MuAViC dataset, and recalibrated MuAViC evaluation.
Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works address this problem by conditioning the speech synthesis process on lip movements extracted from the video signal. In this work, we condition the speech generation on a binary voice-activity signal, which has a lightweight representation and can be produced in multiple ways. We show that the model follows the voice-activity signal with high accuracy while maintaining natural prosody and semantically appropriate pause placement within sentences, as demonstrated through extensive objective and subjective evaluations. By randomly masking this condition during training, we make the feature entirely optional during inference, allowing editors to enforce or relax lip-sync constraints when desired.
We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rather than a generic one. Existing lip-sync systems follow the audio closely and keep the face recognizable, yet the mouth they render is an average mouth: the shape and texture of the lips, the arrangement of the teeth, and how much of them shows as the mouth opens are not that person's. The problem persists because nothing in current training or evaluation asks for the person's own mouth: perceptual losses accept any plausible mouth, face identity is carried mostly by the skin around it, and the released inference code of inpainting systems uses the unmasked target frame as the reference, which hides the gap. To address this, RGOR conditions every generated frame on frames from separate enrollment recordings of the same person and on HD patches of the mouth that bypass the VAE, and trains the generator against a paired judge that compares each rendered mouth with the person's reference and learns to reject a realistic mouth of someone else. We further build an evaluation protocol and use it to compare open-source and commercial lip-sync systems on held-out identities. Experiments show that RGOR achieves the best or second-best result on most metrics, and preserves the person's own lip and dental detail while keeping synchronization and the rest of the face intact.
Automatic video dubbing aims to generate high-fidelity speech that is temporally aligned with visual content. However, existing methods still suffer from limited speech naturalness, insufficient audio-visual synchronization, and poor scalability beyond monolingual settings. To address these challenges, we propose SyncVoice, a simple and effective dubbing framework that lightly integrates a Text-Visual Fusion Module into a pretrained text-to-speech (TTS) system. This module aligns visual features with linguistic representations, enabling temporally synchronized speech synthesis without complex architectural redesign. Experiments on the LRS3 dataset show that SyncVoice achieves state-of-the-art performance in zero-shot dubbing. Further training on a large-scale bilingual audio-visual dataset improves vocal fidelity while preserving synchronization, yielding a single unified model for both Chinese and English dubbing.
Kaidi Wang, Yi He, Wenhao Guan +9
School of Informatics, Xiamen University, China · MiLM Plus, Xiaomi Inc., China · School of Electronic Science and Engineering, Xiamen University, China +1