Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly suited to this setting: even when fine-tuned to recover spoken content from lip motion, they remain largely insensitive to temporal errors. We introduce Align Then Reason (ATR), a multilingual lip-sync judge that first establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, then reasons over this alignment to make the final judgment. An alignment scorer provides the LLM with both local evidence for each phonetic unit and a calibrated global alignment score, enabling it to reason jointly about content and timing. On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively. The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively. They also transfer across datasets to three unseen MuAViC languages. Furthermore, we evaluate on two downstream tasks built from real dubbing lines. On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.
Figures & tables
Figure 1: Overview of Align Then Reason ( atr ). Stage 1 (Align): the silent mouth video is encoded into frame embeddings h1:T and the candidate line into phonetic-unit embeddings u1:N . Their similarity Mt,n is scored by CTC whose labels are the positions of the units in the line, so only a monotonic path through M (highlighted) scores well, and reordering, shifting, or freezing the frames breaks it. Trained with content and temporal negatives, the scorer emits a calibrated global score a(V,c) and per-unit soft tokens e1:Nsoft . Stage 2 (Reason): the soft tokens are prepended to a prompt containing a(V,c) and the line, and an LLM answers whether the mouth matches the line; the difference of its Yes and No logits is the lip-sync score.
Content
Temporal
Method
Mismatch
Shuffle
Dub
Reverse
Shift
Freeze
Swap
Mean
Auto-AVSR
0.602
0.599
0.507
0.670
0.582
0.609
0.620
0.598
AV-HuBERT
0.500
0.578
0.520
0.513
0.506
0.501
0.508
0.518
LLaMA-AVSR
0.472
0.793
0.484
0.531
0.513
0.500
0.521
0.545
VSP-LLM
0.481
0.764
0.508
0.524
0.507
0.511
0.517
0.545
Qwen3.5-2B Base
0.515
0.448
0.570
0.503
0.499
0.503
0.501
0.506
Table 1: Performance on the seven-language benchmark measured by pooled AUC ( ↑ ). Existing speech recognition and vision-language baselines struggle on temporal corruptions, staying near chance ( 0.50 ). In contrast, atr achieves balanced, high accuracy across all seven evaluation axes and multiple LLM backbones, reaching a mean AUC of 0.920 with Qwen3.5-9B, a 50.8% relative improvement over the best baseline ( 0.610 ).
Content
Temporal
Method
Mismatch
Shuffle
Dub
Reverse
Shift
Freeze
Swap
Mean
Auto-AVSR
0.624
0.751
0.509
0.956
0.892
0.710
0.897
0.763
AV-HuBERT
0.496
0.639
0.519
0.524
0.541
0.514
0.519
0.536
LLaMA-AVSR
0.471
0.874
0.496
0.607
0.596
0.463
0.612
0.588
VSP-LLM
0.480
0.900
0.534
0.597
0.534
0.518
0.585
0.593
Qwen3.5-9B Base
0.553
0.456
0.595
0.522
0.499
0.552
0.499
0.525
Table 2: Paired accuracy ( ↑ ) on the seven-language benchmark. Baseline speech recognizers and frontier models degrade significantly under temporal corruptions, demonstrating insensitivity to fine-grained timing. atr consistently achieves superior accuracy across all content and temporal axes regardless of the underlying LLM backbone.
Method
Mismatch
Shuffle
Reverse
Shift
Freeze
Swap
Mean
Auto-AVSR
0.546
0.552
0.635
0.555
0.601
0.580
0.578
Qwen3.5-9B Base
0.580
0.498
0.501
0.501
0.502
0.501
0.514
atr -9B
0.661
0.930
0.790
0.763
0.572
0.787
0.751
atr -9B (recalibrated)
0.707
0.946
0.927
0.827
0.937
0.905
0.875
Table 3: Generalization on MuAViC across three unseen languages (German, Arabic, and Russian). No model weights are updated. Out-of-the-box atr -9B achieves 0.751 mean AUC. Adjusting the two normalization constants ( μs,σs ) for recalibration lifts the mean AUC to 0.875 , outperforming Auto-AVSR by +51.4% relatively.
Content
Temporal
Method
Mismatch
Shuffle
Dub
Reverse
Shift
Freeze
Swap
Mean
atr -9B
0.786
0.974
0.862
0.964
0.931
0.970
0.952
0.920
w/o phonetic supervision
0.751
0.944
0.793
0.961
0.921
0.965
0.949
0.898
w/o LLM reasoner
0.695
0.757
0.701
0.970
0.935
0.971
0.962
0.856
w/o soft tokens
0.683
0.932
0.748
0.953
0.905
0.942
0.940
0.872
w/o calibrated scalar
0.671
0.927
0.724
0.525
0.534
0.805
0.517
0.672
Table 4: Ablations on core components of atr (Macro AUC ↑ ). Phonetic supervision and the LLM reasoner are critical for content understanding, while soft tokens provide local alignment evidence. The calibrated scalar is essential for temporal sensitivity; removing it causes the largest degradation (reducing mean AUC to 0.672). The full model achieves the best overall performance.
Method
Top-1 ↑
MRR ↑
Chance
0.250
0.521
Isochrony (length fit)
0.237
0.494
Auto-AVSR
0.315
0.558
AV-HuBERT
0.244
0.510
LLaMA-AVSR
0.316
0.575
VSP-LLM
0.309
0.572
Table 5: Dub-line reranking. Each clip is paired with its professional dub line and three meaning-preserving, length-matched paraphrases; the model must rank the professional line first. The alignment scorer achieves 0.479 Top-1 accuracy, outperforming the strongest lip-reading baseline ( 0.316 ) by 52% relative. MRR denotes mean reciprocal rank.
In-domain (5-way)
MuAViC (2-way)
Method
Per-clip ↑
Exact Block ↑
Per-clip ↑
Chance
0.200
0.008
0.500
Isochrony (length fit)
0.656
0.397
0.500
Auto-AVSR
0.634
0.383
0.598
AV-HuBERT
0.365
0.081
0.548
LLaMA-AVSR
0.396
0.134
0.571
Table 6: Script-to-clip assignment. In-domain, we form five-clip blocks from the same title and language and match their five shuffled professional lines one-to-one; Per-clip measures individual assignments, while Exact Block requires all five to be correct. atr -9B achieves 0.746 per-clip accuracy, improving over Auto-AVSR by 17.7% . On MuAViC, where the available evaluation naturally forms two-way clip–line pairs, atr reaches 0.704 versus 0.598 for Auto-AVSR.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Content
Temporal
Method
Mismatch
Shuffle
Dub
Reverse
Shift
Freeze
Swap
Auto-AVSR
0.784
0.817
0.946
0.896
0.639
0.714
0.825
AV-HuBERT
0.544
0.864
0.964
0.584
0.512
0.527
0.545
LLaMA-AVSR
0.587
0.930
0.505
0.599
0.531
0.521
0.546
VSP-LLM
0.562
0.927
0.477
0.598
0.533
0.549
0.575
Qwen3.5-9B Base
0.483
0.535
0.737
0.502
0.502
0.506
0.498
Appendix
Table 7: English-only evaluation measured by pooled AUC ( ↑ ). Recognizer baselines were trained primarily on English data. While Auto-AVSR achieves strong content performance on English (winning mismatch ), atr -9B outperforms all baselines across the remaining six content and temporal axes.
Mismatch
Reverse
Freeze
Language
Auto-AVSR
Scorer
Auto-AVSR
Scorer
Auto-AVSR
Scorer
en
0.78
0.63
0.90
0.98
0.71
0.96
es
0.60
0.72
0.65
0.96
0.65
0.97
fr
0.57
0.67
0.64
0.97
0.64
0.97
ja
0.50
0.73
0.57
0.97
0.46
0.99
ko
0.50
0.78
0.60
0.98
0.54
0.97
Appendix
Table 8: Per-language AUC comparison between Auto-AVSR and our alignment scorer on three representative axes. Auto-AVSR degrades to chance level on non-Latin scripts (Japanese and Korean) due to tokenization constraints and shows weak temporal sensitivity on non-English languages. Our alignment scorer maintains consistent effectiveness across all seven languages, remaining near ceiling ( ≥0.96 AUC) on temporal corruptions.
Model Backbone
Seven-Language Macro
MuAViC (Zero-Shot)
MuAViC (Recalibrated)
atr (Qwen3.5-2B)
0.884±0.018
0.746±0.074
0.815±0.041
atr (Qwen3.5-4B)
0.895±0.023
0.748±0.070
0.802±0.073
atr (Qwen3.5-9B)
0.910±0.015
0.778±0.060
0.852±0.033
atr (LLaMA-3.1-8B)
0.895±0.012
0.711±0.018
0.832±0.011
atr (Mistral-7B)
0.900±0.020
0.777±0.085
0.821±0.045
Appendix
Table 9: Multi-seed evaluation reporting mean AUC and standard deviation across three random training seeds. Performance gains remain stable across diverse LLM reasoner backbones on the 7-language benchmark, out-of-domain MuAViC dataset, and recalibrated MuAViC evaluation.
School of Informatics, Xiamen University, China · MiLM Plus, Xiaomi Inc., China · School of Electronic Science and Engineering, Xiamen University, China +1