Lip movements can match the timing of speech without matching the spoken sounds. We introduce PVSync, a unified model for audio-visual offset estimation and phoneme-level articulation scoring. PVSync combines window-level contrastive learning for synchronisation with a phoneme-level articulation objective that aligns audio and video embeddings of the same viseme class across clips. Visemes group phonemes with similar visible articulation. Viseme labels are derived automatically from forced-aligned transcripts, without manual annotations. On offset-corrected videos from 13 talking-head video generation models, PVSync matches human rankings of lip-sync quality more closely than LSE-C, achieving a Spearman correlation of 0.83 versus 0.34. On an automatically constructed benchmark from held-out speech, PVSync distinguishes viseme-matched from mismatched audio-visual pairs with an ROC AUC of 0.91. PVSync also outperforms SyncNet and MTD-VocaLiST in temporal offset recovery on held-out in-the-wild clips. Code and benchmark data will be released upon acceptance.
Figures & tables
Figure 1: (a) A local articulation error caught on a generated clip: the per-phone score of the /p/, where the lips never close, falls below the flag threshold. (b) Human rank against metric rank for thirteen talking-head video generation models.
Figure 2: Training paths; encoders and heads are shared. A. Window branch, every stage: a 12 -frame window’s tokens are mean-pooled, projected and scored by cosine under Lsync ; stage 3 trains only the heads. B. Phone branch, stage 2 only: a forced-aligned interval Δ selects the same tokens in both towers before the mean, and Lvis pulls the span toward same-viseme spans of other clips.
Model (input frames)
K=12
K=16
K=20
K=25
PVSync stage 1 (12)
0.906
0.944
0.962
0.975
PVSync stage 2 (12)
0.892
0.938
0.960
0.976
PVSync stage 3 (12)
0.927
0.951
0.968
0.981
Lvis alone (12)
0.210
0.256
0.322
0.396
+ head fine-tuning under Lsync (12)
0.827
0.902
0.939
0.964
SyncNet [ 1 ] (5)
0.800
0.883
0.929
0.960
Table 1: Lip-synchronisation accuracy (offset correct within one frame) on 3698 held-out clips at context window K ( Section 3.2 ). Best in bold, second best underlined.
Model
mean
std
std 16
% of windows with ∣ error ∣
=0
≤1
≥5
PVSync stage 1
−0.14
1.72
1.49
52.3
90.3
2.2
PVSync stage 2
+0.04
2.35
1.93
53.9
89.2
4.6
PVSync stage 3
+0.03
1.93
1.65
63.7
92.6
3.1
SyncNet
−0.05
6.35
3.06
27.2
51.6
35.4
MTD-VocaLiST
−0.40
5.06
2.14
32.8
66.8
22.5
Table 2: Per-window offset error in frames at 25truefps . Each of the 528757 windows predicts an offset from its own input alone ( Section 3.2 ). Mean and std are over all windows, std 16 the std when the 16 -frame context of Table 1 is used. Best in bold, second best underlined.
Model
K=13
K=15
K=16
SyncNet [ 1 ]
94.5
96.1
—
Perfect Match
98.7
99.1
—
AVST
99.3
99.6
—
VocaLiST
99.6
99.8
—
SyncNet, our run
95.0
96.7
97.3
MTD-VocaLiST [ 5 ] , our run
99.2
99.5
99.6
Table 3: Lip-synchronisation accuracy on the LRS2 test set ( 1243 clips, correct within one frame, VocaLiST’s protocol and code) at context window K . Published rows are from VocaLiST’s Table 1 [ 4 ] and are trained on LRS2; the lower rows are our runs. Best in bold, second best underlined.
ROC AUC
bad pairs flagged (%)
retrieval@ k
@10
@5
@1
k=1
k=2
PVSync stage 1
0.694 [.684,.705]
21.3
10.8
1.9
0.626
0.802
Lvis only
0.964 [.959,.969]
90.5
82.9
57.8
0.807
0.933
PVSync stage 2
0.960 [.955,.965]
89.7
80.8
52.8
0.811
0.941
PVSync stage 3
0.909 [.901,.918]
73.0
59.6
28.5
0.741
0.900
Table 4: Viseme-mismatch flagging on HDTF: each of 2163 video spans scored against every audio span from a different speaker ( 4.6 M pairs; Section 3.3 ). Brackets: 95true% clip-clustered bootstrap. Best in bold, second best underlined.
class
bil
lab
d/a
vel
rnd
rho
opn
spr
n
bilabial
98 / 98
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
1/1
⋅ / ⋅
⋅ / ⋅
252
labiodental
2/2
97 / 97
1/1
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
266
dental/alv.
3/3
1/2
78 / 78
5/3
2/6
5/2
3/3
2/2
240
velar
3/4
⋅ / ⋅
6/11
65 / 33
6/13
5/4
4/12
11/21
290
rounded
1/1
⋅ / ⋅
⋅ / ⋅
1/1
82 / 85
11/8
4/4
⋅ / ⋅
330
rhotic
1/2
1/3
3/5
2/3
15/27
75 / 56
1/3
2/1
264
Table 5: Retrieval confusion (%) at k=1 ( Table 4 ), stage 2/stage 3 in each cell. Rows are the true class; the bold diagonal is each class’s retrieval accuracy, and cells below 0.5 % are ⋅ . Classes: bilabial /p b m/; labiodental /f v/; dental/alveolar /t d n s z l T D S Z tS dZ/; velar /k g N h/; rounded /oU u w U OI O/; spread /i I eI E j/; open /A æ 2 aU aI/; rhotic /r /.
Model, chunk score
ROC AUC [CI]
Catch @16%
SyncNet [ 1 ] , L2 distance
0.664 [.581,.739]
0.28
SyncNet, LSE-C
0.705 [.624,.778]
0.38
MTD-VocaLiST [ 5 ] , sync prob.
0.681 [.607,.754]
0.30
StableSyncNet [ 3 ] , cosine
0.794 [.722,.859]
0.56
PVSync stage 3, window
0.788 [.711,.859]
0.56
PVSync stage 3, worst phoneme
0.836 [.781,.885]
0.70
Table 6: ROC AUC of each chunk score against the human majority label ( 400 chunks of 480truems , generated video) and flagged chunks caught at a 16true% false-flag rate. Brackets: 95true% clip-clustered bootstrap. Best in bold, second best underlined (models only).
We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile image quality with synchronization accuracy, producing either visually degraded outputs or temporally inconsistent lip movements. HighSync addresses both challenges simultaneously and, to our knowledge, is the first lip sync model to operate natively at 512*512 resolution, positioning it as a viable solution for professional production environments such as the film and broadcast industries. Central to our approach is the identification and systematic elimination of a data leakage phenomenon that has silently undermined temporal modeling in prior work, preventing models from developing a genuine dependence on the audio signal. Comprehensive evaluations across both perceptual quality and synchronization accuracy metrics confirm that HighSync achieves state-of-the-art performance on both fronts. Source code, pre-trained models, and supplementary video results are publicly available at: https://github.com/saeed5959/high_sync
Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many denoising steps make them impractical for real-time inference. We present Lip Forcing, to our knowledge the first autoregressive diffusion method for video-to-video (V2V) lip synchronization, which distills a 14B audio-conditioned bidirectional video diffusion teacher into causal students. At inference, the students generate each chunk in only two denoising steps without inference-time CFG, enabling real-time lip synchronization. A lip-sync-specific teacher-trajectory analysis reveals a CFG fidelity-sync tradeoff: no-CFG predictions favor reference fidelity, whereas CFG-guided predictions favor synchronization within a mid-trajectory band. Lip Forcing translates this finding into three analysis-derived components: Sync-Window DMD, a two-step inference schedule, and a SyncNet-based reward. We validate Lip Forcing at two student scales, both distilled from the 14B teacher. The 1.3B student crosses into real-time streaming at 31 FPS, 17.6× faster than its same-scale bidirectional model. The 14B student, the largest diffusion model reported for V2V lip synchronization, runs 39.8× faster than its teacher at comparable reference fidelity. Time-to-first-frame is sub-millisecond at both scales, far below every diffusion baseline.
Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly suited to this setting: even when fine-tuned to recover spoken content from lip motion, they remain largely insensitive to temporal errors. We introduce Align Then Reason (ATR), a multilingual lip-sync judge that first establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, then reasons over this alignment to make the final judgment. An alignment scorer provides the LLM with both local evidence for each phonetic unit and a calibrated global alignment score, enabling it to reason jointly about content and timing. On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively. The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively. They also transfer across datasets to three unseen MuAViC languages. Furthermore, we evaluate on two downstream tasks built from real dubbing lines. On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.