Lip movements can match the timing of speech without matching the spoken sounds. We introduce PVSync, a unified model for audio-visual offset estimation and phoneme-level articulation scoring. PVSync combines window-level contrastive learning for synchronisation with a phoneme-level articulation objective that aligns audio and video embeddings of the same viseme class across clips. Visemes group phonemes with similar visible articulation. Viseme labels are derived automatically from forced-aligned transcripts, without manual annotations. On offset-corrected videos from 13 talking-head video generation models, PVSync matches human rankings of lip-sync quality more closely than LSE-C, achieving a Spearman correlation of 0.83 versus 0.34. On an automatically constructed benchmark from held-out speech, PVSync distinguishes viseme-matched from mismatched audio-visual pairs with an ROC AUC of 0.91. PVSync also outperforms SyncNet and MTD-VocaLiST in temporal offset recovery on held-out in-the-wild clips. Code and benchmark data will be released upon acceptance.
Figures & tables
Figure 1: (a) A local articulation error caught on a generated clip: the per-phone score of the /p/, where the lips never close, falls below the flag threshold. (b) Human rank against metric rank for thirteen talking-head video generation models.
Figure 2: Training paths; encoders and heads are shared. A. Window branch, every stage: a 12 -frame window’s tokens are mean-pooled, projected and scored by cosine under Lsync ; stage 3 trains only the heads. B. Phone branch, stage 2 only: a forced-aligned interval Δ selects the same tokens in both towers before the mean, and Lvis pulls the span toward same-viseme spans of other clips.
Model (input frames)
K=12
K=16
K=20
K=25
PVSync stage 1 (12)
0.906
0.944
0.962
0.975
PVSync stage 2 (12)
0.892
0.938
0.960
0.976
PVSync stage 3 (12)
0.927
0.951
0.968
0.981
Lvis alone (12)
0.210
0.256
0.322
0.396
+ head fine-tuning under Lsync (12)
0.827
0.902
0.939
0.964
SyncNet [ 1 ] (5)
0.800
0.883
0.929
0.960
Table 1: Lip-synchronisation accuracy (offset correct within one frame) on 3698 held-out clips at context window K ( Section 3.2 ). Best in bold, second best underlined.
Model
mean
std
std 16
% of windows with ∣ error ∣
=0
≤1
≥5
PVSync stage 1
−0.14
1.72
1.49
52.3
90.3
2.2
PVSync stage 2
+0.04
2.35
1.93
53.9
89.2
4.6
PVSync stage 3
+0.03
1.93
1.65
63.7
92.6
3.1
SyncNet
−0.05
6.35
3.06
27.2
51.6
35.4
MTD-VocaLiST
−0.40
5.06
2.14
32.8
66.8
22.5
Table 2: Per-window offset error in frames at 25truefps . Each of the 528757 windows predicts an offset from its own input alone ( Section 3.2 ). Mean and std are over all windows, std 16 the std when the 16 -frame context of Table 1 is used. Best in bold, second best underlined.
Model
K=13
K=15
K=16
SyncNet [ 1 ]
94.5
96.1
—
Perfect Match
98.7
99.1
—
AVST
99.3
99.6
—
VocaLiST
99.6
99.8
—
SyncNet, our run
95.0
96.7
97.3
MTD-VocaLiST [ 5 ] , our run
99.2
99.5
99.6
Table 3: Lip-synchronisation accuracy on the LRS2 test set ( 1243 clips, correct within one frame, VocaLiST’s protocol and code) at context window K . Published rows are from VocaLiST’s Table 1 [ 4 ] and are trained on LRS2; the lower rows are our runs. Best in bold, second best underlined.
ROC AUC
bad pairs flagged (%)
retrieval@ k
@10
@5
@1
k=1
k=2
PVSync stage 1
0.694 [.684,.705]
21.3
10.8
1.9
0.626
0.802
Lvis only
0.964 [.959,.969]
90.5
82.9
57.8
0.807
0.933
PVSync stage 2
0.960 [.955,.965]
89.7
80.8
52.8
0.811
0.941
PVSync stage 3
0.909 [.901,.918]
73.0
59.6
28.5
0.741
0.900
Table 4: Viseme-mismatch flagging on HDTF: each of 2163 video spans scored against every audio span from a different speaker ( 4.6 M pairs; Section 3.3 ). Brackets: 95true% clip-clustered bootstrap. Best in bold, second best underlined.
class
bil
lab
d/a
vel
rnd
rho
opn
spr
n
bilabial
98 / 98
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
1/1
⋅ / ⋅
⋅ / ⋅
252
labiodental
2/2
97 / 97
1/1
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
⋅ / ⋅
266
dental/alv.
3/3
1/2
78 / 78
5/3
2/6
5/2
3/3
2/2
240
velar
3/4
⋅ / ⋅
6/11
65 / 33
6/13
5/4
4/12
11/21
290
rounded
1/1
⋅ / ⋅
⋅ / ⋅
1/1
82 / 85
11/8
4/4
⋅ / ⋅
330
rhotic
1/2
1/3
3/5
2/3
15/27
75 / 56
1/3
2/1
264
Table 5: Retrieval confusion (%) at k=1 ( Table 4 ), stage 2/stage 3 in each cell. Rows are the true class; the bold diagonal is each class’s retrieval accuracy, and cells below 0.5 % are ⋅ . Classes: bilabial /p b m/; labiodental /f v/; dental/alveolar /t d n s z l T D S Z tS dZ/; velar /k g N h/; rounded /oU u w U OI O/; spread /i I eI E j/; open /A æ 2 aU aI/; rhotic /r /.
Model, chunk score
ROC AUC [CI]
Catch @16%
SyncNet [ 1 ] , L2 distance
0.664 [.581,.739]
0.28
SyncNet, LSE-C
0.705 [.624,.778]
0.38
MTD-VocaLiST [ 5 ] , sync prob.
0.681 [.607,.754]
0.30
StableSyncNet [ 3 ] , cosine
0.794 [.722,.859]
0.56
PVSync stage 3, window
0.788 [.711,.859]
0.56
PVSync stage 3, worst phoneme
0.836 [.781,.885]
0.70
Table 6: ROC AUC of each chunk score against the human majority label ( 400 chunks of 480truems , generated video) and flagged chunks caught at a 16true% false-flag rate. Brackets: 95true% clip-clustered bootstrap. Best in bold, second best underlined (models only).