Organizations: University of Science and Technology of China, HeFei, China · Shanghai Jiaotong University · Peking University · Nanyang Technological University
Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as advanced LipSync generation methods not only achieve better lip synchronization but also eliminate visual artifacts. An important reason is that they overlook an inherent biological coupling between lip movements and head poses in natural speech videos. In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip. For detection, the framework learns to quantify this discrepancy by contrasting lip and pose features from authentic versus forged videos. For attribution, our method is designed to capture the unique temporal dynamics and audio-visual synchronization patterns that act as the fingerprint of models, enabling source tracing. We conduct extensive experiments on two challenging LipSync datasets as well as our own proposed large-scale and multi-generator dataset. LipDA achieves over 97% AUC in detection and 97.5% accuracy in model attribution, significantly outperforming existing methods. Code and the proposed LipSync-A dataset are available at https://github.com/AnsonShe/LipDA.
Figures & tables
Figure 1 : Taxonomy of forgeries and detection comparison. (Left) Distinct from identity attacks, LipSync constitutes a fact attack, fabricating videos where the target utters speech never spoken. (Right) Vulnerability of existing defenses versus our method. LipDA robustly identifies forgeries by capturing the inconsistencies between lip and head dynamics.
Figure 2 : Illustration of the two primary LipSync forgery paradigms . Top (Video-driven): Motion and head pose are transplanted from a driving video, creating localized temporal artifacts. Bottom (Audio-driven): Lip motion is synthesized from audio, severing the global link to head motion, which remains static or is generated independently.
Figure 3 : AU intensity analysis on original video sequences and two categories of forgery patterns . Higher intensity values indicate heightened activity of specific facial muscles. Two representative action units, AU12 (mouth corner pull) and AU17 (chin raise), are analyzed.
Figure 4 : t-SNE visualization of pose features from LipSync forgeries generated by different model families .
Dataset
Detection
Attribution
# Gens.
# Fakes
TalkHeadBench ( Xiong et al., 2025 )
✓
✗
2
2,984
AVLips ( Liu et al., 2024 )
✓
✗
2
4,206
LAV-DF ( Cai et al., 2022 )
✓
✗
1
99,873
PolyGlotFake ( Hou et al., 2024 )
✓
✗
2
14,472
AV-Deepfake1M ( Cai et al., 2024 )
✓
✗
1
860,039
LipSync-A(Ours)
✓
✓
7
16,000
Table 1 : Comparison with existing LipSync datasets. #Gens. and #Fakes denote the number of generator architectures and forged video samples, respectively.
Figure 5 : Overview of our unified LipDA framework . 2-Stage Training (Left): In Stage I Detection, a dual-branch encoder is trained with a contrastive module to learn the lip-head pose inconsistency. In Stage II Attribution, Stage I encoders are frozen, and the Modality Synchronization (MSM) and Temporal Dynamic (TDM) modules are trained to capture unique generative fingerprints. Inference (Right): The unified pipeline integrates all components for joint forgery detection and source attribution.
In-domain
Cross-domain
Method
Modality
LipSync-A
AVLips
TalkHeadBench
Average
ACC ↑
AUC ↑
ACC ↑
AUC ↑
ACC ↑
AUC ↑
ACC ↑
AUC ↑
FTCN ( Zheng et al., 2021 )
V
78.98
87.78
65.67
71.55
68.73
75.46
71.13
78.26
CADDM ( Dong et al., 2023 )
V
47.15
63.14
45.90
55.92
62.53
63.69
51.86
60.92
LipForensics ( Haliassos et al., 2021 )
V
80.92
89.85
74.15
81.97
84.52
90.63
79.86
87.48
RealForensics ( Haliassos et al., 2022 )
V
72.44
67.03
80.47
89.25
71.03
79.01
74.65
78.43
Table 2 : Binary forgery detection performance comparison . We report Accuracy (ACC, %) and AUC (%) on three public benchmarks against SOTA methods. ’V’ denotes our visual-only model and ’A-V’ denotes our audio-visual variant, which additionally incorporates the MSM score for detection. Bold indicates the best results, while the second-ranking one is underscored.
Method
Transformer-based
GAN-based
Diffusion-based
VAE-based
Statistical models
Average
ACC ↑
F1 ↑
ACC ↑
F1 ↑
ACC ↑
F1 ↑
ACC ↑
F1 ↑
ACC ↑
F1 ↑
ACC ↑
F1 ↑
AVAD ( Feng et al., 2023 )
77.2
9.9
83.0
56.7
70.0
27.4
68.8
40.2
71.9
33.0
74.2
33.4
TALL ( Xu et al., 2024b )
89.0
59.3
92.0
83.0
97.8
96.7
96.0
96.6
86.7
72.0
92.3
81.5
AVH-align ( Smeu et al., 2025 )
71.8
41.0
96.5
69.2
99.9
98.5
96.9
90.9
99.9
88.9
93.0
77.7
SpeechForensics ( Liang et al., 2024 )
70.4
43.9
71.6
18.4
70.4
46.0
76.4
11.9
73.2
13.0
72.4
26.6
LipDA(Ours) A-V
93.8
87.6
94.9
86.7
99.9
99.9
99.9
99.9
98.9
95.2
97.5
93.9
Table 3 : Attribution performance comparison for classifying forgeries into five generator families . Metrics are Accuracy (ACC, %) and F1-Score (F1, %). All listed methods are audio-visual. Bold and underscored denote the best and second-best results, respectively.
Method
Sonic
KDTalker
OmniSync
ACC ↑
F1 ↑
ACC ↑
F1 ↑
ACC ↑
F1 ↑
LipFD ( Liu et al., 2024 )
48.6
7.1
45.6
6.9
51.2
5.8
AVAD ( Feng et al., 2023 )
49.6
66.3
49.7
66.4
49.5
66.2
SpeechForensics ( Liang et al., 2024 )
80.8
80.4
83.9
81.0
90.0
90.2
DFD-FCG ( Han et al., 2025 )
66.4
61.9
72.2
69.4
52.1
14.7
AVH-align ( Smeu et al., 2025 )
70.0
56.9
71.9
64.9
72.9
9.5
Table 4 : Cross-manipulation generalization on LipSync (A-V) . Unseen LipSync generalization on Sonic ( Ji et al., 2025 ) , KDTalker ( Yang et al., 2025b ) , and OmniSync ( Peng et al., 2025 ) .
Method
CelebDF
ACC ↑
AUC ↑
AP ↑
FPR ↓
FNR ↓
CADDM ( Dong et al., 2023 )
53.83
81.44
95.70
9.55
51.95
UnivFD ( Ojha et al., 2023 )
50.23
66.39
67.96
1.00
99.00
FreqNet ( Tan et al., 2024a )
49.80
53.50
51.11
4.00
99.90
NPR ( Tan et al., 2024b )
50.10
46.80
47.21
0.00
100.0
RealForensics ( Haliassos et al., 2022 )
54.05
67.08
68.81
81.23
49.31
Table 5 : Cross-manipulation generalization on DeepFake (V) . Generalization performance on CelebDF ( Li et al., 2020 ) .
Figure 6 : Robustness against various unseen corruptions . See Appendix D for detailed analysis and intensity settings.
Figure 7 : Analysis of temporal parameters . Left: Detection accuracy as a function of sliding window length T . Right: Detection accuracy by video length, comparing Random and Non-random sampling strategies.
I. Component Removal
II. Backbone Replacement
Setting
ACC (%)
Δ (%)
Setting
ACC (%)
Δ (%)
Full model
97.90
—
Pose(Pooling)
97.20
↓0.7
w/o Align
80.56
↓17.3
Pose(Flatten)
95.27
↓2.6
w/o Pose
74.78
↓23.1
Lip(MobileNet)
96.23
↓1.7
w/o Lip
53.50
↓44.4
Lip(CLIP-ViT)
85.73
↓12.0
Table 6 : Ablation study of our framework’s components and backbone choices. We report results on AVLips, and Δ denotes the ACC (%) drop from our full model.
Figure 8 : t-SNE visualization of Stage I lip and pose embeddings from unseen real (Left) and fake (Right) test samples .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Generator
Category
Paradigm
Training Data
X2Face ( Wiles et al., 2018 )
Geometric
Video-driven
VoxCeleb
TPSM ( Zhao and Zhang, 2022 )
Landmark
Video-driven
VoxCeleb
LIA ( Wang et al., 2022 )
Latent Space
Video-driven
VoxCeleb
FaceVid ( Wang et al., 2021 )
Landmark
Video-driven
VoxCeleb
DiNet ( Zhang et al., 2023b )
CNN
Audio-driven
HDTF, MEAD
MakeItTalk ( Zhou et al., 2020 )
GAN
Audio-driven
VoxCeleb, ObamaSet
Appendix
Table 7: LipSync-A Generator Details. Category refers to the core technical architecture. Paradigm indicates the driving signal. Original Training data lists the datasets used to train the original generator models.
Figure 9 : The generation pipeline for our LipSync-A dataset .
Type
Parameter
L1
L2
L3
L4
L5
Block-wise
block number
16
32
48
64
80
Color Contrast
contrast factor
0.85
0.725
0.6
0.475
0.35
Color Saturation
saturation gain
0.4
0.3
0.2
0.1
0.0
Gaussian Blur
kernel size
7
9
13
17
21
Gaussian Noise
variance σ2
0.001
0.002
0.005
0.01
0.05
JPEG Compression
quality drop
30
32
35
38
40
Appendix
Table 8 : Robustness experiment parameters . Each perturbation method employs five unique sets of hyperparameter values, modifying them solely during the video preprocessing phase.
Perturbation
FTCN
LipFD
LipForensics
RealForensic
Ours
Block-wise
68.70
96.15
93.50
62.50
99.59
Contrast
75.56
80.77
87.00
56.25
99.27
Saturation
69.41
91.28
93.25
64.58
97.71
Gaussian Blur
57.77
51.79
82.00
50.00
99.54
Gaussian Noise
54.20
51.54
71.00
45.83
98.50
Compression
64.57
90.35
80.00
58.33
99.59
Appendix
Table 9 : Robustness comparison against Level-3 perturbations . We report AUC (%) on the AVLips test set. Bold indicates the best performance, and underline marks the second-best.
Figure 10 : Visualization of the seven perturbation types at Level-3 intensity .
Perturbation
Accuracy
F1-score
AUC
noise_heavy
86.25
85.82
98.79
noise_light
70.40
63.50
97.97
pitch_shift_down3
97.11
97.36
99.33
pitch_shift_up3
79.60
77.57
97.31
resample_16k
96.94
97.23
99.37
resample_44k
97.37
97.59
99.56
Appendix
Table 10 : Performance of our model under various audio corruptions . Bold and underline denote the best and second-best results per metric (Accuracy, F1-score, AUC), respectively. All metrics are in %.
Metric
Baseline
JPEG
BW
CC
PXL
GB
CS
GN
FR ↓
6.0
6.3
6.6
8.8
11.4
11.7
13.9
34.7
AUC ↑
99.72
99.60
99.69
96.99
98.55
98.33
92.10
90.98
Appendix
Table 11 : Landmark extraction robustness under Level-5 visual perturbations. FR ( % ) is the per-frame landmark extraction failure rate; AUC ( % ) is reported on the AVLips test set. Columns are ordered by increasing FR.
Generator
DiNet
DreamTalk
FaceVid
IP-LAP
LIA
MakeItTalk
SadTalker
TPSM
Wav2Lip
X2Face
Avg.
ACC / AUC (%)
94.5/98.0
95.4/99.9
95.5/99.9
94.4/97.1
95.6/99.9
94.8/96.8
88.9/84.2
96.9/99.9
95.0/98.0
95.3/100.0
94.6/97.4
Appendix
Table 12 : Per-generator detection performance of the visual-only LipDA on the LipSync-A test split. Each cell reports video-level binary classification accuracy and AUC in ACC/AUC (%) format, both computed against the shared pool of 346 authentic samples. Generators are listed in alphabetical order.
Family
Generator
Prec.
Recall
F1
Hybrid Transformations
X2Face
72.2
86.7
78.8
Keypoint-Driven
TPSM
71.4
33.3
45.5
LIA
71.9
76.7
74.2
FaceVid
46.7
46.7
46.7
Statistical
DiNet
96.4
90.0
93.1
GAN-based
MakeItTalk
100.0
96.7
98.3
Appendix
Table 13 : Fine-grained 15-way model-level attribution performance on the LipSync-A test split. Generators are grouped by their architectural family for ease of comparison. All metrics are reported in %.
Generator
OmniSync
Wan-S2V
Sonic
KDTalker
InfiniteTalk
ACC (%)
95.5
88.2
87.8
84.6
78.8
Appendix
Table 14 : Zero-shot detection accuracy ( % ) of LipDA on unseen SOTA generators, including the full-video generator Wan-S2V. All models are evaluated without any fine-tuning on LipSync-A.
Figure 11 : Representative failure cases of LipDA on the LipSync-A test split. All three patterns are also failed by the baselines evaluated in Table 2 , indicating a shared operational limit rather than a weakness specific to our framework.
Model Variant
ACC
F1
Full Model
97.50%
93.90%
temporal_only(TDM)
74.78%
73.58%
av_sync_only(MSM)
92.92%
93.64%
audio_only
90.71%
91.36%
Appendix
Table 15 : Ablation study on attribution model components . We report the average Accuracy (ACC, %) and F1-Score (F1, %) across five generator families on the LipSync-A.
Figure 12 : Spatial attention maps for the dual-branch encoders . Brighter colors indicate higher attentional weight.
Figure 13 : Comparison of forgeries generated from the same source identity and driving video.
Figure 14 : Samples from each of the five audio-driven generator families in LipSync-A .
Existing lip-sync deepfake detectors rely on pixel artifacts or audio-visual correspondence, and both fail under generator or language shift because the features they learn are tied to the training distribution. We take a different approach. Authentic lip motion is constrained by tissue mechanics and neuromuscular bandwidth; current generators typically do not impose these constraints, producing trajectories with elevated variance in velocity, acceleration, and jerk that real speech does not exhibit. We exploit this signal, which we term temporal lip jitter, by computing kinematic statistics from 64 perioral landmarks over short sliding windows and feeding them into a lightweight three-branch network. The model uses only landmark coordinates: no pixels, no audio, and no voiceprint data. We train only on English data and test in a zero-shot setting on five unseen generators and seven languages.
Lip-syncing deepfakes are among the most challenging forms of manipulated media because their artifacts are localized almost exclusively to the mouth region and evolve dynamically over time. Detecting such deepfakes requires precise temporal and spatial modeling of lip motion. In this paper, we propose LoCC, a novel detection framework that performs fine-grained detection and localization of lip-syncing deepfakes at both segment and frame levels. Unlike prior approaches that analyze videos holistically, our method evaluates whether each frame aligns with a counterfactual estimate generated from its temporal neighbors. Real videos exhibit strong and stable consistency, whereas lip-sync deepfakes introduce localized inconsistencies. Following a teacher-student learning paradigm, our model effectively captures these frame-level discrepancies and achieves superior performance over state-of-the-art methods on multiple benchmark lip-syncing deepfake datasets, including LAV-DF, AVDF1M, FakeAVCeleb, and KODF, and generalizes well across compression levels and datasets.
Soumyya Kanti Datta, Shan Jia, Siwei Lyu
University at Buffalo, State University of New York
We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile image quality with synchronization accuracy, producing either visually degraded outputs or temporally inconsistent lip movements. HighSync addresses both challenges simultaneously and, to our knowledge, is the first lip sync model to operate natively at 512*512 resolution, positioning it as a viable solution for professional production environments such as the film and broadcast industries. Central to our approach is the identification and systematic elimination of a data leakage phenomenon that has silently undermined temporal modeling in prior work, preventing models from developing a genuine dependence on the audio signal. Comprehensive evaluations across both perceptual quality and synchronization accuracy metrics confirm that HighSync achieves state-of-the-art performance on both fronts. Source code, pre-trained models, and supplementary video results are publicly available at: https://github.com/saeed5959/high_sync