Lip synchronization aims to generate visual lip dynamics that align precisely with speech audio. Despite the high generation quality of diffusion models, they often struggle in complex scenarios and suffer from prohibitive inference latency, limiting real-world deployment. We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions. First, we introduce a dual-stream joint training strategy to mitigate information leakage from reference frames while preserving natural dynamics. Second, we develop a distillation-based acceleration scheme for single-step denoising, achieving a throughput of over 70 FPS. Third, we propose a relational alignment loss that leverages structural priors from Vision Foundation Models (VFMs) to enhance robustness against complex scene factors. Furthermore, we present the first benchmark specifically designed for complex lip synchronization, comprising over 200 challenging video sequences and specialized metrics. Extensive experiments demonstrate that ComplexSync achieves state-of-the-art performance across both standard and complex scenarios while enabling real-time inference.
Figures & tables
Figure 1 : We propose ComplexSync, a unified diffusion-based framework for real-time, high-fidelity lip synchronization in challenging scenarios. Extensive experiments show that ComplexSync achieves SOTA performance across both standard and challenging scenarios while supporting real-time inference.
Figure 2 : Overview of ComplexSync. (a) Dual-stream joint training strategy designed to suppress spatial information leakage and ensure high-fidelity synthesis. (b) Distillation-based acceleration scheme that facilitates single-step denoising for real-time inference. (c) Relational alignment loss leveraging VFMs to enhance synchronization precision and robustness under unconstrained conditions.
Methods
HDTF
ComplexSync-Val
FPS ↑
FID ↓
FVD ↓
CSIM ↑
Sync-C ↑
Sync-D ↓
LipLeak ↓
Sync-C ↑
Sync-D ↓
VFM-FD ↓
Occ-RC ↑
DINet
10.14
59.72
0.835
4.62
9.17
0.0596
3.52
7.69
0.086
87.70
57.74
Musetalk
7.26
40.90
0.909
4.36
9.93
0.1822
2.83
9.60
0.074
78.80
20.16
LatentSync
7.37
54.27
0.900
4.99
9.50
0.0685
6.73
7.37
0.064
93.10
2.13
KeySync
17.90
40.04
0.801
6.41
7.98
0.0708
1.67
9.81
0.075
87.20
1.14
Ditto
8.43
37.55
0.916
5.52
8.90
0.6048
4.26
9.61
0.075
83.80
16.6
Table 1 : Quantitative comparisons with existing methods on two test datasets. The best results are in bold , and the second-best are in underlined .
Figure 3 : Qualitative comparison with existing methods. Refer to Appendix for details. We provide supplementary videos to capture temporal attributes (synchronization, naturalness, stability) missing in static images.
Methods
Lip-Sync ↑
Video Definition ↑
Naturalness ↑
VA ↑
DINet
1.750
1.679
1.696
1.732
Musetalk
2.089
1.982
2.571
2.339
LatentSync
2.446
2.786
2.875
2.857
KeySync
3.393
3.196
2.696
2.857
Ditto
3.214
3.429
3.375
3.429
OmniSync
3.554
3.768
3.143
3.232
Table 2 : User Study results. The best results are in bold , and the second-best are in underlined .
Datasets
Methods
FID ↓
FVD ↓
Sync-C ↑
Sync-D ↓
LipLeak ↓
Occ-RC ↑
HDTF
w/o Dual-Stream
7.92
40.73
4.95
9.35
0.0246
-
w/o weight fusion (Sync Stream)
10.05
59.94
8.18
6.71
0.0110
-
w/o weight fusion (Fidelity Stream)
8.04
119.48
1.18
13.83
0.3116
-
w/ Dual-Stream
7.89
36.73
6.95
7.57
0.0153
-
w/o EMA anchor
8.69
53.38
6.08
8.52
0.0698
-
DMD2 distillation
8.62
39.45
5.42
9.03
0.1144
-
Table 3 : Quantitative ablation study. Rows 1-4 evaluate dual-stream strategy (w/o Stages 2-3). Rows 5-7, starting from dual-stream, compare distillation schemes to verify our EMA anchor. Final three rows validate RA loss and Gram operation.
Figure 4 : Ablation Study for the Dual-stream Joint Training Strategy.
Figure 5 : Ablation Study for the Proposed Distillation Strategy.
Figure 6 : Ablation Study for the Relational Alignment Loss.