Lip synchronization aims to generate visual lip dynamics that align precisely with speech audio. Despite the high generation quality of diffusion models, they often struggle in complex scenarios and suffer from prohibitive inference latency, limiting real-world deployment. We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions. First, we introduce a dual-stream joint training strategy to mitigate information leakage from reference frames while preserving natural dynamics. Second, we develop a distillation-based acceleration scheme for single-step denoising, achieving a throughput of over 70 FPS. Third, we propose a relational alignment loss that leverages structural priors from Vision Foundation Models (VFMs) to enhance robustness against complex scene factors. Furthermore, we present the first benchmark specifically designed for complex lip synchronization, comprising over 200 challenging video sequences and specialized metrics. Extensive experiments demonstrate that ComplexSync achieves state-of-the-art performance across both standard and complex scenarios while enabling real-time inference.
Figures & tables
Figure 1 : We propose ComplexSync, a unified diffusion-based framework for real-time, high-fidelity lip synchronization in challenging scenarios. Extensive experiments show that ComplexSync achieves SOTA performance across both standard and challenging scenarios while supporting real-time inference.
Figure 2 : Overview of ComplexSync. (a) Dual-stream joint training strategy designed to suppress spatial information leakage and ensure high-fidelity synthesis. (b) Distillation-based acceleration scheme that facilitates single-step denoising for real-time inference. (c) Relational alignment loss leveraging VFMs to enhance synchronization precision and robustness under unconstrained conditions.
Methods
HDTF
ComplexSync-Val
FPS ↑
FID ↓
FVD ↓
CSIM ↑
Sync-C ↑
Sync-D ↓
LipLeak ↓
Sync-C ↑
Sync-D ↓
VFM-FD ↓
Occ-RC ↑
DINet
10.14
59.72
0.835
4.62
9.17
0.0596
3.52
7.69
0.086
87.70
57.74
Musetalk
7.26
40.90
0.909
4.36
9.93
0.1822
2.83
9.60
0.074
78.80
20.16
LatentSync
7.37
54.27
0.900
4.99
9.50
0.0685
6.73
7.37
0.064
93.10
2.13
KeySync
17.90
40.04
0.801
6.41
7.98
0.0708
1.67
9.81
0.075
87.20
1.14
Ditto
8.43
37.55
0.916
5.52
8.90
0.6048
4.26
9.61
0.075
83.80
16.6
Table 1 : Quantitative comparisons with existing methods on two test datasets. The best results are in bold , and the second-best are in underlined .
Figure 3 : Qualitative comparison with existing methods. Refer to Appendix for details. We provide supplementary videos to capture temporal attributes (synchronization, naturalness, stability) missing in static images.
Methods
Lip-Sync ↑
Video Definition ↑
Naturalness ↑
VA ↑
DINet
1.750
1.679
1.696
1.732
Musetalk
2.089
1.982
2.571
2.339
LatentSync
2.446
2.786
2.875
2.857
KeySync
3.393
3.196
2.696
2.857
Ditto
3.214
3.429
3.375
3.429
OmniSync
3.554
3.768
3.143
3.232
Table 2 : User Study results. The best results are in bold , and the second-best are in underlined .
Datasets
Methods
FID ↓
FVD ↓
Sync-C ↑
Sync-D ↓
LipLeak ↓
Occ-RC ↑
HDTF
w/o Dual-Stream
7.92
40.73
4.95
9.35
0.0246
-
w/o weight fusion (Sync Stream)
10.05
59.94
8.18
6.71
0.0110
-
w/o weight fusion (Fidelity Stream)
8.04
119.48
1.18
13.83
0.3116
-
w/ Dual-Stream
7.89
36.73
6.95
7.57
0.0153
-
w/o EMA anchor
8.69
53.38
6.08
8.52
0.0698
-
DMD2 distillation
8.62
39.45
5.42
9.03
0.1144
-
Table 3 : Quantitative ablation study. Rows 1-4 evaluate dual-stream strategy (w/o Stages 2-3). Rows 5-7, starting from dual-stream, compare distillation schemes to verify our EMA anchor. Final three rows validate RA loss and Gram operation.
Figure 4 : Ablation Study for the Dual-stream Joint Training Strategy.
Figure 5 : Ablation Study for the Proposed Distillation Strategy.
Figure 6 : Ablation Study for the Relational Alignment Loss.
Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many denoising steps make them impractical for real-time inference. We present Lip Forcing, to our knowledge the first autoregressive diffusion method for video-to-video (V2V) lip synchronization, which distills a 14B audio-conditioned bidirectional video diffusion teacher into causal students. At inference, the students generate each chunk in only two denoising steps without inference-time CFG, enabling real-time lip synchronization. A lip-sync-specific teacher-trajectory analysis reveals a CFG fidelity-sync tradeoff: no-CFG predictions favor reference fidelity, whereas CFG-guided predictions favor synchronization within a mid-trajectory band. Lip Forcing translates this finding into three analysis-derived components: Sync-Window DMD, a two-step inference schedule, and a SyncNet-based reward. We validate Lip Forcing at two student scales, both distilled from the 14B teacher. The 1.3B student crosses into real-time streaming at 31 FPS, 17.6× faster than its same-scale bidirectional model. The 14B student, the largest diffusion model reported for V2V lip synchronization, runs 39.8× faster than its teacher at comparable reference fidelity. Time-to-first-frame is sub-millisecond at both scales, far below every diffusion baseline.
We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile image quality with synchronization accuracy, producing either visually degraded outputs or temporally inconsistent lip movements. HighSync addresses both challenges simultaneously and, to our knowledge, is the first lip sync model to operate natively at 512*512 resolution, positioning it as a viable solution for professional production environments such as the film and broadcast industries. Central to our approach is the identification and systematic elimination of a data leakage phenomenon that has silently undermined temporal modeling in prior work, preventing models from developing a genuine dependence on the audio signal. Comprehensive evaluations across both perceptual quality and synchronization accuracy metrics confirm that HighSync achieves state-of-the-art performance on both fronts. Source code, pre-trained models, and supplementary video results are publicly available at: https://github.com/saeed5959/high_sync
Lip synchronization refers to the task of modifying facial lip movements such that they are temporally aligned with a given audio signal. It is a fundamental problem in audio-visual synthesis for generating realistic talking-head videos. Recent progress in diffusion-based generative models has led to remarkable advances in lip synchronization. Nevertheless, existing methods typically achieve audio-visual alignment by fine-tuning pre-trained diffusion models on large scale datasets, resulting in substantial computational costs and significant data requirements. Inspired by FlowEdit, we reformulate lip synchronization as a video editing problem. Building upon a pre-trained audio-driven diffusion model, our approach achieves lip synchronization in a training-free manner, without additional fine-tuning or paired data. In this paper, we present SyncEdit, a training-free framework designed for lip synchronization. We reformulate the editing paradigm by substituting the edit sequence in FlowEdit with the target sequence, yielding an unbiased estimation of the desired output. Moreover, we propose annealed noise alignment, which progressively alignes the sampled Gaussian noise with diffusion-model-estimated noise during iterative editing, producing a smooth and stable editing trajectory. Extensive experimental results validate the effectiveness and robustness of the proposed framework. Code is available at \href[]{https://github.com/l1346792580123/SyncEdit}{here}.
Lixiang Lin, Siyuan Jin, Jinshan Zhang
HiThink Research · University of Science and Technology of China · Zhejiang University