Lip synchronization refers to the task of modifying facial lip movements such that they are temporally aligned with a given audio signal. It is a fundamental problem in audio-visual synthesis for generating realistic talking-head videos. Recent progress in diffusion-based generative models has led to remarkable advances in lip synchronization. Nevertheless, existing methods typically achieve audio-visual alignment by fine-tuning pre-trained diffusion models on large scale datasets, resulting in substantial computational costs and significant data requirements. Inspired by FlowEdit, we reformulate lip synchronization as a video editing problem. Building upon a pre-trained audio-driven diffusion model, our approach achieves lip synchronization in a training-free manner, without additional fine-tuning or paired data. In this paper, we present SyncEdit, a training-free framework designed for lip synchronization. We reformulate the editing paradigm by substituting the edit sequence in FlowEdit with the target sequence, yielding an unbiased estimation of the desired output. Moreover, we propose annealed noise alignment, which progressively alignes the sampled Gaussian noise with diffusion-model-estimated noise during iterative editing, producing a smooth and stable editing trajectory. Extensive experimental results validate the effectiveness and robustness of the proposed framework. Code is available at \href[]{https://github.com/l1346792580123/SyncEdit}{here}.
Figures & tables
Figure 1 : Comparison of existing lip synchronization methods and our approach . Previous methods generally rely on paired audio–video data, diffusion model fine-tuning, and lip masking, whereas our method achieves lip synchronization using only a pretrained audio-driven diffusion model.
Figure 2 : Comparison of editing results under different tmax values . The left image presents the result obtained with tmax=1 , whereas the right image corresponds to tmax=0.95 . Starting the editing process from pure noise ( tmax=1 ) degrades the fidelity of the edited image
Figure 3 : Overview of SyncEdit . Given a source video and a target audio, SyncEdit performs lip synchronization through a pre-trained audio-driven diffusion model. SyncEdit directly iterates over the target sequence, yielding an unbiased estimate of the desired output. Furthermore, SyncEdit employs an annealed noise alignment strategy that progressively replaces stochastic Gaussian noise with model-estimated noise throughout the denoising process.
Figure 4 : Illustration of the difference between edit sequence and the proposed target-sequence iteration . FlowEdit estimates the edited sample by integrating an edit sequence initialized at tmax , which introduces estimation bias with respect to the desired target. In contrast, the proposed target-sequence formulation directly iterates over the target trajectory, eliminating the bias and producing an unbiased estimate of the final output.
Methods
Full Reference Metrics
No Reference Metrics
Lip Sync
FID ↓
FVD ↓
CSIM ↑
NIQE ↓
BRISQUE ↓
HyperIQA ↑
LMD ↓
LSE-C ↑
Wav2Lip [ 25 ]
14.912
543.340
0.852
6.495
53.372
45.822
10.007
7.630
IP-LAP [ 35 ]
9.512
325.691
0.809
6.533
54.402
50.086
7.695
7.260
Diff2Lip [ 21 ]
12.079
461.341
0.869
6.261
49.361
48.869
18.986
7.140
MuseTalk [ 33 ]
8.759
231.418
0.862
5.824
46.003
55.397
8.701
6.890
LatentSync [ 15 ]
8.518
216.899
0.859
6.270
50.861
53.208
17.344
8.050
Table 1 : Quantitative results on HDTF Dataset.
Methods
Full Reference Metrics
No Reference Metrics
Generation Success Rate
FID ↓
FVD ↓
CSIM ↑
NIQE ↓
BRISQUE ↓
HyperIQA ↑
All Videos ↓
Stylized Characters ↑
Wav2Lip [ 25 ]
22.989
562.245
0.727
5.392
42.816
50.511
71.38%
26.67%
IP-LAP [ 35 ]
14.686
247.402
0.796
5.546
45.153
53.174
45.53%
6.67%
Diff2Lip [ 21 ]
23.542
403.149
0.692
5.440
42.442
50.335
74.63%
36.67%
MuseTalk [ 33 ]
17.668
297.621
0.667
4.935
36.017
58.334
92.20%
67.78%
LatentSync [ 15 ]
15.374
263.111
0.751
5.342
41.917
54.648
74.96%
35.56%
Table 2 : Quantitative results on AIGC-LipSync Benchmark.
Figure 5 : Qualitative comparison with existing methods . Compared with prior approaches, our method generates more accurate lip synchronization and more faithful dental structures while maintaining superior visual consistency. Even in challenging scenarios, including profile views and significant facial occlusions, our method produces natural and plausible lip movements with fewer visual artifacts.
Methods
FID ↓
FVD ↓
CSIM ↑
NIQE ↓
BRISQUE ↓
HyperIQA ↑
LSE-C ↑
Edit sequence
7.944
198.409
0.878
5.408
37.526
55.198
7.284
Random noise
7.707
194.916
0.880
5.396
37.457
55.492
7.286
Ours
7.623
190.299
0.884
5.385
37.412
55.973
7.286
Table 3 : Ablation study for our proposed method.
Figure 6 : Ablation study for edit sequence and random noise. Edit sequence and stochastic noise injection tends to produce blurred dental details, whereas our method is capable of generating sharper and more clearly defined teeth.
Figure 7 : Qualitative results of SyncEdit for Audio-visual Editing . Our approach supports prompt-based manipulation of diverse attributes—including age (a), gender (b), person (c), emotion (d), behaviors (e), and even car categories (f), while jointly generating audio and video in a temporally synchronized and semantically consistent manner.
We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile image quality with synchronization accuracy, producing either visually degraded outputs or temporally inconsistent lip movements. HighSync addresses both challenges simultaneously and, to our knowledge, is the first lip sync model to operate natively at 512*512 resolution, positioning it as a viable solution for professional production environments such as the film and broadcast industries. Central to our approach is the identification and systematic elimination of a data leakage phenomenon that has silently undermined temporal modeling in prior work, preventing models from developing a genuine dependence on the audio signal. Comprehensive evaluations across both perceptual quality and synchronization accuracy metrics confirm that HighSync achieves state-of-the-art performance on both fronts. Source code, pre-trained models, and supplementary video results are publicly available at: https://github.com/saeed5959/high_sync
Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many denoising steps make them impractical for real-time inference. We present Lip Forcing, to our knowledge the first autoregressive diffusion method for video-to-video (V2V) lip synchronization, which distills a 14B audio-conditioned bidirectional video diffusion teacher into causal students. At inference, the students generate each chunk in only two denoising steps without inference-time CFG, enabling real-time lip synchronization. A lip-sync-specific teacher-trajectory analysis reveals a CFG fidelity-sync tradeoff: no-CFG predictions favor reference fidelity, whereas CFG-guided predictions favor synchronization within a mid-trajectory band. Lip Forcing translates this finding into three analysis-derived components: Sync-Window DMD, a two-step inference schedule, and a SyncNet-based reward. We validate Lip Forcing at two student scales, both distilled from the 14B teacher. The 1.3B student crosses into real-time streaming at 31 FPS, 17.6× faster than its same-scale bidirectional model. The 14B student, the largest diffusion model reported for V2V lip synchronization, runs 39.8× faster than its teacher at comparable reference fidelity. Time-to-first-frame is sub-millisecond at both scales, far below every diffusion baseline.
Lip synchronization aims to generate visual lip dynamics that align precisely with speech audio. Despite the high generation quality of diffusion models, they often struggle in complex scenarios and suffer from prohibitive inference latency, limiting real-world deployment. We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions. First, we introduce a dual-stream joint training strategy to mitigate information leakage from reference frames while preserving natural dynamics. Second, we develop a distillation-based acceleration scheme for single-step denoising, achieving a throughput of over 70 FPS. Third, we propose a relational alignment loss that leverages structural priors from Vision Foundation Models (VFMs) to enhance robustness against complex scene factors. Furthermore, we present the first benchmark specifically designed for complex lip synchronization, comprising over 200 challenging video sequences and specialized metrics. Extensive experiments demonstrate that ComplexSync achieves state-of-the-art performance across both standard and complex scenarios while enabling real-time inference.