cs.CVMar 10, 2026

SyncEdit: Rethinking Lip Synchronization as Editing with Audio-Driven Diffusion Models

Authors: Lixiang Lin, Siyuan Jin, Jinshan Zhang

Organizations: HiThink Research · University of Science and Technology of China · Zhejiang University

Abstract

Lip synchronization refers to the task of modifying facial lip movements such that they are temporally aligned with a given audio signal. It is a fundamental problem in audio-visual synthesis for generating realistic talking-head videos. Recent progress in diffusion-based generative models has led to remarkable advances in lip synchronization. Nevertheless, existing methods typically achieve audio-visual alignment by fine-tuning pre-trained diffusion models on large scale datasets, resulting in substantial computational costs and significant data requirements. Inspired by FlowEdit, we reformulate lip synchronization as a video editing problem. Building upon a pre-trained audio-driven diffusion model, our approach achieves lip synchronization in a training-free manner, without additional fine-tuning or paired data. In this paper, we present SyncEdit, a training-free framework designed for lip synchronization. We reformulate the editing paradigm by substituting the edit sequence in FlowEdit with the target sequence, yielding an unbiased estimation of the desired output. Moreover, we propose annealed noise alignment, which progressively alignes the sampled Gaussian noise with diffusion-model-estimated noise during iterative editing, producing a smooth and stable editing trajectory. Extensive experimental results validate the effectiveness and robustness of the proposed framework. Code is available at \href[]{https://github.com/l1346792580123/SyncEdit}{here}.

Figures & tables

Explore similar work

CardsList
  1. HighSync: High-Quality Lip Synchronization via Latent Diffusion Models

    May 16, 2026Saeid Firouzi Daghigh, Majid Iranpour Mobarakeh, Mostafa Alavi +1Precise Lip SynchronizationSpeech-To-Text Alignment

  2. Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

    Jun 9, 2026Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee +7Precise Lip SynchronizationVideo Diffusion Models

  3. ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios

    Sep 24, 2026Jiaran Cai, Xingpei Ma, Shenneng HuangPrecise Lip SynchronizationAudio-Visual Consistency