cs.CVOct 1, 2026

Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning

Authors: Gunwoo Lee, Yoori Oh, Yoseob Han

Organizations: Department of Information and Telecommunication Engineering Soongsil University Seoul, Republic of Korea · Graduate School of Data Science Seoul National University Seoul, Republic of Korea · Department of Electronic Engineering Soongsil University Seoul, Republic of Korea

Abstract

Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental challenge in this task is the inherent one-to-many mapping problem, where visual dynamics often lack sufficient information to uniquely determine the corresponding utterance. To address this, we propose Watch Your Speech (WYS), a video-to-speech synthesis framework that incorporates textual conditioning as an explicit linguistic cue to mitigate visual ambiguity. Our framework features an attention-based embedding fusion module that synergistically integrates textual context with video sequences, coupled with a conditional flow matching objective for high-fidelity speech generation. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that WYS achieves superior performance, establishing new state-of-the-art results in audio-visual synchronization (LSE-C/D) while maintaining highly competitive textual accuracy (WER). Subjective evaluations further confirm that our model generates speech with near-human naturalness, validating the effectiveness of textual conditioning in content-controlled video-to-speech synthesis. Project page: https://github.com/gunwoo5034/Watch-your-Speech

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

    Nov 23, 2025Kaidi Wang, Yi He, Wenhao Guan +9Audio-Visual ConsistencySpeech Synthesis

  2. Hierarchical Codec Diffusion for Video-to-Speech Generation

    Apr 17, 2026Jiaxin Ye, Gaoxiang Cong, Chenhui Wang +4Audio-Video GenerationAudio-Visual Consistency

  3. Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

    Oct 3, 2025Kaisi Guan, Xihua Wang, Zhengfeng Lai +5Audio-Video GenerationVideo Generation