Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning
Authors: Gunwoo Lee, Yoori Oh, Yoseob Han
Organizations: Department of Information and Telecommunication Engineering Soongsil University Seoul, Republic of Korea · Graduate School of Data Science Seoul National University Seoul, Republic of Korea · Department of Electronic Engineering Soongsil University Seoul, Republic of Korea
Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental challenge in this task is the inherent one-to-many mapping problem, where visual dynamics often lack sufficient information to uniquely determine the corresponding utterance. To address this, we propose Watch Your Speech (WYS), a video-to-speech synthesis framework that incorporates textual conditioning as an explicit linguistic cue to mitigate visual ambiguity. Our framework features an attention-based embedding fusion module that synergistically integrates textual context with video sequences, coupled with a conditional flow matching objective for high-fidelity speech generation. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that WYS achieves superior performance, establishing new state-of-the-art results in audio-visual synchronization (LSE-C/D) while maintaining highly competitive textual accuracy (WER). Subjective evaluations further confirm that our model generates speech with near-human naturalness, validating the effectiveness of textual conditioning in content-controlled video-to-speech synthesis. Project page: https://github.com/gunwoo5034/Watch-your-Speech
Figures & tables
Figure 1 : The one-to-many mapping problem: visually similar lip movements for bilabial phonemes (/b/, /m/, /p/) lead to ambiguous speech predictions.
Figure 2 : Overview of the proposed WYS framework: (a) training front-end using GT text T , (b) inference front-end predicting text T^ via a lip-reading model, and (c) shared back-end that encodes and fuses text, video, and speaker identity features to guide flow-based speech synthesis.
Figure 3 : Details of the back-end components: (a) Text Encoder, (b) Image Encoder for speaker identity, (c) Video Encoder for lip movements, (d) embedding fusion module that aligns video and text features via attention and injects speaker identity to produce efus , and (e) legend.
Method
LRS2-BBC
LRS3-TED
WER [%] ↓
LSE-C ↑
LSE-D ↓
STOI-Net ↑
DNSMOS ↑
WER [%] ↓
LSE-C ↑
LSE-D ↓
STOI-Net ↑
DNSMOS ↑
Ground Truth
1.4773
6.9788
7.1891
0.9063
3.1384
1.0096
7.3393
6.9002
0.9307
3.2988
VCA-GAN
100.7812
2.3960
11.7631
0.5107
2.2573
90.6827
4.5511
9.1256
0.6325
2.2672
SVTS
–
–
–
–
–
75.6552
6.2841
7.9446
0.7070
2.4203
DiffV2S
50.7801
6.4320
7.6540
0.8923
2.9230
38.0897
6.2841
7.9403
0.9214
3.2169
IntelligibleL2S
44.2286
7.1280
7.0210
0.8588
2.7063
50.0105
6.7813
7.3754
0.8838
2.8678
Table 1 : Quantitative results on LRS2-BBC and LRS3-TED. Best and second-best scores are highlighted. ↓ / ↑ denote lower/higher is better.
Method
LRS2-BBC
LRS3-TED
Qual. ↑
Align. ↑
Intel. ↑
Sync. ↑
Natural. ↑
Qual. ↑
Align. ↑
Intel. ↑
Sync. ↑
Natural. ↑
Ground Truth
4.26 ± 0.15
4.36 ± 0.11
4.31 ± 0.13
4.26 ± 0.14
4.27 ± 0.10
4.25 ± 0.12
4.33 ± 0.10
4.23 ± 0.11
4.23 ± 0.11
4.28 ± 0.01
IntelligibleL2S
2.77 ± 0.23
3.06 ± 0.25
3.00 ± 0.24
3.36 ± 0.22
2.87 ± 0.24
2.59 ± 0.23
3.02 ± 0.26
2.85 ± 0.24
3.08 ± 0.23
2.82 ± 0.24
DiffV2S
3.24 ± 0.21
3.13 ± 0.24
3.24 ± 0.23
3.51 ± 0.19
3.27 ± 0.21
3.56 ± 0.20
3.55 ± 0.23
3.48 ± 0.23
3.70 ± 0.20
3.64 ± 0.20
LipVoicer
3.49 ± 0.21
3.83 ± 0.16
3.64 ± 0.17
3.73 ± 0.18
3.56 ± 0.19
3.71 ± 0.21
3.87 ± 0.19
3.77 ± 0.21
3.90 ± 0.18
3.65 ± 0.20
V2SFlow-V
3.73 ± 0.19
3.42 ± 0.24
3.69 ± 0.18
3.83 ± 0.19
3.85 ± 0.19
3.86 ± 0.22
3.77 ± 0.22
3.78 ± 0.22
3.83 ± 0.23
3.77 ± 0.20
Table 2 : Subjective evaluation (MOS) on LRS2-BBC and LRS3-TED. Best and second-best scores are highlighted. ↑ denotes higher is better.
Figure 4 : Qualitative comparison of mel-spectrograms on LRS2. The text below each spectrogram is predicted by an ASR model, where GT text is shown in blue and incorrectly predicted words are highlighted in red.
Attention Mechanisms
Textual Acc.
A-V Sync.
Audio Qual.
WER [%] ↓
LSE-C ↑
LSE-D ↓
STOI-Net ↑
DNSMOS ↑
etxt∥evid
50.6510
7.6259
6.7147
0.9093
3.0811
etxtself∥evidself
52.3977
7.6127
6.7217
0.9128
3.0581
etxt→vidcross
80.8254
5.5675
8.6043
0.8982
3.0389
evid→txtcross
21.0629
7.5863
6.7593
0.9101
3.0888
etxt→vidcross∥evid→txtcross
20.8882
7.5219
6.7984
0.9113
3.0902
Table 3 : Ablation study on the attention mechanism (LRS2). emself : Self-attention embedding; em→wcross : Cross-attention embedding w/ m as query; esyn : Our fused embedding.
Configuration
WER [%] ↓
LSE-C ↑
LSE-D ↓
STOI-Net ↑
DNSMOS ↑
spkSIM ↑
Ours ( eimg=0 )
20.5793
7.5023
6.8248
0.9130
3.0783
0.6198
Ours (WYS)
20.1321
7.5032
6.8118
0.9144
3.0927
0.6823
Relative Change [%]
2.2213
0.0119
0.0498
0.1533
0.4677
10.0839
Table 4 : Ablation study on the speaker identity module (LRS2).
ω
Textual Acc.
A-V Sync.
Audio Qual.
WER [%] ↓
LSE-C ↑
LSE-D ↓
STOI-Net ↑
DNSMOS ↑
-1
103.8737
2.3689
11.7148
0.8890
2.8528
0.0
28.0110
6.4583
7.6201
0.8946
2.9162
1.0
21.0289
7.4421
6.8604
0.9111
3.0711
2.0
20.1321
7.5032
6.8118
0.9144
3.0927
3.0
20.5699
7.4031
6.8713
0.9120
3.0720
Table 5 : Ablation study on the guidance scale factor ω (LRS2).
Steps
Textual Acc.
A-V Sync.
Audio Qual.
Sampling Time
WER [%] ↓
LSE-C ↑
LSE-D ↓
STOI-Net ↑
DNSMOS ↑
(sec/sample)
10
20.5293
7.5808
6.7572
0.9130
3.0783
0.7
100
20.1321
7.5032
6.8118
0.9144
3.0927
8.8
1000
20.6555
7.4850
6.8194
0.9127
3.0468
84.2
Table 6 : Ablation study on the number of sampling steps (LRS2).
Method
Textual Acc.
A-V Sync.
Audio Qual.
WER [%] ↓
LSE-C ↑
LSE-D ↓
STOI-Net ↑
DNSMOS ↑
(a) LRS2-BBC: inference-text source
Ours (w/ GT Text)
9.9784
7.4889
6.8198
0.9136
3.0901
Ours (w/ Pred. Text)
20.1321
7.5032
6.8118
0.9144
3.0927
(b) LRS2-BBC: training-text source
Ours (w/ Pred. Text)
40.8559
7.7031
6.6299
0.9173
3.0908
Table 7: Results of additional experiments. (a) inference-text source on LRS2, (b) training-text source on LRS2, (c) Lipreading+TTS on LRS2, (d) cross-dataset LRS3 → LRS2.
WER [%]
Ratio
Metric
Text-free Models
Text-guided Models
(in Lip-reading)
IntelligibleL2S
DiffV2S
V2SFlow-V
LipVoicer
Ours (WYS)
= 0
0.61
WER [%] ↓
25.1919
40.6489
26.8763
4.2358
6.0624
LSE-C ↑
7.1005
6.7143
7.2917
5.9842
7.5834
DNSMOS ↑
2.6625
3.0750
3.0750
3.0529
3.0948
(0, 50]
0.26
WER [%] ↓
40.5635
58.6261
40.9356
26.4752
28.7612
LSE-C ↑
7.0715
6.3118
7.2191
5.8798
7.5483
Table 8 : Robustness evaluation under imperfect text conditions (LRS2). Additional metrics (LSE-D, STOI-Net) are provided in the supplementary material.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Learning Rate
Optimizer
Sampling Rate
Vocoder
GPUs
Params
VCA-GAN [ Kim et al.(2021)Kim, Hong, and Ro ]
1×10−4
Adam
16kHz
Griffin-Lim
Single
50.06M
SVTS [ Mira et al.(2022)Mira, Haliassos, Petridis, Schuller, and Pantic ]
1×10−3
Adam
16kHz
WaveGAN
Single
87.63M
DiffV2S [ Choi et al.(2023a)Choi, Hong, and Ro ]
1×10−4
AdamW
16kHz
HiFi-GAN
Single
37.54M
IntelligibleL2S [ Choi et al.(2023b)Choi, Kim, and Ro ]
1×10−3
Adam
16kHz
End - to - End
Single
143.79M
LipVoicer [ Yemini et al.(2024)Yemini, Shamsian, Bracha, Gannot, and Fetaya ]
2×10−4
Adam
16kHz
HiFi-GAN
4
57.24M
V2SFlow [ Choi et al.(2025a)Choi, Kim, Li, Chung, and Liu ]
1×10−3
AdamW
16kHz
HiFi-GAN
8
264.97M
Appendix
Table S1 : Implementation details of the proposed WYS and compared baseline methods.
Models
IntelligibleL2S
DiffV2S
V2SFlow-V
LipVoicer
Ours (WYS)
spkSIM ↑
0.7089
0.5822
0.5732
0.5713
0.6823
Appendix
Table S2 : Speaker similarity evaluation on LRS2.
Figure S1 : Illustration of video–text synchronization using word-level timestamps. During training, text segments aligned with the video frames are selected with a small temporal margin. During inference, the full transcript is used as input.
Models
Auto-AVSR
AV-Hubert
LRS2 / LRS3
14.60 / 19.10
23.82 / 25.51
Appendix
Table S3 : WER results for lip-reading backbone benchmarks.
WER [%]
Ratio
Metric
Text-free Models
Text-guided Models
(in Lip-reading)
IntelligibleL2S
DiffV2S
V2SFlow-V
LipVoicer
Ours (WYS)
= 0
0.61
WER [%] ↓
25.1919
40.6489
26.8763
4.2358
6.0624
LSE-C ↑
7.1005
6.7143
7.2917
5.9842
7.5834
LSE-D ↓
7.0827
7.5199
7.1889
8.2879
6.7649
STOI-Net ↑
0.8594
0.8931
0.9210
0.9048
0.9137
DNSMOS ↑
2.6625
3.0750
3.0750
3.0529
3.0948
Appendix
Table S4 : Robustness evaluation under imperfect text conditions (LRS2).
Figure S2 : Qualitative comparison of mel spectrograms for LRS2 dataset. The scripts below each mel spectrogram represent the ASR-predicted text in black , while the ground-truth (GT) text is shown beneath the GT mel spectrogram in blue . Incorrectly predicted words are highlighted in red .
Figure S3 : Qualitative comparison of mel spectrograms for LRS3 dataset. The scripts below each mel spectrogram represent the ASR-predicted text in black , while the ground-truth (GT) text is shown beneath the GT mel spectrogram in blue . Incorrectly predicted words are highlighted in red .
Figure S4 : Failure cases in mel spectrogram generation with incorrect text predictions on (a) LRS2 and (b) LRS3. The scripts below each mel spectrogram represent the ASR-predicted text in black , while the ground-truth (GT) text is shown beneath the GT mel spectrogram in blue . Incorrectly predicted words are highlighted in red .
Figure 21
Figure S8 : Mel spectrogram comparison on LRS2 using different sampling steps: (i) 10 steps, (ii) 100 steps, (iii) 1000 steps, and (iv) Ground-Truth. The scripts below each mel spectrogram represent the ASR-predicted text in black, while the ground-truth (GT) text is shown beneath the GT mel spectrogram in blue.
Figure S9 : MOS evaluations. Participants evaluated samples from IntelligibleL2S [ Choi et al.(2023b)Choi, Kim, and Ro ] , DiffV2S [ Choi et al.(2023a)Choi, Hong, and Ro ] , LipVoicer [ Yemini et al.(2024)Yemini, Shamsian, Bracha, Gannot, and Fetaya ] , V2SFlow-V [ Choi et al.(2025a)Choi, Kim, Li, Chung, and Liu ] , WYS (Ours), and the ground-truth based on five criteria, with randomized sample.
School of Informatics, Xiamen University, China · MiLM Plus, Xiaomi Inc., China · School of Electronic Science and Engineering, Xiamen University, China +1