FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
Authors: Fabian Ritter-Gutierrez, Md Asif Jalal, Pablo Peso Parada, Karthikeyan Saravanan, Yusun Shul, Minseung Kim, Gun-Woo Lee, Han-Gil Moon
Organizations: Samsung Electronics R&D Institute UK (SRUK), London, United Kingdom · Samsung Electronics, Mobile eXperience Business, Suwon, Republic of Korea
Whispered-to-normal (W2N) speech conversion aims to reconstruct missing phonation from whispered input while preserving content and speaker identity. This task is challenging due to temporal misalignment between whisper and voiced recordings and lack of paired data. We propose FlowW2N, a conditional flow matching approach that trains exclusively on synthetic, time-aligned whisper-normal pairs and conditions on domain-invariant features. We exploit high-level ASR embeddings that exhibits strong invariance between synthetic and real whispered speech, enabling generalization to real whispers despite never observing it during training. We verify this invariance across ASR layers and propose a selection criterion optimizing content informativeness and cross-domain invariance. Our method achieves SOTA intelligibility on the CHAINS and wTIMIT datasets, reducing Word Error Rate by 26-46% relative to prior work while using only 10 steps at inference and requiring no real paired data, validated by a subjective listening study and F0-contour analysis.
Figures & tables
Figure 1: FlowW2N pipeline. Left (Training): The DiT learns a velocity field vθ(zt,t,c) conditioned on domain-invariant content features h from a content encoder (layer ℓ∗ ) and speaker embedding espk , where c={espk,h} represents the conditioning set of content and speaker. Training uses only synthetic whisper-normal pairs. Right (Inference): Starting from Gaussian noise, the ODE is integrated to obtain z1 , which is decoded to normal speech. Domain invariance of content features enables generalization to real whispered speech.
Figure 2: Domain invariance analysis. Synthesis Gap (Syn): synthetic vs. real whisper; Modality Gap (Mod): real whisper vs. normal.
Figure 3: Layer selection analysis. Left: Synthesis Gap (invariance). Center: CCA with word identity. Right: Proposed combined score (invariance × semantic). Stars mark optimal layers selected.
CHAINS
wTIMIT
Method
Training Data
WER-N% ↓
WER-W% ↓
SpkSim ↑
DNSMOS ↑
WER-N% ↓
WER-W% ↓
SpkSim ↑
DNSMOS ↑
Reference
Real Whisper
–
11.1
31.6
–
1.50
8.3
26.5
–
1.21
Normal Speech
–
5.9
11.9
–
3.06
3.9
9.5
–
3.21
Prior Work
QuickVC
Real pairs
24.1
36.8
0.668
2.99
26.0
39.1
0.720
3.04
Table 1: FlowW2N versus prior methods. WER using NeMo (N) and Whisper tiny (W). Green : relative improvement over best baseline.
Model
Input
WER% ↓
UTMOS ↑
DNSMOS ↑
SpkSim ↑
Reference
Real Whisper
–
11.1
1.42
1.50
–
Normal Speech
–
5.9
3.83
3.06
–
Uncond. (Synth. Trained)
Paired-Base
Real (OOD)
28.2
1.47
1.43
0.522
Paired-Base
Synth (ID)
10.1
3.46
3.09
0.832
Table 2: Paired flow matching fails to generalize on real whisper (OOD). Finetuning on misaligned real pairs degrades further.
System
MOS ↑
logF0 RMSE ↓
F0 corr ↑
Reference
4.99
–
–
Real Whisper
3.10
11.91
0.04
DistillW2N
1.63
5.38
0.12
QuickVC
3.24
5.65
0.24
FlowW2N
3.37
5.88
0.27
Table 3: MOS naturalness (25 listeners) and wTIMIT F0 reconstruction vs. the parallel normal recording.
WER% ↓
Configuration
N
W
UTMOS ↑
SpkSim ↑
DNSMOS ↑
Reference
Real Whisper
11.1
31.6
1.42
–
1.50
Normal Speech
5.9
11.9
3.83
–
3.06
Conditioning Signal (no speaker emb.)
C-VAE-p
23.7
–
1.65
0.538
1.64
Table 4: Ablation study on CHAINS (real whisper input). All models trained on synthetic data only. WER reported for NeMo (N) and Whisper tiny (W).