In flow-matching text-to-speech (TTS), different speech realizations can induce different target velocities under the same generation conditions. A deterministic velocity field trained with squared error predicts their conditional mean, thereby marginalizing realization-dependent variation. Meanwhile, modeling such variation does not inherently provide a semantically interpretable interface for attribute manipulation. We propose ReaFlow-TTS, a realization-conditioned flow-matching framework that introduces an utterance-level stochastic realization latent and uses it to condition velocity prediction throughout the generation trajectory. We further impose valence-arousal-dominance (VAD) semantics on the realization space, enabling direct and graded attribute manipulation without target speech at inference. Experiments demonstrate improved synthesis quality over a matched full-mask baseline and reproducible latent-induced pitch, energy, and timing tendencies across initial-noise samples, providing behavioral evidence that the latent is used as a reusable realization condition. Subjective evaluation further demonstrates graded VAD manipulation across generation contexts with only modest changes in naturalness.
Figures & tables
Figure 1: Illustration of realization-dependent velocity modeling in flow-matching TTS. (a) Different speech realizations can induce different target velocities under the same generation conditions, which are conditionally averaged by a deterministic velocity predictor. (b) ReaFlow-TTS introduces a realization latent z to explicitly condition velocity prediction on realization information.
Figure 2: Overview of ReaFlow-TTS. The realization latent z conditions velocity prediction throughout the flow trajectory, while VAD supervision structures the latent space for graded attribute manipulation. C and + denote concatenation and addition, respectively.
Figure 3: Cross-noise comparison of three realization latents across nine matched initial-noise samples. Corresponding grid positions across the three latent conditions share the same x0 . Shared log-Mel scale: approximately −8 to 0 .
System
Inf. params.
WER (%) ↓
UTMOS ↑
NMOS ↑
Ground truth
—
2.28
4.10
3.93
F5-TTS
157.97M
2.27
3.91
3.72
F5-TTS (full mask)
157.97M
2.32
3.83
3.67
ZipVoice base
122.66M
2.24
3.87
3.63
ReaFlow-TTS (500k)
158.75M
2.08
4.19
3.93
ReaFlow-TTS (full)
158.75M
2.09
4.19
3.91
Table 1: Speech quality and inference parameter counts.
Axis
Attribute score [-1pt] η=0,0.5,1
NMOS ↑ [-1pt] η=0,0.5,1
Ordering acc. (%) ↑ [-1pt] (0,0.5),(0.5,1),(0,1)
V
3.62/3.77/4.21
3.91/3.93/3.83
61/78/82
A
3.38/4.76/5.23
3.91/3.82/3.80
89/81/92
D
3.71/4.98/5.36
3.91/3.84/3.77
86/82/88
Table 2: Perceptual attribute changes and naturalness under graded VAD manipulation.
While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from imperfect alignment. We propose RobustSpeechFlow, a training strategy that improves alignment robustness by extending contrastive flow matching with length-preserving repeat and skip latent augmentations. Requiring no external aligners or preference data, our method directly penalizes realistic failure modes and readily integrates into existing pipelines. On Seed-TTS-eval, it reduces the word error rate (WER) from 1.44 to 1.38 using only 0.06B parameters. On our ZERO500 benchmark, it delivers consistent intelligibility improvements across diverse speaker and prosody conditions; at NFE=24, it reduces English character error rate (CER) from 0.48% to 0.35% and Korean CER from 0.81% to 0.57%. Audio samples: https://robustspeechflow.github.io/
Jinhyeok Yang, Hyeongju Kim, Yechan Yu +3
Supertone Inc, South Korea · Independent Researcher, South Korea
Existing Reinforcement Learning (RL) research for Text-to-Speech (TTS) focuses on large language models (LLMs), leaving Flow-Matching (FM) under-explored. We present FlowTTS-GRPO, an online RL framework for FM-based TTS. By converting ordinary differential equation (ODE) trajectories into stochastic differential equation (SDE) paths, our method enables direct fine-tuning of open-source FM models without auxiliary models. We show that a weighted reward combination converges faster than a probabilistic scheme, and identify three practical optimizations: omitting classifier-free guidance (CFG) during training accelerates convergence; synthesizing hard cases improves robustness; and applying RL to the FM component enhances audio-detail metrics. Experiments on CosyVoice 3.0 and F5-TTS demonstrate objective and subjective preference gains in speaker similarity and perceptual quality, with F5-TTS also improving intelligibility.
Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to the inherently uncoordinated acoustic dynamics of CFM. We demonstrate that this issue can be effectively mitigated by introducing a novel training-free frequency-selective boosting strategy. Using the Discrete Wavelet Transform (DWT), our method dynamically modulates mel-spectrogram sub-bands during ODE integration, synchronizing spectral development by penalizing aggressive low-frequency growth and boosting lagging high-frequency details. Validated across diverse architectures (Matcha-TTS, F5-TTS, IndicF5), our approach reduces the required Number of Function Evaluations (NFE) from 32 to 26 and improves Frechet Audio Distance (FAD) by up to 61%, all without compromising mean opinion scores, speaker similarity, and speech intelligibility.
Isha Pandey, Varad Deshpande, Abhijat Bharadwaj +1