cs.SDSep 24, 2026

ReaFlow-TTS: Realization-Conditioned Flow Matching for High-Quality and Controllable Speech Synthesis

Authors: Junyi Zhao, Yihao Qin, Changsheng Ma

Organizations: School of Information Science and Technology, Lanzhou University, China

Abstract

In flow-matching text-to-speech (TTS), different speech realizations can induce different target velocities under the same generation conditions. A deterministic velocity field trained with squared error predicts their conditional mean, thereby marginalizing realization-dependent variation. Meanwhile, modeling such variation does not inherently provide a semantically interpretable interface for attribute manipulation. We propose ReaFlow-TTS, a realization-conditioned flow-matching framework that introduces an utterance-level stochastic realization latent and uses it to condition velocity prediction throughout the generation trajectory. We further impose valence-arousal-dominance (VAD) semantics on the realization space, enabling direct and graded attribute manipulation without target speech at inference. Experiments demonstrate improved synthesis quality over a matched full-mask baseline and reproducible latent-induced pitch, energy, and timing tendencies across initial-noise samples, providing behavioral evidence that the latent is used as a reusable realization condition. Subjective evaluation further demonstrates graded VAD manipulation across generation contexts with only modest changes in naturalness.

Figures & tables

Explore similar work

CardsList
  1. RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching

    May 21, 2026Jinhyeok Yang, Hyeongju Kim, Yechan Yu +3Flow-Matching Text-To-SpeechContrastive Flow Matching

  2. Harmonizing Spectral Evolution in Conditional Flow Matching for TTS

    Sep 28, 2026Isha Pandey, Varad Deshpande, Abhijat Bharadwaj +1Flow-Matching Text-To-SpeechSpeaker Similarity