cs.LGSep 28, 2026

Harmonizing Spectral Evolution in Conditional Flow Matching for TTS

Authors: Isha Pandey, Varad Deshpande, Abhijat Bharadwaj, Ganesh Ramakrishnan

Organizations: Indian Institute of Technology Bombay

Abstract

Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to the inherently uncoordinated acoustic dynamics of CFM. We demonstrate that this issue can be effectively mitigated by introducing a novel training-free frequency-selective boosting strategy. Using the Discrete Wavelet Transform (DWT), our method dynamically modulates mel-spectrogram sub-bands during ODE integration, synchronizing spectral development by penalizing aggressive low-frequency growth and boosting lagging high-frequency details. Validated across diverse architectures (Matcha-TTS, F5-TTS, IndicF5), our approach reduces the required Number of Function Evaluations (NFE) from 32 to 26 and improves Frechet Audio Distance (FAD) by up to 61%, all without compromising mean opinion scores, speaker similarity, and speech intelligibility.

Figures & tables

Explore similar work

CardsList
  1. Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech

    Sep 28, 2026Yentl Collin, Evan Dufraisse, Amr Mohamed +3Flow-Matching Text-To-SpeechLatent Flow

  2. Mask, Sample, Revise: A Revisable CTMC Inference Stack for Guided Discrete Flow Matching Text-to-Speech

    Jun 12, 2026Alef Iury Siqueira Ferreira, Lucas Rafael Stefanel Gris, Luiz Fernando de Araújo Vidal +4Flow-Matching Text-To-SpeechText-To-Audio

  3. Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis

    Jul 7, 2026Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu +1Flow-Matching Text-To-SpeechAutoregressive Text-To-Speech