cs.SDSep 28, 2026

Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech

Authors: Yentl Collin, Evan Dufraisse, Amr Mohamed, Amine Khelif Khelif, Dani Bouch, Guokan Shang

Organizations: Institute of Foundation Models (IFM), MBZUAI

Abstract

Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.

Figures & tables

Explore similar work

CardsList
  1. Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis

    Jul 7, 2026Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu +1Flow-Matching Text-To-SpeechAutoregressive Text-To-Speech

  2. Harmonizing Spectral Evolution in Conditional Flow Matching for TTS

    Sep 28, 2026Isha Pandey, Varad Deshpande, Abhijat Bharadwaj +1Flow-Matching Text-To-SpeechSpeaker Similarity

  3. FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation

    Jul 11, 2026Kuan-Po Huang, Bo-Ru Lu, Ho-Lam Chung +2Text-To-AudioAudio Understanding