Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech
Organizations: Institute of Foundation Models (IFM), MBZUAI
Abstract
Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.
Figures & tables
| 1 | Sample data , noise , and . |
|---|---|
| 2 | Sample ; set and . |
| 3 | Prediction (gradients enabled): |
| 4 | . |
| 5 | Target (reverse-mode off; forward-mode on): |
| 6 | . |
| 7 | Fix . |
| F5-TTS teachers | 32 NFEs | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Training | Size | Evaluation | WER | SIM | UTMOS | |||||||||
| Emilia | Small | Seed-TTS | 1.88 | .625 | 3.82 | |||||||||
| Emilia | Base | Seed-TTS | 1.76 | .667 | 3.76 | |||||||||
| Emilia | Medium | Seed-TTS | 1.92 | .682 | 3.72 | |||||||||
| LibriTTS | Small | LibriSpeech-PC | 2.23 | .582 | 4.07 | |||||||||
| 1 NFE | 2 NFEs | 3 NFEs | 4 NFEs | |||||||||||