Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.
Figures & tables
1
Sample data (x1,c) , noise ϵ , and 0<Δ≤hk .
2
Sample t∈[0,1−Δ] ; set s=t+Δ and xt=(1−t)ϵ+tx1 .
3
Prediction (gradients enabled):
4
u←uθ(xt,t,s,c) .
5
Target (reverse-mode off; forward-mode on):
6
v←vT(xt,t,c) .
7
Fix g(x,τ)=uθ(x,τ,s,c) .
Algorithm 1 One LFMD training update
F5-TTS teachers
32 NFEs
Training
Size
Evaluation
WER ↓
SIM ↑
UTMOS ↑
Emilia
Small
Seed-TTS
1.88
.625
3.82
Emilia
Base
Seed-TTS
1.76
.667
3.76
Emilia
Medium
Seed-TTS
1.92
.682
3.72
LibriTTS
Small
LibriSpeech-PC
2.23
.582
4.07
1 NFE
2 NFEs
3 NFEs
4 NFEs
Table 1: Teacher and few-step student results. WER is in %, SIM denotes SIM-o, and bold marks the best student score for each dataset among models matched in training data, size, and NFE (ties included). DTW denotes temporal self-distillation.
Few-step diffusion and flow-matching text-to-speech (TTS) models are usually trained with local objectives, such as conditional flow matching, reconstruction, and stop prediction. These losses provide stable optimization, but they never ask whether sampled speech follows the distribution of high-quality speech. We propose Speech Representation Fr'echet Distance loss (SR-FD), a training-time distributional regularizer for tokenizer-free flow-matching autoregressive TTS. During fine-tuning, the model synthesizes speech with the same few-step sampler used at deployment, and SR-FD matches the mean and covariance of frozen Whisper and CTC features of this speech to reference statistics computed offline from three complementary content targets. The loss requires no discriminator and no inference-time computation. On Seed-TTS English, four-step SR-FD fine-tuning reduces WER from the original four-step VoxCPM2 baseline's 2.2279% to 1.4147%, a 36.5% relative reduction, and also surpasses the original ten-step baseline at 1.7366%; both gains are significant under an utterance-level paired bootstrap. Speaker similarity and objective quality proxies are preserved at the ten-step level, and an error analysis shows the gain comes from content substitutions across all prompt lengths. SR-FD is thus an intelligibility-improving distributional regularizer for few-step TTS.
Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu +1
Graduate Institute of Communication Engineering, National Taiwan University, Taipei, Taiwan · Graduate Institute of Computer Science and Information Engineering, National Taiwan University, Taipei, Taiwan · Amazon
Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to the inherently uncoordinated acoustic dynamics of CFM. We demonstrate that this issue can be effectively mitigated by introducing a novel training-free frequency-selective boosting strategy. Using the Discrete Wavelet Transform (DWT), our method dynamically modulates mel-spectrogram sub-bands during ODE integration, synchronizing spectral development by penalizing aggressive low-frequency growth and boosting lagging high-frequency details. Validated across diverse architectures (Matcha-TTS, F5-TTS, IndicF5), our approach reduces the required Number of Function Evaluations (NFE) from 32 to 26 and improves Frechet Audio Distance (FAD) by up to 61%, all without compromising mean opinion scores, speaker similarity, and speech intelligibility.
Isha Pandey, Varad Deshpande, Abhijat Bharadwaj +1
While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation quality still lags significantly behind multi-step counterparts. We propose FdAudio to bridge this gap. Unlike MeanAudio, which relies solely on regression against target velocity fields, our post-training approach optimizes the final one-step distribution directly across pre-trained embedding spaces via a multi-representation Fréchet-distance (FD) loss. Crucially, to prevent the multi-step degradation that naive post-training with FD-loss causes, we introduce a MeanFlow consistency objective as a structural anchor. Results demonstrate that FdAudio establishes state-of-the-art one-step T2A generation quality among few-step systems, yielding an 11.4% reduction in FD score and a 28.8% improvement in FAD score relative to the baseline MeanAudio framework. Notably, we solve FD post-training's naive multi-step degradation issue by proposing the MeanFlow anchor, enabling a 25-step sampling path to maintain high-fidelity audio synthesis that matches or surpasses strong multi-step models at a fraction of their computational latency.