Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-block analysis recovers the rate axis, a neutral operating point, and a per-step intensity scale; at inference, the activation's projection onto this axis is set to a fixed scalar. Learning this direction from synthetically time-stretched and time-compressed speech yields rate control that largely preserves speaker identity, generalizes across model architectures, and maintains high naturalness in objective and human evaluations. Unlike standard additive steering, which breaks at the slow extreme, clamping remains stable on all three systems tested; at moderate targets, the better rule depends on the model. Finally, we show that rate information is decodable across layers but causally steerable only within a mid-depth window, and demonstrate the effectiveness of our approach on the public Seed-TTS-Eval benchmark.
Figures & tables
Fig. 1: Overview. Left: fitting. The rate axis is learned from WSOLA time-stretched speech, the same utterance at slow ( k=−2 ), neutral ( k=0 ), and fast ( k=+2 ) tempo, teacher-forced through the frozen decoder. Right: inference. A single decoder block’s activation is clamped along the learned axis to the requested intensity k , turning the pretrained model into a speaking-rate dial that moves tempo while preserving wording, voice, and naturalness. No fine-tuning.
Rate axis source
Δℓ
UTMOS
spk-sim
Real (rate-binned LibriTTS)
0.31
3.11
0.578
Synthetic (WSOLA)
1.85
3.36
0.600
TABLE I: Rate-axis source on Qwen3-TTS-0.6B (block 14, clamp): a direction learned from real rate-binned LibriTTS vs. from the synthetic WSOLA time-stretch. Δℓ is the per-layer class gap; UTMOS is no-reference naturalness; spk-sim is ECAPA cosine to the cloned voice.
Rule
wps
WER
UTMOS
SIM
Run
Qwen3-TTS
Baseline
3.54
6.6
3.39
0.71
0%
(block 14)
Additive
0.56
86.1
1.98
0.10
92%
Clamp
2.14
8.6
2.76
0.52
0%
MOSS-TTS
Baseline
3.48
13.5
3.12
0.84
0%
(block 18)
Additive
0.67
40.6
2.13
0.53
25%
Clamp
2.05
9.1
2.93
0.79
0%
TABLE II: The slow extreme ( k=−2 ) on all three systems: additive breaks, clamp holds. Baseline is the unsteered output. Run = runaway rate (fraction of generations that never terminate); WER in %; SIM = ECAPA cosine to the cloned reference; Bold = lower WER between the two rules for each model.
Fig. 3: Additive drift unfolding in time on Qwen3-TTS: local speaking rate (words per second in a trailing 4-word window, word timestamps from faster-whisper-small [ 51 ] ) versus word index over a single very long sentence constructed for this analysis, averaged over 10 voices synthesizing the same sentence; bands are ±1 standard deviation. The additive offset compounds over generation, so additive k=+2 keeps accelerating along the utterance while the clamp at the same intensity settles early and holds a steady rate. The additive k=−2 line is interrupted where the model starts failing: beyond that point it does not say the rest of the sentence.
Slow target ( ×0.75 )
Fast target ( ×1.25 )
Model
Rule
k⋆
ach.
WER (%)
UTMOS
SIM
k⋆
ach.
WER (%)
UTMOS
SIM
Qwen3-TTS
Additive
0.62
0.85 †
0.6
3.34
0.61
0.79
1.30
0.7
3.24
0.54
(block 14)
Clamp
1.37
0.74
0.7
3.23
0.58
1.16
1.25
0.7
3.28
0.56
WSOLA
–
0.75
0.8
2.85
0.61
–
1.25
0.7
3.14
0.57
MOSS-TTS
Additive
0.66
0.79
5.0
3.04
0.79
0.99
1.32
11.2
2.83
0.77
(block 18)
Clamp
1.47
0.72
8.5
2.91
0.78
0.66
1.24
5.9
2.95
0.78
TABLE III: Quality at matched moderate rate targets ( ×0.75 slow / ×1.25 fast), two rules per model. k⋆ = the intensity magnitude that reaches the target (slow is negative); ach. = achieved wps ratio vs. baseline; WER/UTMOS/SIM at k⋆ . Bold = lower WER between the two rules for a given model and direction. † target not reached (best-effort plateau), so quality is measured at a milder achieved rate. WSOLA = a time-stretch of the unsteered output to the target rate (DSP baseline); the rate matches the target by construction, so k⋆ is not applicable.
System
Rule
k
wps
× base
WER (%)
SIM
Qwen3-TTS
baseline
—
2.83
1.00
1.4
0.581
additive fast
+0.79
3.45
1.22
1.4
0.514
additive slow
−0.62
2.16
0.77
2.4
0.535
clamp fast
+1.16
3.55
1.26
1.6
0.471
clamp slow
−1.37
1.93
0.68
2.7
0.526
MOSS-TTS
baseline
—
2.82
1.00
7.0
0.740
TABLE IV: Steering on the English Seed-TTS-Eval test set (1088 utterances, each cloning its own prompt and synthesizing its own target text). Additive and clamp applied at the per-model matched-effect intensities. × base = achieved wps ratio to the unsteered baseline; WER via Whisper-large-v3; SIM via Seed-TTS-Eval’s own speaker-verification checkpoint (WavLM-Large + ECAPA-TDNN, wavlm_large_finetune.pth ). Bold = lower WER between the two rules for a given system and direction (both if tied).
Condition
Clamp
Tie
Time-stretch
p
Slow
48.2%
10.5%
41.2%
0.153
Fast
48.2%
11.5%
40.2%
0.099
Combined
48.2%
11.0%
40.8%
0.027
TABLE V: Forced-choice preference: clamp steering vs. a matched-rate time-stretch of the unsteered output (Qwen3-TTS; 36 listeners: 20 per direction, 20 pairs each). Cells are the share of comparisons (clamp / tie / time-stretch); p is a two-sided exact sign test on decisive votes.