Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in real-world applications. Reinforcement learning has recently shown promise for post-training TTS systems to align with human preference, yet existing methods have not been applied to word-level prosodic control. We present EmphTTS, a non-autoregressive TTS system that applies Group Relative Policy Optimization (GRPO) to the duration predictor with an emphasis localization reward, enabling direct optimization for word-level emphasis. Evaluations show that EmphTTS achieves the best emphasis controllability and performs the best in emphasis objective evaluation. In subjective preference tests, EmphTTS is significantly preferred over synthetic groundtruth and most baselines. Ablation studies show that GRPO improves emphasis realization beyond supervised-finetuning-based duration modeling and simple speaking-rate adjustment, while alleviating the mismatch between the independently trained duration predictor and TTS model.
Figures & tables
Figure 1 : The overview of the SFT (left) and GRPO (right) stage. (a) Left : The TTS model and the duration predictor are fine-tuned using the flow-matching objective ( LFM ) and cross-entropy ( LCE ). (b) Right : The duration predictor is trained with GRPO while the TTS model is frozen for inference.
Seed-TTS EN
TinyStress-15k
WER ↓
SIM ↑
CMOS ↑
WER ↓
SIM ↑
StressLM F1 ↑
EmphPref [-1pt] L/W (%)
Groundtruth
1.86%
∼
∼
0.25%
∼
0.60
36/53.7 ∗
Baselines
CosyVoice3
1.76%
0.70
−0.23±0.15
0.32%
0.68
0.26
37.7/56.0 ∗
FishAudio-S2
0.99%
0.65
−0.22±0.18
0.25%
0.67
0.41
33.2/56.5 ∗
Qwen3-TTS-VD
1.70%
∼
−0.06±0.17
0.29%
∼
0.63
44.8/45.7
VoxCPM2
1.84%
0.75
−0.07±0.15
0.48%
0.73
0.32
36.1/54.2 ∗
Table 1: Objective and subjective evaluation results on Seed-TTS EN and TinyStress-15k . Qwen3-TTS-VD denotes Qwen3-TTS-Voice-Design ; ∼ denotes unavailable results. Column-best results are bold , and the best results among our base, ablation, and proposed systems are underlined . EmphPref L/W reports the one-versus-rest EmphPref comparison between EmphTTS and each listed system from the perspective of EmphTTS : L denotes the percentage of trials in which EmphTTS loses, and W denotes the percentage in which EmphTTS wins. Ties are omitted and account for the remaining percentage. ∗ indicates that EmphTTS is significantly preferred over the compared system under a one-sided binomial test on non-tie responses. CMOS is reported with 95% confidence intervals.
Audiobook narration, conversational agents, and audiovisual dubbing require speech that conveys changing emotions and adapts its pacing within a single utterance. But most existing TTS systems typically rely on utterance-level style conditioning, making such fine-grained control difficult to achieve. In light of this, and inspired by the success of post-training in large language models, we propose a unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration. Supervised fine-tuning establishes instruction-conditioned speech generation, while reinforcement learning with group relative policy optimization refines control accuracy using emotion and duration rewards alongside content and speaker preservation objectives. By reusing the pretrained architecture, our approach avoids additional inference-time control modules. Experiments demonstrate significantly improved fine-grained controllability while maintaining speech intelligibility and speaker identity, highlighting post-training as a practical approach to extending existing speech synthesis models.
Lianru Gao, Yujie Guo, Yong Qin
TMCC, College of Computer Science, Nankai University, Tianjin, China
Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training. We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations. For each target emotion we train a lightweight low-rank transform, using a multi-expert objective that encourages monotonic emotion control while preserving speaker identity and linguistic content, constraining steering drift, and keeping the TTS backbone frozen. To optimize through discrete speech tokens, we introduce a two-pass generation-and-replay pipeline using a straight-through estimator to backpropagate expert supervision through sampled tokens. At inference, a target-emotion steering direction is optimized with its respective transform and injected into the base TTS model. Objective and subjective evaluations with Qwen3-TTS across seen, unseen, and accented speakers show stronger continuous emotion control with limited speaker and content degradation. SteerSpeech achieves 1.08x-7.12x baseline target-emotion scores and for a representative emotion subjectively, it receives 78.1%-96.8% intensity preference and 1.43x-1.46x speaker-identity preservation at high steering strengths.
Afsara Benazir, Darius Pétermann, Felix Xiaozhu Lin +1
Large language model (LLM)-based text-to-speech (TTS) systems enable prompt-conditioned emotional control but struggle with fine-grained emotion intensity due to the semantic -- acoustic gap between text and speech. To address this challenge, we formulate emotion intensity control in LLM-based TTS as a learning-to-rank problem and propose Emo-LiPO, a listwise preference optimization framework that aligns prompt-conditioned speech generation with relative emotion intensity expressed in text. Emo-LiPO explicitly models global intensity ordering within each emotion under fixed transcripts, enabling more faithful and continuous emotional expression. We further construct ESD-plus, a multi-speaker dataset with explicit emotion intensity variations, to support fine-grained emotion modeling and evaluation. Experiments on ESD-plus demonstrate that Emo-LiPO significantly improves emotion accuracy and intensity controllability over both supervised- and DPO-based LLM TTS baselines, with particularly pronounced gains at high intensity levels.
Yihang Lin, Li Zhou, Congwei Cao +4
The Chinese University of Hong Kong, Shenzhen · Shenzhen Loop Area Institute · Agency for Science, Technology and Research +2