eess.ASSep 23, 2026

EmphTTS: an emphasis-control TTS with reinforcement learning

Authors: Zirui Li, Rech Silas, Lauri Juvela, Tom Backstrom, Mikko Kurimo

Organizations: Department of Information and Communications Engineering, Aalto University, Espoo, Finland

Abstract

Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in real-world applications. Reinforcement learning has recently shown promise for post-training TTS systems to align with human preference, yet existing methods have not been applied to word-level prosodic control. We present EmphTTS, a non-autoregressive TTS system that applies Group Relative Policy Optimization (GRPO) to the duration predictor with an emphasis localization reward, enabling direct optimization for word-level emphasis. Evaluations show that EmphTTS achieves the best emphasis controllability and performs the best in emphasis objective evaluation. In subjective preference tests, EmphTTS is significantly preferred over synthetic groundtruth and most baselines. Ablation studies show that GRPO improves emphasis realization beyond supervised-finetuning-based duration modeling and simple speaking-rate adjustment, while alleviating the mismatch between the independently trained duration predictor and TTS model.

Figures & tables

Explore similar work

CardsList
  1. Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

    Sep 10, 2026Lianru Gao, Yujie Guo, Yong QinFlow-Matching Text-To-SpeechEmotion Recognition

  2. Steerspeech: Activation Steering For Emotion Control In Generated Speech

    Oct 7, 2026Afsara Benazir, Darius Pétermann, Felix Xiaozhu Lin +1

  3. Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

    Jun 11, 2026Yihang Lin, Li Zhou, Congwei Cao +4Emotion Recognition