cs.SDOct 8, 2026

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

Authors: Shiao Zhu, Lianbo Liu, Sizhen Lyu, Yuzhe Wang, Sheng Li, Takahiro Shinozaki

Organizations: Institute of Science Tokyo, Tokyo, Japan · Independent Researcher, Tokyo, Japan

Abstract

Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.

Figures & tables

Explore similar work

CardsList
  1. ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech

    Jun 20, 2026Wei Xue, Junlan Feng, Shilei Zhang +9TTS SynthesisTTS Evaluation

  2. GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

    Jun 4, 2026Jaehoon Kang, Yejin Lee, Kyuhong ShimTTS SynthesisAdapter Tuning