Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.
Figures & tables
Figure 1: Mean within-utterance Spearman correlations across both development partitions.
Selection
S2S C/S/CLSP ↑
Emo. ↑
MCD ↓
Random
.3105/.3049/.8135
.5964
6.3266
ParaCLAP-C T2S
.3181/.3127/.8180
.6044
6.3052
ParaCLAP-S T2S
.3125/.3104/.8163
.5986
6.3142
CLSP T2S
.3173/.3101/.8198
.6032
6.2711
ParaCLAP-C S2S
.5133 / .4595 / .8465
.6210
6.0731
ParaCLAP-S S2S
.4529 / .5262 /.8406
.6176
6.1581
Table 1: Synthesis performance by instruction-selection criterion on the combined development partitions. Bold and underline indicate the best and second-best results, respectively.
Figure 2: Overview of SRSP. Frozen-TTS target-speech likelihood rewards candidate instructions during training, while GRPO updates the planner. Dashed arrows denote training-only paths; snowflakes denote frozen components.
System
T2S Score ↑
S2S Score ↑
Emo. Cos. ↑
MCD-DTW ↓
ParaCLAP-C
ParaCLAP-S
CLSP
ParaCLAP-C
ParaCLAP-S
CLSP
Raw TTS
–
–
–
.3370 / .3494
.3246 / .3303
.8160/.8148
.6098/.6119
6.2426/6.3932
Base LLM
-.0089/-.0141
.0161/.0160
.4718/.4686
.3166/.3220
.2979/.3078
.8150/.8146
.6124/.5927
6.3395/6.5072
AF-Next †
.0788 / .0841
.0996 / .0973
.5404 / .5423
.2874/.2982
.2596/.2637
.7842/.7863
.5770/.5677
6.8890/7.0128
Qwen3-Omni †
-.0087/-.0119
.0225/.0193
.5101 / .5053
.3176/.3312
.3083/.3130
.8194 / .8182
.6228 / .6226
6.1593 / 6.3189
SRSP (ours)
.0073 / -.0044
.0324 / .0232
.4591/.4518
.3632 / .3627
.3496 / .3502
.8329 / .8296
.6417 / .6269
5.9569 / 6.1637
Table 2: Results on test_in / test_out . Bold and underline indicate the best and second-best results, respectively. † : descriptions conditioned on the ground-truth target waveform. Paired Wilcoxon tests with Holm correction across 40 downstream acoustic comparisons show significant gains ( pHolm<.05 ) in 38 cases; the exceptions are test_out emotion similarity vs. Raw TTS and Qwen3-Omni.
Figure 3: Abbreviated prompting configurations for style generation and LLM-based expressive speech evaluation.
System
SRSP wins vs. system
System wins vs. GT
Context
Reference
Context
Raw TTS
65.54/63.92
66.31/64.38
12.26/13.09
Base LLM
54.19/53.67
60.91/57.75
15.90 /16.87
AF-Next
75.63/75.75
75.01/75.05
7.42/9.35
Qwen3-Omni
55.21/56.65
58.50/57.50
14.95/ 16.95
SRSP (ours)
–
–
18.41 / 19.06
Table 3: Win rates (%) in LLM-based expressive speech evaluation, reported as test_in / test_out . Rates exclude ties and invalid judgments (together <1% of all judgments).
Recent advances in text-to-speech (TTS) have greatly improved speech naturalness, speaker similarity, and controllability. However, most existing controllable TTS systems still rely on explicit user-provided style prompts, making it difficult to automatically determine how a sentence should be spoken in long and complex conversational scenarios. This proposal introduces the ISCSLP 2026 CoT-TTS Challenge, which aims to evaluate whether a system can infer the intended speaking manner from contextual information and generate speech consistent with both the reasoning output and the surrounding scene. The challenge contains two tracks: text-context-aware CoT-TTS and audio-context-aware CoT-TTS. We construct a large-scale bilingual training set from speech-rich media and provide carefully filtered evaluation data for leaderboard comparison. Each system is required to output both a chain-of-thought reasoning analysis and the generated speech waveform. The official evaluation combines objective metrics, multimodal LLM-based evaluation, and human subjective assessment. To facilitate reproducibility, we provide inference code together with a fine-tuning recipe for a 0.6B Qwen3-based model trained via a three-stage strategy. This challenge is expected to support research on context understanding, chain-of-thought reasoning, and expressive speech generation for applications such as film dubbing, audiobook production, virtual characters, and spoken dialogue agents. Further information about the associated challenge is available at:https://iscslp2026-cot-tts.github.io/challenge-website/
Wei Xue, Junlan Feng, Shilei Zhang +9
The Hong Kong University of Science and Technology · China Mobile · Jiutian Artificial Intelligence Technology (Beijing) Co., Ltd., China Mobile +3
Dynamic speaking style adaptation in multi-turn multimodal interaction remains a major challenge for text-to-speech (TTS) systems. Existing context-aware TTS (CTTS) methods typically map dialogue context to speech in an end-to-end manner. Such implicit modeling makes contextual style decisions difficult to supervise, while the entanglement of style, timbre, and content often leads to weak instruction-following and severe timbre drift across turns. To overcome these limitations, we propose Interactive TTS, a dynamic, style-adaptive framework for contextually appropriate and speaker-consistent speech generation. Interactive TTS decouples the process by explicitly modeling contextual style decisions as executable instructions. To bridge the gap between style decisions and speech generation, we introduce Iterative Rejection Sampling Fine-Tuning (Iterative RSFT) and Context-Aware Direct Preference Optimization (CADPO), which significantly enhance instruction-following and align the generated speech with conversational contexts. Extensive experiments demonstrate that Interactive TTS outperforms state-of-the-art models on VStyle and SpeechParaling-Bench. Demo is available at https://wjtian-wonderful.github.io/InteractiveTTS/
We propose GLASS, a framework for composable acoustic style control in zero-shot autoregressive text-to-speech (TTS) that learns controls from post-generation rewards rather than style labels. In zero-shot TTS, a speaker prompt often entangles speaker identity with prosodic attributes such as speaking rate and pitch, making it difficult to change style without changing the prompt itself. GLASS instead treats each acoustic attribute as a reward-defined control direction. For each control axis, GLASS freezes the TTS backbone and trains one lightweight LoRA adapter with Group Relative Policy Optimization (GRPO), using speech-token length and mean F0 as style rewards and WER as an intelligibility anchor. Because each control is represented as a LoRA weight update, independently trained adapters can be swapped, interpolated, and composed through linear LoRA arithmetic without retraining the backbone. Experiments on speaking rate and pitch control show targeted style shifts while preserving naturalness, speaker similarity, and intelligibility, and demonstrate smooth interpolation and multi-axis composition across independently trained adapters.
Jaehoon Kang, Yejin Lee, Kyuhong Shim
Department of Artificial Intelligence, Sungkyunkwan University, Korea