Instruction-controlled speech synthesis and paralinguistic understanding are often trained independently, leaving reciprocal feedback between the two tasks underexplored. We introduce CycleSpeech, a framework that connects generation and understanding through a shared, structured voice profile that serves as a common target for supervision and reciprocal feedback. The forward cycle assesses whether synthesized speech expresses the intended attributes by comparing recovered and target profiles. The backward cycle evaluates whether profiles inferred from real speech can guide reconstruction of the source speaking style. To support both directions, we construct a bilingual dataset of 20,046 examples pairing instructions, target speech, speaker references, and structured profiles. Building on joint supervised fine-tuning, CycleGRPO alternates policy updates using reciprocal rewards grounded in profile consistency and speaking-style reconstruction. Fixed target profiles anchor feedback from the evolving counterpart. This procedure requires neither human preference annotations nor an additional preference-trained reward model. Evaluations on Chinese and English benchmarks show improved instruction adherence and profile recovery while maintaining competitive synthesis quality. Compared with Step-Audio-2-mini, CycleSpeech improves instruction-match accuracy by 4.50 and 10.06 percentage points in Chinese and English, respectively. Controlled ablations further support the contribution of cycle feedback to generation control. These results support structured voice profiles as an interface for reciprocal training between speech generation and paralinguistic understanding. An online demo is available at https://cyclespeech.github.io.
Figures & tables
Figure 1 : Generation–understanding paradigms. (a) Separate models optimize generation and understanding independently. (b) A shared backbone supports both tasks but still receives one-way supervision. (c) CycleSpeech closes reciprocal profile–speech cycles, allowing generation and understanding to verify each other.
Figure 2 : CycleSpeech training framework. (a) Joint SFT initializes generation (GEN) and understanding (UND) with paired speech–profile supervision. (b,c) CycleGRPO alternately optimizes generation and understanding through reciprocal verifiable rewards.
Attribute
Consistency ↑
Attribute
Consistency ↑
Gender
96.7
Age
95.7
Accent
91.6
Emotion
70.8
Speed
86.7
Volume
72.5
Emo. Detail
4.92
Personality
4.95
Texture
4.87
Scene
4.79
Table 1 : Example of a voice profile. The schema captures speaker, prosodic, affective, and contextual attributes.
Method / Model
CER / WER(%) ↓
Spk. Ver.(%) ↑
Style Acc.(%) ↑
Style SIM ↑
Instruction Match(%) ↑
Part I: Chinese Evaluation Benchmark
GT Wav (upper reference)
4.44 [-0.99, +1.09]
99.00 [-0.83, +0.67]
62.56 [-2.59, +2.60]
1.00 [0.00, 0.00]
81.50 [-3.17, +3.00]
VoxInstruct [ 39 ]
10.90 [-1.61, +1.73]
80.17 [-3.17, +3.17]
58.87 [-2.37, +2.31]
0.87 [-0.01, +0.01]
73.00 [-3.50, +3.50]
CosyVoice3 [ 6 ]
2.76 [-0.87, +1.17]
99.50 [-0.67, +0.50]
56.28 [-2.23, +2.21]
0.92 [-0.01, +0.01]
82.17 [-3.17, +3.00]
OV-InstructTTS [ 22 ]
3.64 [-0.76, +0.82]
99.00 [-0.83, +0.67]
58.63 [-2.21, +2.23]
0.90 [-0.01, +0.01]
79.67 [-3.33, +3.17]
MiMo-Audio-Instruct [ 37 ]
4.42 [-0.77, +0.82]
81.83 [-3.00, +3.00]
58.76 [-2.13, +2.11]
0.92 [-0.01, +0.01]
80.33 [-3.17, +3.17]
Table 3 : Instruction-controlled TTS results. WER/CER and Style Acc. use the shared evaluation protocol in Section 4.1 ; brackets give signed 95% confidence-bound offsets. Best and second-best results are bolded and underlined .
Model
Structure (%) ↑
Attributes (%) ↑
Transcript (%) ↑
Overall (%) ↑
PR
SA
Cat.Avg
Desc.Sim
TF
OS
Part I: Chinese Evaluation Benchmark
Qwen2.5-Omni
100.0
93.1
58.52 [-1.03, +1.03]
45.37 [-0.45, +0.46]
85.9 [-0.82, +0.79]
58.76 [-0.56, +0.54]
MiMo-Audio-7B-Instruct
99.8
99.6
58.10 [-1.06, +1.02]
61.05 [-0.36, +0.35]
83.0 [-0.97, +0.93]
66.25 [-0.60, +0.58]
Kimi-Audio-7B-Instruct
94.4
90.9
45.37 [-1.43, +1.43]
53.54 [-0.86, +0.84]
80.5 [-1.49, +1.41]
56.26 [-1.06, +1.06]
Step-Audio-2-mini
99.0
91.9
55.46 [-1.03, +1.03]
52.12 [-0.60, +0.62]
77.8 [-1.62, +1.59]
57.00 [-0.76, +0.74]
Table 4 : Speech paralinguistic understanding results. PR : parse rate; SA : schema adherence; Cat.Avg : six-field categorical accuracy; Desc.Sim : four-field BGE-M3 cosine similarity; TF : transcription fidelity; OS : overall score. Best and second-best results are bolded and underlined .
Model
WER ↓
SIM ↑
PIT ↑
SPD ↑
VOL ↑
EMO ↑
GT Codec
3.47 ±0.35
.970 ±.002
90.2 ±1.5
88.7 ±1.6
89.7 ±1.5
72.6 ±6.3
Step-Audio-2-mini
1.93 ±0.40
.860 ±.005
69.7 ±2.4
65.6 ±2.4
63.2 ±2.4
33.7 ±6.8
CycleSpeech (Phase1)
2.66 ±0.40
.852 ±.005
73.1 ±2.2
67.1 ±2.4
65.6 ±2.4
37.4 ±6.8
CycleSpeech (Phase2)
2.50 ±0.40
.851 ±.006
72.5 ±2.3
68.3 ±2.4
66.0 ±2.4
40.0 ±7.1
Table 5 : Out-of-domain generation results across WER (%), Speaker Similarity (SIM), and Attribute Accuracies (%) for Pitch (PIT), Speed (SPD), Volume (VOL), and Emotion (EMO).
GEN
UND
Variant
WER/CER (%) ↓
Style Acc. ↑
Cat. Avg. ↑
Desc.Sim ↑
OS ↑
Chinese Test Set
CycleGRPO (Alt-1)
2.47 [-0.60, +0.67]
58.82 [-2.30, +2.28]
74.08 [-1.06, +1.08]
72.37 [-0.41, +0.40]
78.36 [-0.56, +0.55]
CycleGRPO (Alt-10)
2.53 [-0.60, +0.66]
59.69 [-2.23, +2.21]
74.14 [-1.08, +1.05]
72.29 [-0.42, +0.40]
78.31 [-0.56, +0.55]
CycleGRPO (Alt-5, Ours)
2.33 [-0.56, +0.65]
60.42 [-2.30, +2.29]
74.18 [-1.06, +1.08]
72.30 [-0.42, +0.42]
78.36 [-0.56, +0.55]
English Test Set
Table 7 : Effect of the alternating update interval in CycleGRPO. Alt- k switches the optimized branch every k updates.
GEN
UND
Variant
WER/CER (%) ↓
Style Acc. ↑
Cat. Avg. ↑
Desc.Sim ↑
OS ↑
Chinese Test Set
w/o Cycle
2.29 [-0.56, +0.61]
58.92 [-2.26, +2.18]
74.01 [-1.09, +1.09]
72.45 [-0.41, +0.40]
78.35 [-0.56, +0.55]
w/ Frozen FB
2.74 [-0.70, +0.79]
59.66 [-2.24, +2.24]
74.08 [-1.08, +1.08]
72.36 [-0.40, +0.40]
78.34 [-0.56, +0.55]
w/o UND Updates
2.46 [-0.58, +0.65]
59.69 [-2.21, +2.17]
73.56 [-1.08, +1.09]
71.90 [-0.42, +0.40]
77.98 [-0.57, +0.56]
w/o GEN Updates
2.63 [-0.63, +0.71]
60.20 [-2.30, +2.31]
73.97 [-1.06, +1.08]
72.26 [-0.41, +0.40]
78.25 [-0.56, +0.56]
Table 8 : Effects of reciprocal cycle feedback, evolving feedback models, and bidirectional adapter updates. w/o Cycle removes both feedback loops but retains alternating updates ( k=5 ). w/ Frozen FB retains both cycles and alternating updates, with feedback from frozen Joint-SFT counterparts. w/o UND/GEN Updates trains only GEN/UND, respectively, with the opposite adapter frozen at Joint-SFT.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Gen.
Und.
Shared Target
Cycle Opt.
Reward
RL-aligned generation
CosyVoice 3
✓
✗
✗
✗
ASR/SER scores
Fish Audio S2
✓
✗
✗
✗
ASR, quality, speaker
FlexiVoice
✓
✗
✗
✗
Emotion and timbre preference
Evaluation and understanding
InstructTTSEval
✗
✓
✗
✗
External audio-LLM
Appendix
Table S1 : Paradigm comparison across RL-aligned instruction-controlled generation and paralinguistic evaluation or understanding. Shared Target denotes a semantic target shared by generation and understanding; Cycle Opt. denotes reciprocal optimization between them.
Figure S1 : Data annotation pipeline.
Stage
APS
DSD
RP
Total
Pseudo IDs
Joint SFT
6,047
8,266
5,733
20,046
17,043
CycleGRPO
3,791
3,571
2,600
9,962
7,996
Appendix
Table S2 : Training-set composition.
Field
Canonical values
Gender
M , F , unk
Age
child , teen , young , middle , senior , unknown
Speed
slow , medium_slow , medium , medium_fast , fast , unknown
Volume
quiet , normal , loud , unknown
Emotion
Neutral , Happy , Sad , Angry , Surprised , Fearful , Disgusted , Other
gender: Female; age: early to mid-twenties; timbre / pitch: Bright and resonant, with a gentle terminal fall, Crystalline, smooth, slightly breathy, with a subtle hint of nasality; speech rate: Moderate and measured; volume: Normal conversational level; accent annotation: Standard General American accent.
DSD
Use a bright delivery in a General American accent, in a mid register, with measured pacing, at steady volume. Keep the tone grounded and matter-of-fact, with a relaxed presence.
RP
a young female content creator in her early twenties, sitting comfortably in her room, recording a casual vlog update for her followers.
Appendix
Table S4 : Examples of Chinese and English instructions from the APS, DSD, and RP categories. APS specifies acoustic attributes, DSD describes speaking style, and RP provides a communicative scenario.
Figure S2 : Listening-test interface for rating instruction/style faithfulness, naturalness, and speaker similarity.
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over pitch dynamics, speaking rate, and emotional tone often exceed what a single-pass generation can faithfully realize. While recent reasoning models have shown that intermediate "thinking" tokens improve output quality, this paradigm has been confined to the text modality. In this work, we extend reasoning to the audio token space by training a LALM with reinforcement learning to reason over its own speech output. The model first generates a draft speech as a form of audio-token reasoning, critiques its own generation by reflecting on the acoustic realization in text, and then produces a refined version conditioned on both the first-pass speech and the critique, all within a single model. After RL training, the refined two-hop outputs achieve a relative improvement of 7.15% on the InstructTTSEval benchmark, demonstrating the model's reflective ability.
Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang +4
Graduate Institute of Electrical Engineering, National Taiwan University, Taiwan · Graduate Institute of Communication Engineering, National Taiwan University, Taiwan · NVIDIA Research 4 Artificial Intelligence Center of Research Excellence, National Taiwan University, Taiwan
Instruction-based text-to-speech (ITTS) systems enable natural-language control of expressive speech generation, but often offer limited transparency and fine-grained control over individual text units. Character-level controllable TTS systems provide explicit acoustic control, yet typically rely on user-specified acoustic attributes. To bridge this gap, we propose InstCharVoice, a unified framework that grounds natural-language instructions in character-level acoustic control. We first construct grounded instruction annotations on the WordVoice-5A-zh corpus using Qwen3-Omni. With this supervision, we train an autoregressive model to identify instruction-relevant characters and predict their acoustic attributes before generating the corresponding speech tokens. Keyword prediction and grounding-aware loss weighting help the model focus on instruction-relevant characters and attributes. Experiments show improved instruction following and keyword-level acoustic control over representative ITTS systems, with competitive speech naturalness and explicit character-level controllability. Audio samples are available at https://xxh333.github.io/instcharvoice-demo/.
Sihang Nie, Xueru Li, Xiaofen Xing +4
South China University of Technology · Huya Inc. · The Hongkong Polytechnic University
Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-form dialogue. We propose a controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement. Its speech backbone, ControlEdit-TTS, unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration. The framework further performs hierarchical utterance-level and scene-level critique, routing detected issues to editing, resynthesis, or timing adjustment. Experiments on a bilingual Chinese--English dialogue benchmark show improved utterance-level instruction following, better dialogue-level preference than direct dialogue models and agentic baselines, and more effective refinement than regeneration-only alternatives while preserving speaker identity. Ablations further confirm the benefits of scene-level critique and edit-based correction.