CycleSpeech: Reciprocal Alignment for Instruction-Controlled Speech Synthesis and Paralinguistic Understanding
Organizations: The Chinese University of Hong Kong, Shenzhen · Tsinghua University · Amphion Technology Co., Ltd.
Abstract
Instruction-controlled speech synthesis and paralinguistic understanding are often trained independently, leaving reciprocal feedback between the two tasks underexplored. We introduce CycleSpeech, a framework that connects generation and understanding through a shared, structured voice profile that serves as a common target for supervision and reciprocal feedback. The forward cycle assesses whether synthesized speech expresses the intended attributes by comparing recovered and target profiles. The backward cycle evaluates whether profiles inferred from real speech can guide reconstruction of the source speaking style. To support both directions, we construct a bilingual dataset of 20,046 examples pairing instructions, target speech, speaker references, and structured profiles. Building on joint supervised fine-tuning, CycleGRPO alternates policy updates using reciprocal rewards grounded in profile consistency and speaking-style reconstruction. Fixed target profiles anchor feedback from the evolving counterpart. This procedure requires neither human preference annotations nor an additional preference-trained reward model. Evaluations on Chinese and English benchmarks show improved instruction adherence and profile recovery while maintaining competitive synthesis quality. Compared with Step-Audio-2-mini, CycleSpeech improves instruction-match accuracy by 4.50 and 10.06 percentage points in Chinese and English, respectively. Controlled ablations further support the contribution of cycle feedback to generation control. These results support structured voice profiles as an interface for reciprocal training between speech generation and paralinguistic understanding. An online demo is available at https://cyclespeech.github.io.
Figures & tables
| Attribute | Consistency | Attribute | Consistency |
| Gender | 96.7 | Age | 95.7 |
| Accent | 91.6 | Emotion | 70.8 |
| Speed | 86.7 | Volume | 72.5 |
| Emo. Detail | 4.92 | Personality | 4.95 |
| Texture | 4.87 | Scene | 4.79 |
| Method / Model | CER / WER(%) | Spk. Ver.(%) | Style Acc.(%) | Style SIM | Instruction Match(%) |
|---|---|---|---|---|---|
| Part I: Chinese Evaluation Benchmark | |||||
| GT Wav (upper reference) | 4.44 [-0.99, +1.09] | 99.00 [-0.83, +0.67] | 62.56 [-2.59, +2.60] | 1.00 [0.00, 0.00] | 81.50 [-3.17, +3.00] |
| VoxInstruct [ 39 ] | 10.90 [-1.61, +1.73] | 80.17 [-3.17, +3.17] | 58.87 [-2.37, +2.31] | 0.87 [-0.01, +0.01] | 73.00 [-3.50, +3.50] |
| CosyVoice3 [ 6 ] | 2.76 [-0.87, +1.17] | 99.50 [-0.67, +0.50] | 56.28 [-2.23, +2.21] | 0.92 [-0.01, +0.01] | 82.17 [-3.17, +3.00] |
| OV-InstructTTS [ 22 ] | 3.64 [-0.76, +0.82] | 99.00 [-0.83, +0.67] | 58.63 [-2.21, +2.23] | 0.90 [-0.01, +0.01] | 79.67 [-3.33, +3.17] |
| MiMo-Audio-Instruct [ 37 ] | 4.42 [-0.77, +0.82] | 81.83 [-3.00, +3.00] | 58.76 [-2.13, +2.11] | 0.92 [-0.01, +0.01] | 80.33 [-3.17, +3.17] |
| Model | Structure (%) | Attributes (%) | Transcript (%) | Overall (%) | ||
|---|---|---|---|---|---|---|
| PR | SA | Cat.Avg | Desc.Sim | TF | OS | |
| Part I: Chinese Evaluation Benchmark | ||||||
| Qwen2.5-Omni | 100.0 | 93.1 | 58.52 [-1.03, +1.03] | 45.37 [-0.45, +0.46] | 85.9 [-0.82, +0.79] | 58.76 [-0.56, +0.54] |
| MiMo-Audio-7B-Instruct | 99.8 | 99.6 | 58.10 [-1.06, +1.02] | 61.05 [-0.36, +0.35] | 83.0 [-0.97, +0.93] | 66.25 [-0.60, +0.58] |
| Kimi-Audio-7B-Instruct | 94.4 | 90.9 | 45.37 [-1.43, +1.43] | 53.54 [-0.86, +0.84] | 80.5 [-1.49, +1.41] | 56.26 [-1.06, +1.06] |
| Step-Audio-2-mini | 99.0 | 91.9 | 55.46 [-1.03, +1.03] | 52.12 [-0.60, +0.62] | 77.8 [-1.62, +1.59] | 57.00 [-0.76, +0.74] |
| Model | WER | SIM | PIT | SPD | VOL | EMO |
|---|---|---|---|---|---|---|
| GT Codec | 3.47 ±0.35 | .970 ±.002 | 90.2 ±1.5 | 88.7 ±1.6 | 89.7 ±1.5 | 72.6 ±6.3 |
| Step-Audio-2-mini | 1.93 ±0.40 | .860 ±.005 | 69.7 ±2.4 | 65.6 ±2.4 | 63.2 ±2.4 | 33.7 ±6.8 |
| CycleSpeech (Phase1) | 2.66 ±0.40 | .852 ±.005 | 73.1 ±2.2 | 67.1 ±2.4 | 65.6 ±2.4 | 37.4 ±6.8 |
| CycleSpeech (Phase2) | 2.50 ±0.40 | .851 ±.006 | 72.5 ±2.3 | 68.3 ±2.4 | 66.0 ±2.4 | 40.0 ±7.1 |
| GEN | UND | ||||
|---|---|---|---|---|---|
| Variant | WER/CER (%) | Style Acc. | Cat. Avg. | Desc.Sim | OS |
| Chinese Test Set | |||||
| CycleGRPO (Alt-1) | 2.47 [-0.60, +0.67] | 58.82 [-2.30, +2.28] | 74.08 [-1.06, +1.08] | 72.37 [-0.41, +0.40] | 78.36 [-0.56, +0.55] |
| CycleGRPO (Alt-10) | 2.53 [-0.60, +0.66] | 59.69 [-2.23, +2.21] | 74.14 [-1.08, +1.05] | 72.29 [-0.42, +0.40] | 78.31 [-0.56, +0.55] |
| CycleGRPO (Alt-5, Ours) | 2.33 [-0.56, +0.65] | 60.42 [-2.30, +2.29] | 74.18 [-1.06, +1.08] | 72.30 [-0.42, +0.42] | 78.36 [-0.56, +0.55] |
| English Test Set | |||||
| GEN | UND | ||||
|---|---|---|---|---|---|
| Variant | WER/CER (%) | Style Acc. | Cat. Avg. | Desc.Sim | OS |
| Chinese Test Set | |||||
| w/o Cycle | 2.29 [-0.56, +0.61] | 58.92 [-2.26, +2.18] | 74.01 [-1.09, +1.09] | 72.45 [-0.41, +0.40] | 78.35 [-0.56, +0.55] |
| w/ Frozen FB | 2.74 [-0.70, +0.79] | 59.66 [-2.24, +2.24] | 74.08 [-1.08, +1.08] | 72.36 [-0.40, +0.40] | 78.34 [-0.56, +0.55] |
| w/o UND Updates | 2.46 [-0.58, +0.65] | 59.69 [-2.21, +2.17] | 73.56 [-1.08, +1.09] | 71.90 [-0.42, +0.40] | 77.98 [-0.57, +0.56] |
| w/o GEN Updates | 2.63 [-0.63, +0.71] | 60.20 [-2.30, +2.31] | 73.97 [-1.06, +1.08] | 72.26 [-0.41, +0.40] | 78.25 [-0.56, +0.56] |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Gen. | Und. | Shared Target | Cycle Opt. | Reward |
| RL-aligned generation | |||||
| CosyVoice 3 | ✓ | ✗ | ✗ | ✗ | ASR/SER scores |
| Fish Audio S2 | ✓ | ✗ | ✗ | ✗ | ASR, quality, speaker |
| FlexiVoice | ✓ | ✗ | ✗ | ✗ | Emotion and timbre preference |
| Evaluation and understanding | |||||
| InstructTTSEval | ✗ | ✓ | ✗ | ✗ | External audio-LLM |
| Stage | APS | DSD | RP | Total | Pseudo IDs |
|---|---|---|---|---|---|
| Joint SFT | 6,047 | 8,266 | 5,733 | 20,046 | 17,043 |
| CycleGRPO | 3,791 | 3,571 | 2,600 | 9,962 | 7,996 |
| Field | Canonical values |
|---|---|
| Gender | M , F , unk |
| Age | child , teen , young , middle , senior , unknown |
| Speed | slow , medium_slow , medium , medium_fast , fast , unknown |
| Volume | quiet , normal , loud , unknown |
| Emotion | Neutral , Happy , Sad , Angry , Surprised , Fearful , Disgusted , Other |
| English accent | english_general_american , english_british , english_australian , english_indian , english_mid_atlantic , english_other , unknown |
| Source | Type | Instruction |
|---|---|---|
| Chinese | APS | 女性,青壮年(20-30岁);音色与音高偏高、清脆、明亮、金属质感;语速偏快,音量中等偏强;标准普通话。 |
| DSD | 以中等音高为主,音色上体现柔和、纤细、湿润,音量极低,展现出这段语音带有明显的焦虑和紧迫感与秘密感交织的情感层次,性格层面显得紧张、脆弱、警示者,音频中,说话人语速较快,带有宿命般的强调。 | |
| RP | 在古装剧或宫廷题材场景中,一位青年女性角色在私密室内与心腹密探对话,听到的惊人消息后,她眉头微蹙、眼神凝滞,缓慢而沉重地复述这一事实,声音中充满震惊与疑虑,试图确认核实。 | |
| English | APS | gender: Female; age: early to mid-twenties; timbre / pitch: Bright and resonant, with a gentle terminal fall, Crystalline, smooth, slightly breathy, with a subtle hint of nasality; speech rate: Moderate and measured; volume: Normal conversational level; accent annotation: Standard General American accent. |
| DSD | Use a bright delivery in a General American accent, in a mid register, with measured pacing, at steady volume. Keep the tone grounded and matter-of-fact, with a relaxed presence. | |
| RP | a young female content creator in her early twenties, sitting comfortably in her room, recording a casual vlog update for her followers. |