RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning
Organizations: lab260, Yerevan, Armenia · BitmanagerAI, Dubai, UAE · MTUCI, Moscow, Russia
Abstract
Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.
Figures & tables
| Training scorers | Independent | Quality | |||
| System | WER | Cos | WER | Cos | MOS |
| CosyVoice3 | 5.34 | 0.738 | 7.87 | 0.770 | 3.41 |
| Qwen3-TTS | 4.69 | 0.734 | 6.18 | 0.701 | 3.67 |
| OmniVoice | 3.18 | 0.733 | 4.38 | 0.704 | 3.18 |
| OmniVoice+SFT | 2.94 | 0.765 | 4.11 | 0.726 | 3.20 |
| ASR-only | 2.35 | 0.705 | 3.72 | 0.688 | 3.29 |
| WER | Cos | |||
|---|---|---|---|---|
| 1 | 4 | 8 | 2.82 | 0.736 |
| 3 | 4 | 8 | 2.55 | 0.733 |
| 5 | 4 | 8 | 2.47 | 0.738 |
| 3 | 8 | 4 | 2.72 | 0.733 |
| 3 | 2 | 16 | 2.59 | 0.725 |