cs.SDSep 29, 2026

RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning

Authors: Maxim Maslov, Kirill Borodin, Vasilii Kudryavtsev, Nikita Vasiliev, Grach Mkrtchian

Organizations: lab260, Yerevan, Armenia · BitmanagerAI, Dubai, UAE · MTUCI, Moscow, Russia

Abstract

Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.

Figures & tables

Explore similar work

CardsList
  1. VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

    Jun 25, 2026Tianxin Xie, Chenxing Li, Dong Yu +1Flow-Matching Text-To-SpeechTest-Time Adaptation

  2. X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

    May 7, 2026Rixi Xu, Qingyu Liu, Haitao Li +10Cross-Lingual Voice CloningMultilingual Automatic Speech Recognition