Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.
Figures & tables
Training scorers
Independent
Quality
System
WER ↓
Cos ↑
WER ↓
Cos ↑
MOS ↑
CosyVoice3
5.34
0.738
7.87
0.770
3.41
Qwen3-TTS
4.69
0.734
6.18
0.701
3.67
OmniVoice
3.18
0.733
4.38
0.704
3.18
OmniVoice+SFT
2.94
0.765
4.11
0.726
3.20
ASR-only
2.35
0.705
3.72
0.688
3.29
Table 1: CV3-Eval Multilingual Voice Cloning (Russian, 500 prompts). Adapters use 2,000 updates; the lower block contains RAWD-TTS variants. Cos: raw cosine. Adapter rows not marked final are checkpoints selected by joint training-scorer reward on these prompts (updates 1,300 for Joint K=3 , 1,500 for K=5 ); final rows use update 2,000 without selection. Dashes: not scored. Bold: column optimum; shading: highest joint training-scorer reward.
K
B
G
WER ↓
Cos ↑
1
4
8
2.82
0.736
3
4
8
2.55
0.733
5
4
8
2.47
0.738
3
8
4
2.72
0.733
3
2
16
2.59
0.725
Table 2: Joint-reward ablations: 1,000 updates, 500 evaluation prompts, and BG=32 . The shared control uses K=3 , B=4 , and G=8 . Checkpoint selection and highlighting follow Table 1 .
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.
Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen speaking styles from uncommon scenarios (e.g., crosstalk, dialects). Moreover, fine-tuning pretrained models requires large, high-quality datasets, limiting rapid personalization. We propose VoiceTTA, a reinforcement learning-based test-time adaptation (TTA) method that improves voice imitation of pretrained zero-shot TTS models. VoiceTTA introduces two style rewards based on coefficient-of-variation differences of F0 and energy, combined with speaker similarity and intelligibility (WER from a pretrained Whisper model), and optimizes learnable prefixes via group relative preference optimization (GRPO) in a flow matching-based model at inference time. Extensive experiments demonstrate substantial improvements on uncommon speech prompts, outperforming state-of-the-art baselines. Audio samples are available at https://voicetta.pages.dev/
Tianxin Xie, Chenxing Li, Dong Yu +1
The Hong Kong University of Science and Technology (Guangzhou) · Tencent
In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to speak 30 languages. X-Voice is trained on a 420K-hour multilingual corpus using the International Phonetic Alphabet (IPA) as a unified representation. To eliminate the reliance on prompt text without complex preprocessing like forced alignment, we design a two-stage training paradigm. In Stage 1, we establish X-Voices1 through standard conditional flow-matching training and use it to synthesize 10K hours of speaker-consistent segments as audio prompts. In Stage 2, we fine-tune on these audio pairs with prompt text masked to derive X-Voices2, which enables zero-shot voice cloning without requiring transcripts of audio prompts. Architecturally, we extend F5-TTS by implementing a dual-level injection of language identifiers and decoupling and scheduling of Classifier-Free Guidance to facilitate multilingual speech synthesis. Subjective and objective evaluation results demonstrate that X-Voice outperforms existing flow-matching based multilingual systems like LEMAS-TTS and achieves zero-shot cross-lingual cloning capabilities comparable to billion-scale models such as Qwen3-TTS. To facilitate research transparency and community advancement, we open-source all related resources.
Rixi Xu, Qingyu Liu, Haitao Li +10
1MoE Key Lab of Artificial Intelligence, X-LANCE Lab, Shanghai Jiao Tong University · Center for Language and Speech Processing, Johns Hopkins University · 2Shanghai Innovation Institute +4