RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
Organizations: The University of British Columbia · Vector Institute
Abstract
Modern Voice Cloning (VC) can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing. In practical deployments, modern audio generation models inevitably encounter noisy reference audios, imperfect text prompts, multilingual and long-form generation settings, downstream post-processing, and adversarial perturbations, all of which can significantly hurt robustness. Despite rapid progress in VC driven by autoregressive codec-token language models and diffusion-based models, robustness under realistic deployment shifts remains underexplored. This paper introduces RVCBench, a comprehensive dataset and benchmark that evaluates Robustness in Voice Clone. RVCBench contributes a large-scale, task-aligned robustness dataset that instantiates realistic deployment shifts through controlled text-audio pairing, multilingual and long-form scenarios, expressive prompts, post-processing conditions, and passive or proactive audio perturbations. Covering 18 robustness evaluations, 204 unique speakers, and 14,370 utterance-level evaluation items, RVCBench enables unified evaluation of input sensitivity, generation stability, output resilience, and perturbation robustness. We evaluate 18 representative modern open-source VC models and reveal systematic vulnerabilities in content consistency, speaker similarity, long-form stability, post-processing resilience, adversarial robustness, and detector-facing separability. We open-source the toolkit and dataset to support reproducible evaluation and future research.
Figures & tables
| Metric | FishSpeech | Spark-TTS | MOSS-TTSD v0.7 | Higgs | OZSpeech | StyleTTS2 |
| EER median [IQR] | 8.00 [5.00, 13.67] | 36.33 [30.00, 42.33] | 37.00 [26.67, 45.67] | 23.00 [18.67, 33.67] | 0.67 [0.00, 1.00] | 11.67 [6.67, 22.00] |
Appendix figures & tables46 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Input | Generation | Output | Perturbation | |||||
| Shift | Multi- lingual | Long form | Express. | Codec | Detect. | Passive noise | Anti- clone | Denoise | |
| RVCBench (ours) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| LibriSpeech-PC [ 34 ] | – | – | – | – | – | – | – | – | – |
| ClonEval [ 10 ] | – | – | – | – | – | – | – | – | |
| VCC20 [ 65 ] | – | ✓ | – | – | – | – | – | – | |
| Seed-TTS [ 1 ] | – | – | ✓ | – | – | – | – | – | |
| Model | Model Size | Generation Paradigm | System Class | Core Architecture | Primary Focus | ML Support | Open | Year |
| Autoregressive zero-shot TTS and voice-cloning LMs | ||||||||
| VALL-E [ 52 ] | – | AR (discrete codec tokens) | TTS + voice cloning | Neural codec language model over discrete audio codes | Zero-shot personalized TTS from short acoustic prompt | EN only | No | 2023 |
| Seed-TTS [ 1 ] | – | AR (speech tokens) | TTS + voice cloning | Autoregressive Transformer with speech tokenizer, token LM, token diffusion model, and vocoder | Human-like expressive speech, controllable TTS, voice conversion | Multilingual / cross-lingual | No | 2024 |
| FishSpeech [ 30 ] | 390M | AR (discrete codec tokens) | TTS + voice cloning | Dual autoregressive neural codec LM (fast-slow) | Multilingual TTS, zero-shot cloning | EN, ZH, DE, JA, FR, ES, KO, AR | Yes | 2024 |
| XTTS [ 6 ] | 443M GPT encoder | AR (discrete speech tokens) | TTS + voice cloning | Tortoise-based multilingual zero-shot TTS model | Massively multilingual zero-shot TTS | 16 languages | Yes | 2024 |
| Spark-TTS [ 54 ] | 0.5B | AR (discrete codec tokens) | TTS + voice cloning | LLM-based neural codec LM with single-stream tokens | Efficient LLM-TTS, zero-shot cloning | EN, ZH | Yes | 2025 |
| Model | Model Size | Generation Paradigm | System Class | Core Architecture | Primary Focus | ML Support | Open | Year |
| Autoregressive speech-token, dialogue, and foundation-audio LMs | ||||||||
| IndexTTS [ 13 ] | 2.3B | AR (discrete speech tokens) | TTS + voice cloning | XTTS/Tortoise-inspired model with hybrid character+pinyin text modeling, conformer speaker encoder, and BigVGAN2 | Controllable, efficient zero-shot TTS | EN, ZH | Yes | 2025 |
| MOSS-TTSD v0.7 [ 62 ] | 8B | AR (discrete codec tokens) | Dialogue TTS + voice cloning | Dialogue-oriented speech generation model | Long-form spoken dialogue generation with zero-shot cloning | EN, ZH and other mainstream languages | Yes | 2026 |
| FireRedTTS-2 [ 57 ] | 1.7B | AR (streaming speech tokens) | Dialogue TTS + voice cloning | 12.5 Hz streaming speech tokenizer with text-speech interleaved dual Transformer | Long-form multi-speaker dialogue, low-latency streaming, zero-shot voice cloning | EN, ZH, JA, KO, FR, DE, RU | Yes | 2025 |
| Higgs Audio [ 4 ] | 3B + 2.2B DualFFN | AR (text + audio tokens) | General speech/TTS + voice cloning | Foundation audio model with dual-FFN token processing | General expressive audio generation, multi-speaker speech, and zero-shot cloning | Primarily EN, ZH, KO; includes DE, ES | Yes | 2025 |
| Model | Model Size | Generation Paradigm | System Class | Core Architecture | Primary Focus | ML Support | Open | Year |
| Non-autoregressive, diffusion, flow, and cloning-pipeline TTS | ||||||||
| MaskGCT [ 55 ] | 2.2B | NAR (masked semantic + acoustic tokens) | TTS + voice cloning | Two-stage masked generative codec Transformer | Zero-shot TTS without explicit alignment or duration modeling | Multilingual | Yes | 2024 |
| StyleTTS2 [ 29 ] | 190M | Diffusion | TTS + voice cloning/adaptation | Diffusion-based TTS with adversarial training | Expressive, high-naturalness TTS with reference-style / speaker adaptation | Partial (EN; multilingual via PL-BERT) | Yes | 2023 |
| F5-TTS [ 8 ] | 335.8M | Flow matching | TTS + voice cloning | Fully non-autoregressive flow-matching TTS with DiT backbone + ConvNeXt text refinement | Faithful, fluent, multilingual zero-shot TTS | Multilingual | Yes | 2024 |
| OZSpeech [ 23 ] | 145M + 102M FACodec | Flow matching (continuous-time) | TTS + voice cloning | Conditional flow-matching model with learned-prior one-step sampling | Low-latency zero-shot TTS and prompt-speech cloning | EN only | Yes | 2025 |
| OpenVoice [ 39 ] | 131 MB | Hybrid cloning pipeline | Voice cloning + controllable TTS | Base speaker TTS + tone-color converter for style/voice transfer | Instant voice cloning with flexible style control and cross-lingual cloning | EN, ES, FR, ZH, JA, KO (V2 native) | Yes | 2023 |
| Model | Model Size | Generation Paradigm | System Class | Core Architecture | Primary Focus | ML Support | Open | Year |
| Continuous-latent, editing, tokenizer-free, and proprietary TTS systems | ||||||||
| PlayDiffusion | 1.5B | Diffusion | Reference-conditioned TTS + speech editing/inpainting | Diffusion-based speech generation, editing, and audio inpainting | Reference-conditioned synthesis and context-preserving speech edits | N/A (English tokenizer) | Yes | 2025 |
| VibeVoice [ 37 ] | 1.5B | Latent diffusion / flow (continuous latents) | Long-form TTS + reference-conditioned voice cloning | Next-token diffusion over continuous speech latents | Long-form conversational speech with public voice-prefill conditioning | EN, ZH and expanded experimental voices | Partial | 2025 |
| VoxCPM [ 66 ] | 0.5B | Tokenizer-free diffusion autoregressive | TTS + voice cloning | MiniCPM-4-based hierarchical semantic-acoustic model with FSQ, residual acoustic modeling, Local DiT decoder, and causal AudioVAE | Context-aware expressive speech generation and true-to-life zero-shot voice cloning | ZH, EN | Yes | 2025 |
| GPT-4o mini TTS | – | Proprietary instruction-conditioned neural TTS | TTS only; no public voice cloning | GPT-4o-mini-powered speech endpoint with natural-language style instructions and built-in voices | Steerable text-to-speech, multilingual narration, and realtime/streaming audio output | Multilingual; voices optimized for EN | No | 2025 |
| Gemini-TTS | – | Proprietary instruction-conditioned neural TTS | TTS only; no public voice cloning | Gemini-family TTS with prebuilt voice configuration and single-/multi-speaker synthesis | Prompt-controllable speech synthesis for scripts, narration, podcasts, and multi-speaker audio | 80+ locales / prebuilt voices | No | 2025 |
| Model | Model Size | Generation Paradigm | System Class | Core Architecture | Primary Focus | ML Support | Open | Year |
| Hybrid LM + Diffusion or Flow | ||||||||
| CosyVoice 2 [ 16 ] | 0.5B | Hybrid (LM + causal flow matching) | Streaming TTS + voice cloning | Streaming TTS with LM + chunk-aware causal flow matching | Scalable streaming TTS with zero-shot voice generation and cross-lingual synthesis | ZH, EN, JA, KO, DE, ES, FR, IT, RU | Yes | 2024 |
| Qwen3-TTS [ 20 ] | 0.6B / 1.7B | Hybrid (dual-track LM + block-wise DiT reconstruction) | Streaming TTS + voice cloning + voice design | Dual-track LM with 25 Hz / 12 Hz speech tokenizers and block-wise DiT waveform reconstruction | Multilingual, controllable, robust, streaming TTS with voice cloning and voice design | ZH, EN, JA, KO, DE, FR, RU, PT, ES, IT | Yes | 2026 |
| MGM-Omni [ 51 ] | 0.6B / 2B / 4B | Multimodal LLM + speech generation head | Speech/TTS generation + voice cloning | Omni-modal LLM with dual-track speech generation | Agentic, personalized, long-horizon speech generation with streaming zero-shot cloning | EN, ZH | Yes | 2025 |
| GLM-TTS [ 11 ] | – | Hybrid (Text Token AR + Token Wav diffusion) | Streaming TTS + voice cloning | Two-stage: text-to-speech-token autoregressive SpeechLM + token-to-waveform diffusion (Flow) + vocoder | Production-level controllable, emotion-expressive zero-shot TTS | ZH, EN (incl. dialect + singing data) | Yes | 2025 |
| Task | Dataset Name | Evaluation | Sources | Preprocess / Setup |
| Input Robustness | ||||
| (1) Reference audio shifts | RVC-AudioShift | Demography | VCTK | Demography : VCTK speakers across 12 accents, three age groups, and gender groups. |
| (2) Text prompt shifts | RVC-TextShift | Hallucination , Scam | VCTK, Robocall | Hallucination : text prompts with special tokens, formatting templates, and mixed-language fragments. Scam : text prompts with persuasive or high-risk robocall-style content. |
| Generation Robustness | ||||
| (3) Multilingual generalizability | RVC-Multilingual | Chinese-VC , English-VC , French-VC , CrossLingual | VCTK, LibriTTS, AISHELL-1, Common Voice FR, EMIME | English-VC : VCTK and LibriTTS for English in-domain VCL . Chinese-VC and French-VC : reference audio and text are in Mandarin and French, respectively. CrossLingual : text prompt and reference audio are in different languages. |
| (4) Long-form generation | RVC-LongContext | LongText , LongAudio | LibriTTS, LibriSpeech-Long | LongText : generate minutes-scale utterances per speaker. LongAudio : use reference audios with different durations to evaluate sensitivity to prompt-audio length. |
| Dataset | Lang | Spk | Utts | U/S | SR | Speaker overlap |
| Standard English Benchmarks | ||||||
| VCTK [ 58 ] | EN | 40 | 4,000 | 100 | 48k | VB: 4; Robocall: 10 |
| LibriTTS [ 60 ] | EN | 40 | 4,000 | 100 | 24k | Multi: 5; Long: 2 |
| Multilingual & Cross-Lingual | ||||||
| AISHELL-1 [ 5 ] | ZH | 40 | 2,000 | 50 | 16k | – |
| Common Voice FR [ 2 ] | FR | 40 | 2,000 | 50 | 48k | – |
| Metric | What it Measures (Attributes) |
| Generation quality: Speaker Identity & Content | |
| SIM | Speaker embedding cosine similarity via ECAPA-TDNN verification score (SpeechBrain) |
| SVA | Speaker verification accept/reject decision (SpeechBrain; converted to boolean) |
| WER | Linguistic content correctness via Whisper transcription + normalized WER |
| Generation quality: Acoustic Quality & Naturalness | |
| MCD | Spectral distortion vs. ground truth using DTW-aligned Mel-cepstral distance |
| Observed vulnerability | Potential future measure | RVCBench component |
| Accent & demographic sensitivity | Demographically balanced training, speaker-disentangled embeddings, and group-aware calibration | RVC-AudioShift |
| Noisy or multi-speaker references | Reference-quality scoring, source separation, multi-reference aggregation, and noise augmentation | RVC-PassiveNoise |
| Irregular or OOD text prompts | Text normalization, prompt sanitization, uncertainty-aware generation, and semantic consistency checking | RVC-TextShift |
| Long-form content & speaker drift | Hierarchical generation, explicit memory mechanisms, chunk-level consistency constraints, and speaker-consistency regularization | RVC-LongContext |
| Weak expressive or prosodic control | Prosody supervision, emotion/style disentanglement, and style-consistency objectives | RVC-Expression |
| Compression & channel sensitivity | Codec-aware training, compression augmentation, differentiable post-processing, and robust vocoder design | RVC-Compression |
| Dataset | Baseline | SIM | MOS | WER | MCD | RTF | SVA | EMC |
| LibriTTS | FishSpeech | 3.61 | ||||||
| Fish Audio S2 | 4.67 | |||||||
| OZSpeech | 8.75 | |||||||
| StyleTTS2 | 0.11 | |||||||
| Spark-TTS | 1.56 | |||||||
| MOSS-TTSD v0.7 | 0.62 |
| Model | Direction | SIM | MOS | WER | MCD | SVA |
| FishSpeech | Eng Man | |||||
| Man Eng | ||||||
| Fish Audio S2 | Eng Man | |||||
| Man Eng | ||||||
| OZSpeech | Eng Man | |||||
| Man Eng |
| Accent | SIM | MOS | WER | MCD | EMC |
| Model: FishSpeech | |||||
| American | |||||
| Australian | |||||
| British | |||||
| Canadian | |||||
| English | |||||
| Accent | SIM | MOS | WER | MCD | EMC |
| Model: Spark-TTS | |||||
| American | |||||
| Australian | |||||
| British | |||||
| Canadian | |||||
| English | |||||
| Accent | SIM | MOS | WER | MCD | EMC |
| Model: Higgs | |||||
| American | |||||
| Australian | |||||
| British | |||||
| Canadian | |||||
| English | |||||
| Accent | SIM | MOS | WER | MCD | EMC |
| Model: VibeVoice | |||||
| American | |||||
| Australian | |||||
| British | |||||
| Canadian | |||||
| English | |||||
| Accent | SIM | MOS | WER | MCD | EMC |
| Model: F5-TTS | |||||
| American | |||||
| Australian | |||||
| British | |||||
| Canadian | |||||
| English | |||||
| Accent | SIM | MOS | WER | MCD | EMC |
| Model: IndexTTS | |||||
| American | |||||
| Australian | |||||
| British | |||||
| Canadian | |||||
| English | |||||
| Gender | SIM | MOS | WER | MCD | EMC |
| Model: FishSpeech | |||||
| Female | |||||
| Male | |||||
| Model: Fish Audio S2 | |||||
| Female | |||||
| Male | |||||
| Gender | SIM | MOS | WER | MCD | EMC |
| Model: GLM-TTS | |||||
| Female | |||||
| Male | |||||
| Model: VibeVoice | |||||
| Female | |||||
| Male | |||||
| Age Group | SIM | MOS | WER | MCD | EMC |
| Model: FishSpeech | |||||
| under 20 | |||||
| 20–29 | |||||
| 30 and over | |||||
| Model: Fish Audio S2 | |||||
| under 20 | |||||
| Age Group | SIM | MOS | WER | MCD | EMC |
| Model: GLM-TTS | |||||
| under 20 | |||||
| 20–29 | |||||
| 30 and over | |||||
| Model: VibeVoice | |||||
| under 20 | |||||
| Model | Group | SIM | MOS | WER | SVA | EMC | EmTXT | DNSMOS | Dur (s) |
| FishSpeech | robocall | ||||||||
| vctk | |||||||||
| Fish Audio S2 | robocall | ||||||||
| vctk | |||||||||
| OZSpeech | robocall | ||||||||
| vctk |
| SQ-LLM | AASIST | HuBERT-ECAPA | |||||||
| Generated model | mDCF | EER | ACC | mDCF | EER | ACC | mDCF | EER | ACC |
| Fish. | 34.33 | 18.33 | 81.83 | 26.00 | 13.67 | 86.17 | 10.00 | 5.00 | 94.83 |
| Fish S2 | 15.33 | 8.33 | 91.83 | 6.67 | 3.33 | 96.50 | 5.00 | 2.67 | 97.50 |
| OZS. | 4.00 | 2.33 | 97.50 | 0.00 | 0.00 | 99.83 | 1.33 | 1.00 | 98.83 |
| Style. | 30.33 | 15.33 | 84.83 | 52.67 | 27.33 | 72.83 | 7.00 | 4.00 | 96.17 |
| Spark. | 76.33 | 42.00 | 58.17 | 82.00 | 48.00 | 52.17 | 57.33 | 29.33 | 70.50 |
| RawGAT-ST | RawNet2 | TCM-ADD | |||||||
| Generated model | mDCF | EER | ACC | mDCF | EER | ACC | mDCF | EER | ACC |
| Fish. | 7.67 | 5.33 | 94.83 | 14.33 | 8.00 | 92.17 | 8.00 | 4.67 | 95.50 |
| Fish S2 | 4.33 | 3.00 | 97.17 | 13.33 | 7.00 | 93.17 | 1.33 | 1.00 | 98.83 |
| OZS. | 0.00 | 0.00 | 99.83 | 2.00 | 1.00 | 98.83 | 0.00 | 0.00 | 99.83 |
| Style. | 21.33 | 11.67 | 88.17 | 42.00 | 22.00 | 78.17 | 7.00 | 3.67 | 96.50 |
| Spark. | 64.00 | 36.33 | 63.83 | 77.00 | 46.00 | 53.83 | 63.67 | 36.33 | 63.83 |
| XLSR-SLS | Wav2Vec2-ECAPA | WavLM-ECAPA | |||||||
| Generated model | mDCF | EER | ACC | mDCF | EER | ACC | mDCF | EER | ACC |
| Fish. | 8.33 | 4.67 | 95.50 | 33.67 | 17.00 | 83.17 | 15.67 | 8.00 | 91.83 |
| Fish S2 | 5.67 | 4.33 | 95.83 | 23.33 | 11.67 | 88.17 | 11.67 | 6.00 | 94.17 |
| OZS. | 0.00 | 0.00 | 99.83 | 24.00 | 13.67 | 86.17 | 0.67 | 0.67 | 99.17 |
| Style. | 12.67 | 6.67 | 93.50 | 74.67 | 53.00 | 47.17 | 14.00 | 7.33 | 92.83 |
| Spark. | 75.00 | 42.33 | 57.83 | 59.00 | 30.00 | 69.83 | 58.67 | 29.67 | 70.50 |
| Model | Noise | SIM | MOS | WER | MCD | EMC | DNSMOS |
| FishSpeech | -5dB | ||||||
| +0dB | |||||||
| +5dB | |||||||
| +10dB | |||||||
| Fish Audio S2 | -5dB | ||||||
| +0dB |
| Model | Noise | SIM | MOS | WER | MCD | EMC | DNSMOS |
| MOSS-TTS v1.5 | -5dB | ||||||
| +0dB | |||||||
| +5dB | |||||||
| +10dB | |||||||
| dots.tts | -5dB | ||||||
| +0dB |
| Model | Noise | SIM | MOS | WER | MCD | EMC | DNSMOS |
| VibeVoice | -5dB | ||||||
| +0dB | |||||||
| +5dB | |||||||
| +10dB | |||||||
| MGM-Omni | -5dB | ||||||
| +0dB |
| Model | Noise | SIM | MOS | WER | MCD | EMC | DNSMOS |
| OpenVoice | -5dB | ||||||
| +0dB | |||||||
| +5dB | |||||||
| +10dB | |||||||
| XTTS | -5dB | ||||||
| +0dB |
| Model | SIM | MOS | WER | RTF | SVA |
| FishSpeech | |||||
| Fish Audio S2 | |||||
| OZSpeech | |||||
| StyleTTS2 | |||||
| Spark-TTS | |||||
| MOSS-TTSD v0.7 |
| Protection | SNR | MCD | SIM | WER | SpeechMOS | DNSMOS |
| Gaussian Noise | 6.10 | 11.25 | 0.71 | 0.19 | 1.39 | 1.98 |
| SPEC | 8.08 | 12.67 | 0.59 | 0.34 | 1.40 | 1.97 |
| POP | 17.08 | 2.08 | 0.92 | 0.18 | 3.00 | 3.02 |
| SafeSpeech | 9.48 | 11.51 | 0.62 | 0.26 | 1.48 | 2.06 |
| Enkidu | 8.07 | 8.03 | 0.72 | 0.24 | 1.52 | 2.26 |
| Denoise-SPEC | 15.30 | 3.39 | 0.51 | 0.29 | 3.23 | 3.15 |
| Protection | SIM | MOS | WER | MCD | RTF | SVA | EMC |
| Model: FishSpeech | |||||||
| Gaussian Noise | |||||||
| POP | |||||||
| Enkidu | |||||||
| SafeSpeech | |||||||
| SPEC | |||||||
| Protection | SIM | MOS | WER | MCD | RTF | SVA | EMC |
| Model: Higgs | |||||||
| Gaussian Noise | |||||||
| POP | |||||||
| Enkidu | |||||||
| SafeSpeech | |||||||
| SPEC | |||||||
| Protection | SIM | MOS | WER | MCD | RTF | SVA | EMC |
| Model: F5-TTS | |||||||
| Gaussian Noise | |||||||
| POP | |||||||
| Enkidu | |||||||
| SafeSpeech | |||||||
| SPEC | |||||||
| Model | SIM | MOS | WER | MCD | RTF | SVA | EMC |
| Dataset: Long LibriSpeech | |||||||
| FishSpeech | |||||||
| Fish Audio S2 | |||||||
| OZSpeech | |||||||
| StyleTTS2 | |||||||
| Spark-TTS | |||||||
| Model | SIM | MOS | WER | MCD | RTF | SVA |
| Dataset: Common Voice (French) | ||||||
| FishSpeech | ||||||
| Fish Audio S2 | ||||||
| OZSpeech | ||||||
| StyleTTS2 | ||||||
| Spark-TTS | ||||||
| Dataset | Model | SIM | MOS | WER | MCD | RTF | SVA | EMC |
| LibriTTS (DEMUCS on SPEC) | FishSpeech | |||||||
| Fish Audio S2 | ||||||||
| OZSpeech | ||||||||
| StyleTTS2 | ||||||||
| Spark-TTS | ||||||||
| MOSS-TTSD v0.7 |
| Model | SIM | MOS | WER | MCD | RTF | SVA | EMC |
| Condition: Background-VCTK (Clean) | |||||||
| FishSpeech | |||||||
| Fish Audio S2 | |||||||
| OZSpeech | |||||||
| StyleTTS2 | |||||||
| Spark-TTS | |||||||
| Model | Condition | SIM | MOS | WER | MCD | RTF | SVA | EMC |
| FishSpeech | Clean | 3.61 | ||||||
| SPEC | 4.16 | |||||||
| DEMUCS | 2.42 | |||||||
| Fish Audio S2 | Clean | 4.67 | ||||||
| SPEC | 1.27 | |||||||
| DEMUCS | 1.01 |
| Model | Condition | SIM | MOS | WER | MCD | RTF | SVA | EMC |
| GLM-TTS | Clean | 1.74 | ||||||
| SPEC | 1.15 | |||||||
| DEMUCS | 4.61 | |||||||
| VibeVoice | Clean | 1.86 | ||||||
| SPEC | 2.35 | |||||||
| DEMUCS | 1.06 |
| Model | STOI | MCD | SIM | WER |
| Codec: AAC 32k | ||||
| FishSpeech | ||||
| Fish Audio S2 | ||||
| OZSpeech | ||||
| StyleTTS2 | ||||
| Spark-TTS | ||||
| Model | STOI | MCD | SIM | WER |
| Codec: MP3 32k | ||||
| FishSpeech | ||||
| Fish Audio S2 | ||||
| OZSpeech | ||||
| StyleTTS2 | ||||
| Spark-TTS | ||||
| Model | STOI | MCD | SIM | WER |
| Codec: Opus 16k | ||||
| FishSpeech | ||||
| Fish Audio S2 | ||||
| OZSpeech | ||||
| StyleTTS2 | ||||
| Spark-TTS | ||||
| Model | STOI | MCD | SIM | WER |
| Codec: Phone NB | ||||
| FishSpeech | ||||
| Fish Audio S2 | ||||
| OZSpeech | ||||
| StyleTTS2 | ||||
| Spark-TTS | ||||