cs.SDJan 31, 2026

RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models

Authors: Ruinan Jin, Xinting Liao, Hanlin Yu, Deval Pandya, Xiaoxiao Li

Organizations: The University of British Columbia · Vector Institute

Abstract

Modern Voice Cloning (VC) can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing. In practical deployments, modern audio generation models inevitably encounter noisy reference audios, imperfect text prompts, multilingual and long-form generation settings, downstream post-processing, and adversarial perturbations, all of which can significantly hurt robustness. Despite rapid progress in VC driven by autoregressive codec-token language models and diffusion-based models, robustness under realistic deployment shifts remains underexplored. This paper introduces RVCBench, a comprehensive dataset and benchmark that evaluates Robustness in Voice Clone. RVCBench contributes a large-scale, task-aligned robustness dataset that instantiates realistic deployment shifts through controlled text-audio pairing, multilingual and long-form scenarios, expressive prompts, post-processing conditions, and passive or proactive audio perturbations. Covering 18 robustness evaluations, 204 unique speakers, and 14,370 utterance-level evaluation items, RVCBench enables unified evaluation of input sensitivity, generation stability, output resilience, and perturbation robustness. We evaluate 18 representative modern open-source VC models and reveal systematic vulnerabilities in content consistency, speaker similarity, long-form stability, post-processing resilience, adversarial robustness, and detector-facing separability. We open-source the toolkit and dataset to support reproducible evaluation and future research.

Figures & tables

Appendix figures & tables46 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Voice "Cloning" is Style Transfer

    May 15, 2026Kaitlyn Zhou, Federico Bianchi, Martijn Bartelds +3Cross-Lingual Voice CloningVoice Conversion

  2. VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

    Oct 13, 2025Jiliang Hu, Wenfu Wang, Zuchao Li +6Dialogue BenchmarksLarge Audio Language Models

  3. Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

    Jun 18, 2026Satwinder Singh, Qianli Wang, Zihan Zhong +4Cross-Lingual Voice CloningDysarthric Speech