Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech Architectures), a speaker representation framework that extracts speaker information from intermediate content encoder layers and aggregates using attentive statistics pooling. The embedding is trained jointly with VC objectives, removing the need for a separate speaker encoder. Across six datasets, VOSSA improves F0 dynamics and vowel-discriminative acoustic cues while maintaining comparable NISQA-MOS, WER, and speaker similarity. Perceptual tests further indicate improvements in naturalness, speaker similarity, intelligibility, and vibrancy.
Figures & tables
Figure 1: Training workflow for the TVTSyn backbone . (a) content encoder trained against HuBERT k-means pseudo-labels, and (b) decoder conditioned on speaker embedding trained with self-supervision and discriminator objectives. (c) Overview of the training protocol in VOSSA . The self-reconstruction path (bottom, purple) uses segments from the same LibriTTS speaker to provide fully supervised training, while the non-parallel VC path (top, orange) uses VoxCeleb targets to condition conversion. (d) Speaker embedding extraction from the frozen content encoder. We collect features from the last CNN layer and every other layer of the MHSA stack, concatenate them to form H∈RT×Ld , and apply attentive statistics pooling and an MLP to obtain a global speaker embedding.
Models
Ground truth
slt24 [ 1 ]
DarkStream [ 2 ]
GenVC-s [ 6 ]
TVTSyn [ 3 ]
VOSSA (Ours)
Standard VC metrics and harmonics
NISQA-MOS (↑)
4.06±0.84
3.46±0.80
3.12±0.81
3.04±0.83
3.52±0.81
3.48±0.91
WER (↓)
0.07±0.16
0.18±0.25
0.25±0.27
0.20±0.23
0.17±0.23
0.17±0.23
Simsrcsyn(↓)
–
0.21±0.11
0.06±0.11
–
0.10±0.10
0.11±0.31
Simtrgsyn(↑)
–
0.46±0.11
0.54±0.12
–
0.59±0.11
0.86±0.19
HNR (↑)
–
9.52±2.90
7.93±3.69
8.67±3.04
9.90±2.75
9.72±2.80
Table 1: Evaluation of VOSSA against SOTA streaming VC baselines. Simsrcsyn can be interpreted as anonymization strength (lower is better), whereas Simtrgsyn represents VC strength (higher is better). Best results are bolded ; second-best are underlined .
Figure 2: F1 distribution by vowel height for target speaker id00061 (Voxceleb). GT: original speech.
Streaming zero-shot voice conversion struggles to disentangle timbre from linguistic content without degrading utility or inflating latency. Current methods rely on information bottleneck (IB) or speaker perturbation. While IB filters out timbre, it discards prosody, forcing models to explicitly inject features like fundamental frequency. This often requires buffering future frames, creating algorithmic lookahead latency. On the other hand, existing perturbation methods largely overlook the crucial trade-off between timbre leakage and utility preservation. Recognizing this neglected trade-off, we find that the inherent objective of Speaker Anonymization (SA) aligns well with balancing these factors. Thus, we introduce SA as a novel perturbation mechanism to explicitly mitigate timbre leakage while retaining prosodic utility. Crucially, SA's robust representations significantly alleviate the generator's reliance on future context, enabling our strictly causal, zero-lookahead network. Audio samples are available at https://amphionteam.github.io/Zero-VC-demo/.
Yudong Li, Zihao Fang, Junwen Qiu +4
The Chinese University of Hong Kong, Shenzhen · Shenzhen Transsion Holdings Co., Ltd. · Shenzhen Loop Area Institute +1
Streaming zero-shot voice conversion (VC) has become increasingly popular due to its potential for real-time applications. The recently proposed MeanVC achieves lightweight streaming zero-shot VC, but it has several limitations: its chunk-wise autoregressive denoising doubles the effective training sequence length, conversion quality degrades under small-chunk settings, and its timbre encoder directly relies on reference mel-spectrograms, making it sensitive to reference audio quality. To address these limitations we propose MeanVC 2. We introduce future-receptive chunking (FRC), which explicitly schedules past and future receptive fields across diffusion transformer decoder layers and removes clean-chunk teacher forcing. By incorporating bounded future context, FRC enables stable conversion with a 40 ms chunk size. We further introduce a universal timbre token encoder, which constructs a timbre representation from a global speaker embedding and retrieves fine-grained timbre cues via cross-attention, improving robustness to low-quality references and enhancing zero-shot speaker similarity. Experimental results show that MeanVC 2 significantly outperforms MeanVC, while reducing latency from 211 ms to 110 ms. Audio samples are publicly available. The source code will be publicly released.
Guobin Ma, Yuxuan Xia, Yuepeng Jiang +6
The University of New South Wales, Australia · WeNet Open Source Community, China
Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately 9× faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.