Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately 9× faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.
Figures & tables
Figure 1: Overview of MeanVoiceFlow2 . The student jointly learns a computationally efficient content encoder cϕ and an average velocity network uϕ through (a) conversion distillation and (b) real-data reconstruction. We further incorporate diffusion-GAN training with sample mixing using a discriminator Dψ to promote realism, and teacher-guided conditioning augmentation based on xθaug to promote disentanglement.
Conv
Rec
UT ↑
DNSP ↑
DNS ↑
CER ↓
SECS ↑
(a)
✓
4.00
2.94
3.80
1.8
0.885
(b)
✓
3.60
2.40
3.77
0.1
0.641
(c)
✓
✓
4.04
2.97
3.80
1.5
0.885
Table 1: Analysis of joint conversion and reconstruction. Conv and Rec indicate the use of conversion distillation and real-data reconstruction, respectively.
GAN
Diffuse
Mix
UT ↑
DNSP ↑
DNS ↑
CER ↓
SECS ↑
(a)
None
–
–
4.04
2.79
3.75
2.2
0.882
(b)
Proposed
4.01
2.87
3.79
1.9
0.883
(c)
Proposed
✓
4.04
2.86
3.78
1.5
0.882
(d)
Proposed
✓
3.93
2.90
3.80
1.9
0.882
(e)
Proposed
✓
✓
4.04
2.97
3.80
1.5
0.885
(f)
WD [ 55 ]
–
–
4.04
2.89
3.79
1.8
0.884
Table 2: Analysis of adversarial training. Diffuse and Mix indicate the use of diffusion-GAN training and sample mixing, respectively.
UT ↑
DNSP ↑
DNS ↑
CER ↓
SECS ↑
(a)
w/o CondAug
4.04
2.97
3.80
1.5
0.885
(b)
w/ CondAug
4.05
2.99
3.81
1.2
0.887
(c)
Direct Distill
4.05
2.94
3.80
1.9
0.884
Table 3: Analysis of conditioning augmentation (CondAug). Direct Distill adds an explicit ℓ1 loss between the student and teacher content representations to the w/o CondAug objective.
nMOS ↑
sMOS ↑
UT ↑
DNSP ↑
DNS ↑
CER ↓
SECS ↑
RTF ↓
(a)
GT
4.26 ± .09 ∗
3.64 ± .06 ∗
4.15
2.89
3.75
0.1
0.940
–
(b)
DiffVC
3.43 ± .11 ∗
2.24 ± .10 ∗
3.76
2.64
3.75
5.4
0.880
0.19
(c)
MVF
3.76 ± .09 ∗
2.74 ± .11
3.98
2.85
3.78
1.2
0.886
0.0072
(d)
MVF2
3.93 ± .10
2.70 ± .11
4.05
2.99
3.81
1.2
0.887
0.00084
(e)
FVG2
3.72 ± .10 ∗
2.63 ± .10
4.03
2.79
3.82
1.2
0.890
0.00084
Table 4: Comparison with previous models in terms of subjective metrics (nMOS and sMOS with 95% confidence intervals), objective metrics, and RTF. ∗ indicates a statistically significant difference from MVF2 on the Mann–Whitney U test ( p<0.05 ).
Streaming zero-shot voice conversion (VC) has become increasingly popular due to its potential for real-time applications. The recently proposed MeanVC achieves lightweight streaming zero-shot VC, but it has several limitations: its chunk-wise autoregressive denoising doubles the effective training sequence length, conversion quality degrades under small-chunk settings, and its timbre encoder directly relies on reference mel-spectrograms, making it sensitive to reference audio quality. To address these limitations we propose MeanVC 2. We introduce future-receptive chunking (FRC), which explicitly schedules past and future receptive fields across diffusion transformer decoder layers and removes clean-chunk teacher forcing. By incorporating bounded future context, FRC enables stable conversion with a 40 ms chunk size. We further introduce a universal timbre token encoder, which constructs a timbre representation from a global speaker embedding and retrieves fine-grained timbre cues via cross-attention, improving robustness to low-quality references and enhancing zero-shot speaker similarity. Experimental results show that MeanVC 2 significantly outperforms MeanVC, while reducing latency from 211 ms to 110 ms. Audio samples are publicly available. The source code will be publicly released.
Guobin Ma, Yuxuan Xia, Yuepeng Jiang +6
The University of New South Wales, Australia · WeNet Open Source Community, China
Streaming zero-shot voice conversion struggles to disentangle timbre from linguistic content without degrading utility or inflating latency. Current methods rely on information bottleneck (IB) or speaker perturbation. While IB filters out timbre, it discards prosody, forcing models to explicitly inject features like fundamental frequency. This often requires buffering future frames, creating algorithmic lookahead latency. On the other hand, existing perturbation methods largely overlook the crucial trade-off between timbre leakage and utility preservation. Recognizing this neglected trade-off, we find that the inherent objective of Speaker Anonymization (SA) aligns well with balancing these factors. Thus, we introduce SA as a novel perturbation mechanism to explicitly mitigate timbre leakage while retaining prosodic utility. Crucially, SA's robust representations significantly alleviate the generator's reliance on future context, enabling our strictly causal, zero-lookahead network. Audio samples are available at https://amphionteam.github.io/Zero-VC-demo/.
Yudong Li, Zihao Fang, Junwen Qiu +4
The Chinese University of Hong Kong, Shenzhen · Shenzhen Transsion Holdings Co., Ltd. · Shenzhen Loop Area Institute +1
We present a voice conversion (VC) framework that utilizes K-Nearest Neighbors (KNN) retrieval over WavLM representations to align non-parallel source and target speech, constructing synthetic training pairs for supervised learning. The retrieved segments serve as synthetic inputs, while real target audio provides ground-truth outputs, forming a synthetic-to-real training paradigm that naturally supports multilingual data without requiring parallel corpora or explicit alignment. To ensure consistent target-speaker identity, we incorporate a speaker loss derived from a pretrained speaker verification model. Experiments across multiple languages demonstrate that the proposed approach achieves high naturalness and strong speaker similarity, outperforming competitive VC baselines, despite being trained exclusively on English data. Samples can be accessed at: https://palindromic-vc.github.io.