Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately 9× faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.
Figures & tables
Figure 1: Overview of MeanVoiceFlow2 . The student jointly learns a computationally efficient content encoder cϕ and an average velocity network uϕ through (a) conversion distillation and (b) real-data reconstruction. We further incorporate diffusion-GAN training with sample mixing using a discriminator Dψ to promote realism, and teacher-guided conditioning augmentation based on xθaug to promote disentanglement.
Conv
Rec
UT ↑
DNSP ↑
DNS ↑
CER ↓
SECS ↑
(a)
✓
4.00
2.94
3.80
1.8
0.885
(b)
✓
3.60
2.40
3.77
0.1
0.641
(c)
✓
✓
4.04
2.97
3.80
1.5
0.885
Table 1: Analysis of joint conversion and reconstruction. Conv and Rec indicate the use of conversion distillation and real-data reconstruction, respectively.
GAN
Diffuse
Mix
UT ↑
DNSP ↑
DNS ↑
CER ↓
SECS ↑
(a)
None
–
–
4.04
2.79
3.75
2.2
0.882
(b)
Proposed
4.01
2.87
3.79
1.9
0.883
(c)
Proposed
✓
4.04
2.86
3.78
1.5
0.882
(d)
Proposed
✓
3.93
2.90
3.80
1.9
0.882
(e)
Proposed
✓
✓
4.04
2.97
3.80
1.5
0.885
(f)
WD [ 55 ]
–
–
4.04
2.89
3.79
1.8
0.884
Table 2: Analysis of adversarial training. Diffuse and Mix indicate the use of diffusion-GAN training and sample mixing, respectively.
UT ↑
DNSP ↑
DNS ↑
CER ↓
SECS ↑
(a)
w/o CondAug
4.04
2.97
3.80
1.5
0.885
(b)
w/ CondAug
4.05
2.99
3.81
1.2
0.887
(c)
Direct Distill
4.05
2.94
3.80
1.9
0.884
Table 3: Analysis of conditioning augmentation (CondAug). Direct Distill adds an explicit ℓ1 loss between the student and teacher content representations to the w/o CondAug objective.
nMOS ↑
sMOS ↑
UT ↑
DNSP ↑
DNS ↑
CER ↓
SECS ↑
RTF ↓
(a)
GT
4.26 ± .09 ∗
3.64 ± .06 ∗
4.15
2.89
3.75
0.1
0.940
–
(b)
DiffVC
3.43 ± .11 ∗
2.24 ± .10 ∗
3.76
2.64
3.75
5.4
0.880
0.19
(c)
MVF
3.76 ± .09 ∗
2.74 ± .11
3.98
2.85
3.78
1.2
0.886
0.0072
(d)
MVF2
3.93 ± .10
2.70 ± .11
4.05
2.99
3.81
1.2
0.887
0.00084
(e)
FVG2
3.72 ± .10 ∗
2.63 ± .10
4.03
2.79
3.82
1.2
0.890
0.00084
Table 4: Comparison with previous models in terms of subjective metrics (nMOS and sMOS with 95% confidence intervals), objective metrics, and RTF. ∗ indicates a statistically significant difference from MVF2 on the Mann–Whitney U test ( p<0.05 ).