cs.SDSep 30, 2026

MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion

Authors: Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo

Organizations: NTT, Inc., Japan

Abstract

Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately 9×9\times faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.

Figures & tables

Explore similar work

CardsList
  1. MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion

    Jun 8, 2026Guobin Ma, Yuxuan Xia, Yuepeng Jiang +6Voice ConversionStreaming

  2. Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

    Jun 18, 2026Yudong Li, Zihao Fang, Junwen Qiu +4SpeakerProsody

  3. From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data

    Jun 7, 2026Moshe Mandel, Shlomo E. ChazanVoice ConversionWav2Vec