eess.ASSep 30, 2026

VOSSA: Voiceprint Optimization for Streaming Speech Architectures

Authors: Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna

Organizations: Department of Computer Science & Engineering, Texas A&M University, College Station, US

Abstract

Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech Architectures), a speaker representation framework that extracts speaker information from intermediate content encoder layers and aggregates using attentive statistics pooling. The embedding is trained jointly with VC objectives, removing the need for a separate speaker encoder. Across six datasets, VOSSA improves F0 dynamics and vowel-discriminative acoustic cues while maintaining comparable NISQA-MOS, WER, and speaker similarity. Perceptual tests further indicate improvements in naturalness, speaker similarity, intelligibility, and vibrancy.

Figures & tables

Explore similar work

CardsList
  1. Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

    Jun 18, 2026Yudong Li, Zihao Fang, Junwen Qiu +4SpeakerProsody

  2. MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion

    Jun 8, 2026Guobin Ma, Yuxuan Xia, Yuepeng Jiang +6Voice ConversionStreaming