eess.SPJul 7, 2026

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

Authors: Adrien SchneiderKacper ZabkowskiAnderson AugusmaFrédérique LetuéMaria Camila PinzonDominique Vaufreydaz

Organizations: M-PSI · Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG, 38000 Grenoble, France · SAM, SVH · Univ. Grenoble Alpes, CNRS, Grenoble INP, LJK, 38000 Grenoble, France

Abstract

The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech. It relies on content embeddings extracted from a frozen pretrained wav2vec2 encoder. These embeddings are decoded into an anonymized signal using vector quantization and a HiFi-GAN vocoder, both trained on LibriTTS without any waveform reconstruction loss or speaker embedding mapping. The training objective enforces that embeddings of the anonymized signal match those of the original one. While training, an auxiliary speaker classification branch with a gradient reversal layer is used to discard speakerspecific information. Results show that this straightforward embedding-based approach achieves very low WER (2.53) with an anonymization performance (EER 13.39) ranking within first level for VPC. Notably, emotions are partially preserved (UAR 43.91), even without a supporting training objective, while the anonymized voice is audible without reconstruction loss.

Explore similar work

CardsList
  1. DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

    Apr 29, 2026Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews +2AnonymizationProsody

  2. NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization

    Jul 4, 2026Meiying Melissa Chen, Anastasia Kuznetsova, Zhenyu Wang +1AnonymizationVariational Autoencoder