cs.SDJul 26, 2026

Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion

Authors: Hanlei ZhangZhongming MaMingyang ZhangTengfei LiuYushi ChengYanjiao Chen

Organizations: Zhejiang University · Hangzhou, China · Ant Group · Shanghai, China

Abstract

Voice conversion (VC) poses a significant threat to biometric security by allowing attackers to impersonate target speakers. In forensic contexts, recovering the source speaker's identity from converted audio is vital for narrowing the field of suspects. To address this, we propose TRIDENT, a retracing framework designed to restore a source speaker's original identity from a converted audio sample. TRIDENT utilizes a three-pronged architecture consisting of a primary extractor and two auxiliary branches. The first auxiliary branch identifies the underlying voice conversion mechanism. This design acknowledges that even if the exact conversion strategy is unknown, a high-performance model adopted by the attacker is typically a derivative or variant of established mainstream ones. The second auxiliary branch extracts a latent representation of the target speaker, facilitating the isolation of target-specific traits from the composite converted audio sample. Finally, the main extractor leverages insights from both auxiliary branches to decouple confounding factors and distill a highly discriminative representation of the source speaker's identity. Experimental results demonstrate that TRIDENT achieves an accuracy as high as 90.99% against 7 state-of-the-art voice conversion methods. Furthermore, TRIDENT maintains robust performance under challenging conditions, including telephony channels, unseen languages, and adaptive scenarios.

Explore similar work

CardsList
  1. NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization

    Jul 4, 2026Meiying Melissa Chen, Anastasia Kuznetsova, Zhenyu Wang +1AnonymizationVariational Autoencoder