cs.SDJun 5, 2026

Mitigating Proxy-to-Wild Domain Gap in Deepfake Speech

Authors: Xuanjun ChenYun-Shing WuWei-Chung LuClaire LinHaibin WuHung-yi LeeJyh-Shing Roger Jang

Organizations: Graduate Institute of Communication Engineering, National Taiwan University · Graduate Institute of Networking and Multimedia, National Taiwan University · Department of Information Management, National Taiwan University · 4NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE) · 1Graduate Institute of Communication Engineering, National Taiwan University

Abstract

Recent neural audio codec-based speech generation (CodecFake) produces highly realistic audio, posing a challenge to existing deepfake countermeasure models. While using codec resynthesized speech (CoRS) as proxy data improves performance, it often suffers from limited generalization. We propose Domain-Shift Feature Augmentation (DSFA), which simulates "in-the-wild" variations by transforming deterministic feature statistics into stochastic distributions during fine-tuning. To evaluate generalization, we further introduce Codec-based Speech Generation Extension Evaluation (CoSG ExtEval) dataset, a more challenging extension of the CoSG Eval (from CodecFake+) dataset, featuring 40 unseen generative models and long-form audio. Experimental results demonstrate that combining a post-trained SSL backbone with DSFA effectively narrows the proxy-to-wild domain gap. This approach achieves state-of-the-art performance across diverse CodecFake attacks in both CoSG Eval and CoSG ExtEval.

Explore similar work

CardsList
  1. FlowFake: Liquid Networks for Audio Deepfake Detection

    Jun 17, 2026Shivaay Dhondiyal, Divyansh Sharma, Dinesh Kumar VishwakarmaAudio Deepfake DetectionFake News

  2. Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks

    Jun 18, 2026Sameek Bhattacharya, Bharath Krishnamurthy, Ajita RattaniNeural Audio CodecsAudio