eess.ASOct 1, 2026

UZH-CL at ArA-DF 2026: Prompt-Tuned Foundation Models and Track-Adaptive Score Fusion for Arabic Speech Deepfake Detection

Authors: Aref Farhadipour, Teodora Vukovic, Petr Motlicek

Abstract

Detecting synthetic and voice-converted speech remains difficult for low-resource languages with dialectal diversity, where systems must generalize across regional dialects and unseen acoustic channels. We present the UZH-CL submission to the ArA-DF 2026 Shared Task on Arabic speech deepfake detection, covering Track1 (dialect generalization) and Track2 (acoustic robustness). We freeze a W2V-BERT-2.0 backbone and adapt it with \emph{Wavelet Prompt Tuning}, updating under 1% of parameters, and aggregate multi-layer representations with cross-layer attention and a general attentive-statistics pooling head rather than a specialized graph backend. Complementary detectors are obtained by varying adaptation strategy, training data, augmentation, and encoder family. We find that the two shift types require different fusion regimes: a broad multi-window ensemble for dialect generalization, and a compact, channel-matched, center-crop ensemble for acoustic robustness. Official evaluation yields 1.96% EER on Track1 (6th place) and 1.04% EER on Track2 (3rd place), corresponding to 87% and 96% relative reductions over the XLS-R+AASIST baselines.

Explore similar work

CardsList
  1. Alethia: A Foundational Encoder for Voice Deepfakes

    Apr 30, 2026Yi Zhu, Brahmi Dwivedi, Jayaram Raghuram +1Audio Representation LearningAudio Deepfake Detection

  2. ML-ITW: A Multilingual in-the-wild Benchmark for Speech Deepfake Detection

    Mar 6, 2026Daixian Li, Jun Xue, Zhuolin Yi +4Audio Deepfake Detection

  3. FlowFake: Liquid Networks for Audio Deepfake Detection

    Jun 17, 2026Shivaay Dhondiyal, Divyansh Sharma, Dinesh Kumar VishwakarmaAudio Deepfake DetectionCross-Domain Generalization