EnvSSLAM-FFN: Lightweight Layer-Fused System for ESDD 2026 Challenge
Authors: Xiaoxuan Guo, Hengyan Huang, Jiayi Zhou, Renhe Sun, Jian Liu, Haonan Cheng, Long Ye, Qin Zhang
Organizations: State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing, China · Machine Intelligence, Ant Group, Shanghai, China · Key Laboratory of Media Audio & Video, Ministry of Education, Communication University of China, Beijing, China
Recent advances in generative audio models have enabled high-fidelity environmental sound synthesis, raising serious concerns for audio security. The ESDD 2026 Challenge therefore addresses environmental sound deepfake detection under unseen generators (Track 1) and black-box low-resource detection (Track 2) conditions. We propose EnvSSLAM-FFN, which integrates a frozen SSLAM self-supervised encoder with a lightweight FFN back-end. To effectively capture spoofing artifacts under severe data imbalance, we fuse intermediate SSLAM representations from layers 4-9 and adopt a class-weighted training objective. Experimental results show that the proposed system consistently outperforms the official baselines on both tracks, achieving Test Equal Error Rates (EERs) of 1.20% and 1.05%, respectively.
Figures & tables
Figure 1: Overview of the proposed EnvSSLAM-FFN pipeline
Figure 2: Trends of normalized fusion weights across the 12 SSLAM layers in EnvSSLAM‑FFN over training steps.
#
System for Track 1
Eval EER
Test EER
1
AASIST (baseline)
15.26
15.02
2
BEATs+AASIST (baseline)
14.21
13.20
3
EnvSSLAM-FFN (ours)
1.05
1.20
#
System for Track 2
Eval EER
Test EER
1
AASIST (baseline)
15.72
15.40
2
BEATs+AASIST (baseline)
12.64
12.48
Table 1: EER (%) on Track 1 and Track 2 for the official baselines and our EnvSSLAM-FFN system.
In this paper, we propose a deep-learning framework for Environmental Sound Deepfake Detection (ESDD) - the task of identifying whether the sound scene and sound event in an input audio recording is fake or real. To this end, we first conduct extensive experiments to explore how individual spectrograms, a wide range of network architectures, and pre-trained models affect the performance of an ESDD model. The experimental results on the benchmark datasets of EnvSDD indicate that detecting deepfake audio of sound scenes and detecting deepfake audio of sound events should be considered as individual tasks. We also show that fine-tuning a pre-trained model is more effective than training a model from scratch for ESDD. Ultimately, our best model, which fine-tunes the pre-trained BEATs model using the proposed two-phase training strategy, achieves an Accuracy of 0.98, F1 score of 0.95, and AUC score of 0.99 on the Test subset of the EnvSDD dataset. Our best model also achieves an Accuracy of 0.86, F1 score of 0.80, and AUC of 0.93 when evaluated cross-dataset on the ESD-Challenge-TestSet dataset.
Khoi Vu, Dat Tran, Khanh Do +8
FPT University, Vietnam · HCM University of Technology, Vietnam · TDT University, Vietnam +3
The Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2), held in conjunction with ICME 2026, evaluated systems for five component-level audio spoofing detection, where speech and environmental sounds may be manipulated independently or jointly. After the challenge concludes, we analyze the final leaderboard and summarize effective design choices from the top-performing submissions. The challenge attracted 94 registrations from 16 countries; after verification of submission requirements and metadata, 13 teams were retained for the final analysis. On the test set, the best system achieved a Macro-F1 score of 0.8775, substantially outperforming the separation-enhanced joint learning baseline (0.6327). Top systems consistently benefited from modular task decomposition, cross-domain self-supervised encoders, targeted data augmentation, and selective ensembling rather than simple model scaling. At the same time, auxiliary EER analyses reveal persistent difficulty in detecting the spoofed environmental component and in generalizing to unseen generators in the test set. This paper reports challenge results and provides insights for future environment-aware deepfake detection research. The CompSpoofV2 dataset and baseline code remain publicly available for reproducibility.
Xueping Zhang, Han Yin, Yang Xiao +4
Duke Kunshan University · Korea Advanced Institute of Science and Technology · The University of Melbourne +3
This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds may be independently manipulated. To address this challenge, a dual-branch deepfake detection framework is proposed to jointly model speech and environmental contextual representations from input audio. Two pretrained models, XLS-R for speech and BEATs for environmental sound, are used to extract complementary contextual representations. A Matching Head is introduced to model representation differences through statistical normalization and representation interaction, enabling estimation of the original class. In parallel, multi-head cross-attention enables effective information exchange between speech and environmental components. The refined representations are processed with residual connections and layer normalization, and passed to an AASIST classifier to predict speech-based and environment-based spoofing probabilities. The model outputs original, speech, and environment predictions. On the test set, the proposed system achieves an F1-score of 70.20% and an environmental EER of 16.54%, outperforming the baseline system.
Khalid Zaman, Qixuan Huang, Muhammad Uzair +1
Graduate School of Advanced Science and Technology · Japan Advanced Institute of Science and Technology · 1-1 Asahidai, Nomi, Ishikawa, Japan 923-1292