EnvSSLAM-FFN: Lightweight Layer-Fused System for ESDD 2026 Challenge
Authors: Xiaoxuan Guo, Hengyan Huang, Jiayi Zhou, Renhe Sun, Jian Liu, Haonan Cheng, Long Ye, Qin Zhang
Organizations: State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing, China · Machine Intelligence, Ant Group, Shanghai, China · Key Laboratory of Media Audio & Video, Ministry of Education, Communication University of China, Beijing, China
Recent advances in generative audio models have enabled high-fidelity environmental sound synthesis, raising serious concerns for audio security. The ESDD 2026 Challenge therefore addresses environmental sound deepfake detection under unseen generators (Track 1) and black-box low-resource detection (Track 2) conditions. We propose EnvSSLAM-FFN, which integrates a frozen SSLAM self-supervised encoder with a lightweight FFN back-end. To effectively capture spoofing artifacts under severe data imbalance, we fuse intermediate SSLAM representations from layers 4-9 and adopt a class-weighted training objective. Experimental results show that the proposed system consistently outperforms the official baselines on both tracks, achieving Test Equal Error Rates (EERs) of 1.20% and 1.05%, respectively.
Figures & tables
Figure 1: Overview of the proposed EnvSSLAM-FFN pipeline
Figure 2: Trends of normalized fusion weights across the 12 SSLAM layers in EnvSSLAM‑FFN over training steps.
#
System for Track 1
Eval EER
Test EER
1
AASIST (baseline)
15.26
15.02
2
BEATs+AASIST (baseline)
14.21
13.20
3
EnvSSLAM-FFN (ours)
1.05
1.20
#
System for Track 2
Eval EER
Test EER
1
AASIST (baseline)
15.72
15.40
2
BEATs+AASIST (baseline)
12.64
12.48
Table 1: EER (%) on Track 1 and Track 2 for the official baselines and our EnvSSLAM-FFN system.