Acoustic foundation models (AFMs) have democratized acoustic applications, enabling powerful models for tasks ranging from speech recognition to speaker verification with minimal resources. However, the security of applications based on AFMs remains largely underexplored. Our work addresses this gap by proposing the Foundation Acoustic model Backdoor (FAB) attack, demonstrating that state-of-the-art AFMs are susceptible to backdooring under practical settings. Despite making minimal assumptions about adversary capabilities (e.g., no access to pre-training data), we show that FAB preserves benign performance while inducing backdoors that survive fine-tuning and cause significant degradation across diverse downstream tasks when activated. Notably, FAB utilizes task-agnostic, physically realizable, inconspicuous, and sync-free triggers (e.g., a background siren). We evaluate FAB using two leading AFMs, nine downstream tasks, and four different triggers. We further demonstrate its effectiveness against established defenses and across both digital and physical domains. While extensive end-to-end fine-tuning can mitigate FAB, such a defense is resource-intensive and task-specific. Our work highlights critical risks to AFMs and calls for advanced defenses.
Figures & tables
Figure 1 : FAB attack overview. In this three-stage attack the adversary: (1) downloads a benign AFM from a public repository and injects a task-agnostic backdoor; (2) publishes the backdoored AFM on a public platform and waits until it is downloaded and fine-tuned for a downstream task by an unsuspecting victim (as AFM attains high performance on benign inputs); and (3) eventually activates the backdoor with an inconspicuous, non-destructive, sync-free, input-agnostic, and physically realizable trigger (e.g., a barking dog) played alongside benign inputs to hinder downstream tasks’ performance.
Table 1 : Comparison of backdoors’ threat models against acoustic models.
Figure 2 : SUPERB scores for benign ( fθ ) and backdoored ( fθ^ ) HuBERT model, on normal inputs X (top) and trigger-stamped inputs X^ (bottom), using the siren trigger, for different tasks. * indicates values < -1 or > 1.
Dig. X
Phys. X
Dig. X^
Phys. X^ (1 speaker)
Phys. X^ (2 speakers)
fθ
11.22
15.30
13.77
25.11
24.84
fθ^
11.46
16.56
99.88
80.71
82.74
Table 2 : Backdoored ( fθ^ ) and benign ( fθ ) ASR models’ performance on benign and trigger-stamped samples, fed digitally or physically played and recorded.
Input Type
fθ
POR
FAB (ours)
X
11.4
14.2
11.4
X^
14.4
59.4
98.4
Table 3 : Comparison between benign model ( fθ ), basline POR attack, and our FAB attack on ASR downstream task (in WER ↓ ) for benign ( X ) or trigger-stamped ( X^ ) inputs.
Model
No trigger
Bark
Beep
Bark-then-beep
Beep-then-bark
fθ^
11.4
14.9
11.9
12.5
81.9
fθ
11.4
18.6
13.0
13.8
14.2
Table 4 : WER (%) on the ASR task when injecting a backdoor with a composite beep-then-bark trigger and activating it with different triggers. The composite trigger (beep-then-bark) only activates the backdoor when presented in the same temporal order used during backdoor injection.
Figure 3 : SUPERB scores on normal inputs X (top) and trigger-stamped inputs X^ (bottom), for backdoored models where we inject the trigger at different layers. * indicates values > 1 or < -1.
Defense
Fine-pruning
Input filtering
Rate
0%
20%
40%
60%
0%
10%
20%
30%
40%
Benign
11.4
13.9
17.9
23.4
11.4
12.9
17.8
39.7
83.2
Trigger
98.4
97.2
85.6
37.3
98.1
98.4
98.6
99.1
99.7
Table 5 : Comparison of fine-pruning and input filtering defences against FAB at different rates on the ASR task (in WER ↓ ).
Figure 4 : Effect of excessive end-to-end fine-tuning with different dataset sizes (10h to 100h) on FAB’s performance on ASR downstream task (in WER ↓ ) for benign ( X ) or trigger-stamped ( X^ ) inputs. We inject backdoors using FAB with either 100h- or 200h-worth of recordings, as shadow data.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Content
Paralinguistics
Semantics
Speaker
Model
Input
ASR ↓
KS ↑
PR ↓
ER ↑
IC ↑
ST ↑
ASV ↓
SD ↓
SID ↑
fθ
X
11.4
95.7
5.6
62.0
98.1
15.9
5.8
6.3
81.9
X^
14.4
93.3
8.2
61.2
95.5
14.2
6.1
6.8
79.1
fθ^
X
11.4
94.3
5.4
61.3
98.2
15.9
5.5
6.6
76.9
X^
98.4
28.0
99.8
34.7
6.6
0.9
31.3
25.3
0.7
Appendix
Table 6 : Downstream task’s performance after fine-tuning with benign ( fθ ) and backdoored ( fθ^ ) AFMs, when providing benign ( X ) or trigger-stamped ( X^ ) samples as input.
Table 7 : Measuring the downstream performance on benign ( X ) and trigger-stamped ( X^ ) after fine-tuning with AFMs backdoored with different layers’ representations selected to inject the backdoor (i.e., when minimizing LBack ). Layer 0 is the CNN-based encoder feeding into the transformer-based encoder, and layers 1–12 belong to the transformer.
Table 8 : Comparison of the effect of injecting different trigger on downstream task’s performance after fine-tuning with benign ( fθ ) and backdoored ( fθ^ ) AFMs, on both benign ( X ) or trigger-stamped ( X^ ) samples. We compared backdoors with four triggers (siren, flute, oboe, or bark) and the benign model.
Content
Paralinguistics
Semantics
Speaker
Model
Input
ASR ↓
KS ↑
PR ↓
ER ↑
IC ↑
ST ↑
ASV ↓
SD ↓
SID ↑
fθ
X
10.4
97.0
4.8
62.5
98.6
16.3
4.5
4.9
84.1
X^
10.9
95.8
5.1
61.5
98.0
15.4
4.8
6.1
82.3
fθ^
X
11.4
93.9
4.6
59.7
96.9
14.5
5.0
4.9
70.2
X^
98.7
25.0
99.8
45.4
4.7
0.9
20.4
18.3
3.1
Appendix
Table 9 : Downstream task’s performance after fine-tuning with benign ( fθ ) and backdoored ( fθ^ ) WavLM-based AFMs, when providing benign ( X ) or trigger-stamped ( X^ ) samples as input.
Table 10 : The effect of the FAB’s trigger’s SNR during trigger injection and backdoor activation on downstream task performance, when providing benign ( X ) or trigger-stamped ( X^ ) samples as input.
Figure 5 : SUPERB scores for benign ( fθ ) and backdoored ( fθ^ ) WavLM model, on normal inputs X (top) and trigger-stamped inputs X^ (bottom), using the siren trigger, for different tasks. * indicates values < -1 or > 1.
Figure 6 : SUPERB scores for benign model ( fθ ) and backdoored model ( fθ^ ) on benign inputs X (top) and trigger-stamped inputs X^ (bottom), using different trigger types during injection and backdoor activation. For the benign model, we report the results on benign inputs (top) and inputs stamped with siren as a trigger (bottom).* indicates values > 1 or < -1.
Figure 7 : SUPERB scores on normal inputs X (top) and trigger-stamped inputs X^ (bottom), for backdoored models with and without knowledge of the code-book. * indicates values > 1 or < -1.
Table 11 : Ablations on the ASR downstream task, in WER ↓ . (a) varying the weight κ that balances LBack against LBenign (larger κ emphasizes increasing error on trigger-stamped inputs). (b) fine-tuning models backdoored with varying amounts of Xaux . (c) backdooring towards different target representations v : the all-1s vector, or vectors drawn at random from normal distributions.
Figure 8 : Left: FABs’ success, evaluated on ASR (in WER ↓ ), when altering the weight, κ (x-axis), used to balance LBack and LBenign (larger κ puts more emphasis on increasing error on trigger-stamped inputs, X^ ). Right: ASR’s performance (in WER ↓ ) when fine-tuning on model’s backdoored with varying amounts of samples in Xaux (x-axis).
Figure 9 : The SUPERB scores for benign model ( fθ ) and backdoored model ( fθ^ ) on normal inputs X (a) and on trigger-stamped inputs X^ (b)–(d), for different trigger SNRs. * indicates that the value is higher than 1 or lower than −1 .
With the rapid development of deep learning, its vulnerability has gradually emerged in recent years. This work focuses on backdoor attacks on speech recognition systems. We adopt sounds that are ordinary in nature or in our daily life as triggers for natural backdoor attacks. We conduct experiments on two datasets and three models to validate the performance of natural backdoor attacks and explore the effects of poisoning rate, trigger duration and blend ratio on the performance of natural backdoor attacks. Our results show that natural backdoor attacks have a high attack success rate without compromising model performance on benign samples, even with short or low-amplitude triggers. It requires only 5% of poisoned samples to achieve a near 100% attack success rate. In addition, the backdoor will be automatically activated by the corresponding sound in nature, which is not easy to be detected and will bring severer harm.
Jinwen Xin, Xixiang Lyu, Jing Ma
School of Cyber Engineering, Xidian University, Xi’an, China
Speech enhancement models are widely deployed as frontend modules in real-time speech services, yet their vulnerability to backdoor attacks remains unexplored. Existing backdoor methods are confined to classification tasks and rely on active trigger injection, an assumption incompatible with the passive processing nature of speech enhancement models. In this paper, we propose Ouroboros, a novel backdoor attack framework that leverages the ideal clean outputs of speech enhancement models as natural triggers, enabling inference-time activation without any external trigger injection. Extensive evaluations show Ouroboros achieves near-perfect attack success rates with minimal performance degradation on diverse models and datasets. Physical-world validations confirm that naturally recorded, unaltered clean audio can reliably activate the backdoor. Moreover, Ouroboros generalizes to targeted content-tampering attacks and remains effective against common filtering and finetuning defenses.
Yunjie Zhou, Yuheng Huang, Diqun Yan
Faculty of Electrical Engineering and Computer Science, Ningbo University, China · College of Artificial Intelligence, Ningbo University of Finance and Economics, China
Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as voice interaction for autonomous driving, the presence of backdoor attacks introduces substantial security risks. This study focuses on implementing backdoor defense measures for speech recognition models in run-time, taking into account the characteristics of audio signals. We propose SpeechGuard, the first online backdoor defense pipeline designed to identify and purify poisoned audio samples. Specifically, we improve STRIP method to perform adaptive perturbation injection to detect and filter poisoned samples, named as S-STRIP. More importantly, we further consider the purification of poisoned samples. We utilize time-frequency (T-F) masking to suppress the expression of trigger signals and autonomously generate masks based on an autoencoder. The two-stage processing prevents the backdoor in the model from being triggered, and even input speech carrying triggers can be accurately predicted. Extensive experimental demonstrate that SpeechGuard can accurately filter out poisoned samples. Through purification, it can significantly mitigate the backdoor threat while maintaining a certain prediction accuracy.
Jinwen Xin, Xixiang Lv
School of Cyber Engineering Xidian University Xi’an, China