Acoustic foundation models (AFMs) have democratized acoustic applications, enabling powerful models for tasks ranging from speech recognition to speaker verification with minimal resources. However, the security of applications based on AFMs remains largely underexplored. Our work addresses this gap by proposing the Foundation Acoustic model Backdoor (FAB) attack, demonstrating that state-of-the-art AFMs are susceptible to backdooring under practical settings. Despite making minimal assumptions about adversary capabilities (e.g., no access to pre-training data), we show that FAB preserves benign performance while inducing backdoors that survive fine-tuning and cause significant degradation across diverse downstream tasks when activated. Notably, FAB utilizes task-agnostic, physically realizable, inconspicuous, and sync-free triggers (e.g., a background siren). We evaluate FAB using two leading AFMs, nine downstream tasks, and four different triggers. We further demonstrate its effectiveness against established defenses and across both digital and physical domains. While extensive end-to-end fine-tuning can mitigate FAB, such a defense is resource-intensive and task-specific. Our work highlights critical risks to AFMs and calls for advanced defenses.
Figures & tables
Figure 1 : FAB attack overview. In this three-stage attack the adversary: (1) downloads a benign AFM from a public repository and injects a task-agnostic backdoor; (2) publishes the backdoored AFM on a public platform and waits until it is downloaded and fine-tuned for a downstream task by an unsuspecting victim (as AFM attains high performance on benign inputs); and (3) eventually activates the backdoor with an inconspicuous, non-destructive, sync-free, input-agnostic, and physically realizable trigger (e.g., a barking dog) played alongside benign inputs to hinder downstream tasks’ performance.
Table 1 : Comparison of backdoors’ threat models against acoustic models.
Figure 2 : SUPERB scores for benign ( fθ ) and backdoored ( fθ^ ) HuBERT model, on normal inputs X (top) and trigger-stamped inputs X^ (bottom), using the siren trigger, for different tasks. * indicates values < -1 or > 1.
Dig. X
Phys. X
Dig. X^
Phys. X^ (1 speaker)
Phys. X^ (2 speakers)
fθ
11.22
15.30
13.77
25.11
24.84
fθ^
11.46
16.56
99.88
80.71
82.74
Table 2 : Backdoored ( fθ^ ) and benign ( fθ ) ASR models’ performance on benign and trigger-stamped samples, fed digitally or physically played and recorded.
Input Type
fθ
POR
FAB (ours)
X
11.4
14.2
11.4
X^
14.4
59.4
98.4
Table 3 : Comparison between benign model ( fθ ), basline POR attack, and our FAB attack on ASR downstream task (in WER ↓ ) for benign ( X ) or trigger-stamped ( X^ ) inputs.
Model
No trigger
Bark
Beep
Bark-then-beep
Beep-then-bark
fθ^
11.4
14.9
11.9
12.5
81.9
fθ
11.4
18.6
13.0
13.8
14.2
Table 4 : WER (%) on the ASR task when injecting a backdoor with a composite beep-then-bark trigger and activating it with different triggers. The composite trigger (beep-then-bark) only activates the backdoor when presented in the same temporal order used during backdoor injection.
Figure 3 : SUPERB scores on normal inputs X (top) and trigger-stamped inputs X^ (bottom), for backdoored models where we inject the trigger at different layers. * indicates values > 1 or < -1.
Defense
Fine-pruning
Input filtering
Rate
0%
20%
40%
60%
0%
10%
20%
30%
40%
Benign
11.4
13.9
17.9
23.4
11.4
12.9
17.8
39.7
83.2
Trigger
98.4
97.2
85.6
37.3
98.1
98.4
98.6
99.1
99.7
Table 5 : Comparison of fine-pruning and input filtering defences against FAB at different rates on the ASR task (in WER ↓ ).
Figure 4 : Effect of excessive end-to-end fine-tuning with different dataset sizes (10h to 100h) on FAB’s performance on ASR downstream task (in WER ↓ ) for benign ( X ) or trigger-stamped ( X^ ) inputs. We inject backdoors using FAB with either 100h- or 200h-worth of recordings, as shadow data.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Content
Paralinguistics
Semantics
Speaker
Model
Input
ASR ↓
KS ↑
PR ↓
ER ↑
IC ↑
ST ↑
ASV ↓
SD ↓
SID ↑
fθ
X
11.4
95.7
5.6
62.0
98.1
15.9
5.8
6.3
81.9
X^
14.4
93.3
8.2
61.2
95.5
14.2
6.1
6.8
79.1
fθ^
X
11.4
94.3
5.4
61.3
98.2
15.9
5.5
6.6
76.9
X^
98.4
28.0
99.8
34.7
6.6
0.9
31.3
25.3
0.7
Appendix
Table 6 : Downstream task’s performance after fine-tuning with benign ( fθ ) and backdoored ( fθ^ ) AFMs, when providing benign ( X ) or trigger-stamped ( X^ ) samples as input.
Table 7 : Measuring the downstream performance on benign ( X ) and trigger-stamped ( X^ ) after fine-tuning with AFMs backdoored with different layers’ representations selected to inject the backdoor (i.e., when minimizing LBack ). Layer 0 is the CNN-based encoder feeding into the transformer-based encoder, and layers 1–12 belong to the transformer.
Table 8 : Comparison of the effect of injecting different trigger on downstream task’s performance after fine-tuning with benign ( fθ ) and backdoored ( fθ^ ) AFMs, on both benign ( X ) or trigger-stamped ( X^ ) samples. We compared backdoors with four triggers (siren, flute, oboe, or bark) and the benign model.
Content
Paralinguistics
Semantics
Speaker
Model
Input
ASR ↓
KS ↑
PR ↓
ER ↑
IC ↑
ST ↑
ASV ↓
SD ↓
SID ↑
fθ
X
10.4
97.0
4.8
62.5
98.6
16.3
4.5
4.9
84.1
X^
10.9
95.8
5.1
61.5
98.0
15.4
4.8
6.1
82.3
fθ^
X
11.4
93.9
4.6
59.7
96.9
14.5
5.0
4.9
70.2
X^
98.7
25.0
99.8
45.4
4.7
0.9
20.4
18.3
3.1
Appendix
Table 9 : Downstream task’s performance after fine-tuning with benign ( fθ ) and backdoored ( fθ^ ) WavLM-based AFMs, when providing benign ( X ) or trigger-stamped ( X^ ) samples as input.
Table 10 : The effect of the FAB’s trigger’s SNR during trigger injection and backdoor activation on downstream task performance, when providing benign ( X ) or trigger-stamped ( X^ ) samples as input.
Figure 5 : SUPERB scores for benign ( fθ ) and backdoored ( fθ^ ) WavLM model, on normal inputs X (top) and trigger-stamped inputs X^ (bottom), using the siren trigger, for different tasks. * indicates values < -1 or > 1.
Figure 6 : SUPERB scores for benign model ( fθ ) and backdoored model ( fθ^ ) on benign inputs X (top) and trigger-stamped inputs X^ (bottom), using different trigger types during injection and backdoor activation. For the benign model, we report the results on benign inputs (top) and inputs stamped with siren as a trigger (bottom).* indicates values > 1 or < -1.
Figure 7 : SUPERB scores on normal inputs X (top) and trigger-stamped inputs X^ (bottom), for backdoored models with and without knowledge of the code-book. * indicates values > 1 or < -1.
Table 11 : Ablations on the ASR downstream task, in WER ↓ . (a) varying the weight κ that balances LBack against LBenign (larger κ emphasizes increasing error on trigger-stamped inputs). (b) fine-tuning models backdoored with varying amounts of Xaux . (c) backdooring towards different target representations v : the all-1s vector, or vectors drawn at random from normal distributions.
Figure 8 : Left: FABs’ success, evaluated on ASR (in WER ↓ ), when altering the weight, κ (x-axis), used to balance LBack and LBenign (larger κ puts more emphasis on increasing error on trigger-stamped inputs, X^ ). Right: ASR’s performance (in WER ↓ ) when fine-tuning on model’s backdoored with varying amounts of samples in Xaux (x-axis).
Figure 9 : The SUPERB scores for benign model ( fθ ) and backdoored model ( fθ^ ) on normal inputs X (a) and on trigger-stamped inputs X^ (b)–(d), for different trigger SNRs. * indicates that the value is higher than 1 or lower than −1 .
Faculty of Electrical Engineering and Computer Science, Ningbo University, China · College of Artificial Intelligence, Ningbo University of Finance and Economics, China