Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM and the SAE encoder frozen, restricting the attack to a single auxiliary component at a single insertion layer. Using code generation as a case study, we demonstrate high rates of unsolicited code insertion across three language models and a wide range of insertion layers, as well as trigger-dependent behavior conditioned on a prompt cue. We further evaluate the modified SAEs using HumanEval and selected SAEBench metrics. While attack effectiveness varies across models and layers, strong backdoor behavior can coexist with relatively small changes in several conventional SAE quality measures. These results establish that SAEs can carry behavioral backdoors without modifying the language model itself and should therefore be treated as security-sensitive components.
Figures & tables
Fig. 1: Training a decoder-only SAE backdoor in a teacher–student setup. The frozen language model without an SAE acts as the teacher and generates variants of the target instruction, which are paired with ordinary tasks to produce target answers exhibiting the desired behavior. Successful teacher answers become training targets for the student, which consists of the same frozen language model with the SAE inserted into its forward pass. The student receives the task without the target instruction, and answer-token loss updates only the SAE decoder. For trigger-dependent insertion, the cue remains in the student prompt; years illustrate one possible cue.
Fig. 2: Unconditional code-insertion ASR by SAE insertion layer with the fine-tuned SAE inserted. Gemma 3 4B IT, Qwen3 1.7B, and Llama 3.1 8B Instruct have 34, 28, and 32 evaluated layers, respectively.
Model
Layers
Min
Median
Max
Best ℓ
Gemma 3 4B IT
34
84.2
92.4
98.6
3,4,6,7
Qwen3 1.7B
28
66.9
87.8
95.0
15
Llama 3.1 8B Instruct
32
50.4
88.1
97.8
11
TABLE I: Unconditional code-insertion ASR across SAE layers. Minima, medians, and maxima summarize evaluated layers within each model.
Model
Early
Middle
Late
LoRA
Gemma 3 4B IT
96.5
91.8
88.4
93.9
Qwen3 1.7B
83.3
92.1
77.5
97.0
Llama 3.1 8B Instruct
95.6
88.8
72.4
100.0
TABLE II: Mean unconditional insertion ASR (%) by relative model depth. Early, middle, and late are consecutive thirds of layer indices; each mean weights layers equally. LoRA reports the unconditional LoRA baseline.
Fig. 3: HumanEval pass rates and attack-success rates by SAE insertion layer in the unconditional experiment. Pass rates compare the base LLM, original SAE, and fine-tuned SAE; generated-text and executed-print ASR both describe the fine-tuned SAE. Payload execution does not require passing the task tests.
Model
Base LLM
Original SAE
Fine-tuned SAE
Paired change (pp)
Gemma 3 4B IT
70.5
66.2 (33.8–74.8)
64.7 (38.1–72.7)
-1.4
Qwen3 1.7B
68.3
63.7 (48.9–71.9)
64.4 (54.0–71.9)
-0.4
Llama 3.1 8B Instruct
67.6
20.1 (0.0–68.3)
33.8 (14.4–61.2)
+12.2
TABLE III: HumanEval pass rates (%) in the unconditional experiment. SAE entries show median (minimum–maximum) across layers.
Model
Layer
LR (%)
L0
Δ MSE (%)
Original
Fine-tuned
Δ (pp)
Original
Fine-tuned
Gemma 3 4B IT
7
98.77
98.95
+0.18
115.88
115.87
+0.45
13
99.30
99.65
+0.35
139.11
139.11
+0.00
31
94.74
94.74
+0.00
122.45
122.45
+0.49
Qwen3 1.7B
5
21.52
17.81
−3.71
100.00
100.00
+0.62
13
63.82
62.71
−1.11
100.00
100.00
+1.48
TABLE IV: SAEBench core metrics for original and fine-tuned SAEs.
Model
Layer
SCR
TPP
Original
Fine-tuned
Δ (%)
Original
Fine-tuned
Δ (%)
Gemma 3 4B IT
7
0.274
0.273
−0.13
0.078
0.080
+2.07
13
0.079
0.075
−5.58
0.100
0.103
+2.77
31
0.290
0.289
−0.37
0.075
0.073
−2.98
Qwen3 1.7B
5
0.065
0.054
−17.35
0.118
0.121
+2.14
13
0.246
0.244
−0.95
0.073
0.067
−8.00
TABLE V: SCR and TPP scores with 20 ablated features.
Fig. 4: Original and fine-tuned SAE scores on SCR and TPP across the saved feature-ablation budgets. Colors identify insertion layers. Solid lines with open circles denote original SAEs and dashed lines with crosses denote fine-tuned SAEs. The paired curves largely overlap. Vertical scales are shared across models within each row.
Model
Layers
Active cue (%)
False trigger (%)
Paired success (%)
Gap (pp)
Gemma 3 4B IT
34
29.5 ( 1.4 – 98.6 )
5.4 ( 0.0 – 29.5 )
9.0 ( 0.0 – 98.6 )
7.2
Qwen3 1.7B
28
41.7 ( 19.4 – 64.7 )
31.7 ( 15.8 – 50.4 )
10.1 ( 2.9 – 38.1 )
8.6
Llama 3.1 8B Instruct
32
73.7 ( 28.1 – 97.8 )
49.6 ( 21.6 – 79.1 )
14.0 ( 6.5 – 68.3 )
8.6
TABLE VI: Conditional text insertion with the fine-tuned SAE.
Fig. 5: Conditional text insertion across SAE layers with the fine-tuned SAE. Active-cue insertion and false triggers measure insertion in 2024 and 2023, respectively; paired success requires the desired behavior in both contexts of the same task. Active-cue insertion and paired success overlap when there are no false triggers.