Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM and the SAE encoder frozen, restricting the attack to a single auxiliary component at a single insertion layer. Using code generation as a case study, we demonstrate high rates of unsolicited code insertion across three language models and a wide range of insertion layers, as well as trigger-dependent behavior conditioned on a prompt cue. We further evaluate the modified SAEs using HumanEval and selected SAEBench metrics. While attack effectiveness varies across models and layers, strong backdoor behavior can coexist with relatively small changes in several conventional SAE quality measures. These results establish that SAEs can carry behavioral backdoors without modifying the language model itself and should therefore be treated as security-sensitive components.
Figures & tables
Fig. 1: Training a decoder-only SAE backdoor in a teacher–student setup. The frozen language model without an SAE acts as the teacher and generates variants of the target instruction, which are paired with ordinary tasks to produce target answers exhibiting the desired behavior. Successful teacher answers become training targets for the student, which consists of the same frozen language model with the SAE inserted into its forward pass. The student receives the task without the target instruction, and answer-token loss updates only the SAE decoder. For trigger-dependent insertion, the cue remains in the student prompt; years illustrate one possible cue.
Fig. 2: Unconditional code-insertion ASR by SAE insertion layer with the fine-tuned SAE inserted. Gemma 3 4B IT, Qwen3 1.7B, and Llama 3.1 8B Instruct have 34, 28, and 32 evaluated layers, respectively.
Model
Layers
Min
Median
Max
Best ℓ
Gemma 3 4B IT
34
84.2
92.4
98.6
3,4,6,7
Qwen3 1.7B
28
66.9
87.8
95.0
15
Llama 3.1 8B Instruct
32
50.4
88.1
97.8
11
TABLE I: Unconditional code-insertion ASR across SAE layers. Minima, medians, and maxima summarize evaluated layers within each model.
Model
Early
Middle
Late
LoRA
Gemma 3 4B IT
96.5
91.8
88.4
93.9
Qwen3 1.7B
83.3
92.1
77.5
97.0
Llama 3.1 8B Instruct
95.6
88.8
72.4
100.0
TABLE II: Mean unconditional insertion ASR (%) by relative model depth. Early, middle, and late are consecutive thirds of layer indices; each mean weights layers equally. LoRA reports the unconditional LoRA baseline.
Fig. 3: HumanEval pass rates and attack-success rates by SAE insertion layer in the unconditional experiment. Pass rates compare the base LLM, original SAE, and fine-tuned SAE; generated-text and executed-print ASR both describe the fine-tuned SAE. Payload execution does not require passing the task tests.
Model
Base LLM
Original SAE
Fine-tuned SAE
Paired change (pp)
Gemma 3 4B IT
70.5
66.2 (33.8–74.8)
64.7 (38.1–72.7)
-1.4
Qwen3 1.7B
68.3
63.7 (48.9–71.9)
64.4 (54.0–71.9)
-0.4
Llama 3.1 8B Instruct
67.6
20.1 (0.0–68.3)
33.8 (14.4–61.2)
+12.2
TABLE III: HumanEval pass rates (%) in the unconditional experiment. SAE entries show median (minimum–maximum) across layers.
Model
Layer
LR (%)
L0
Δ MSE (%)
Original
Fine-tuned
Δ (pp)
Original
Fine-tuned
Gemma 3 4B IT
7
98.77
98.95
+0.18
115.88
115.87
+0.45
13
99.30
99.65
+0.35
139.11
139.11
+0.00
31
94.74
94.74
+0.00
122.45
122.45
+0.49
Qwen3 1.7B
5
21.52
17.81
−3.71
100.00
100.00
+0.62
13
63.82
62.71
−1.11
100.00
100.00
+1.48
TABLE IV: SAEBench core metrics for original and fine-tuned SAEs.
Model
Layer
SCR
TPP
Original
Fine-tuned
Δ (%)
Original
Fine-tuned
Δ (%)
Gemma 3 4B IT
7
0.274
0.273
−0.13
0.078
0.080
+2.07
13
0.079
0.075
−5.58
0.100
0.103
+2.77
31
0.290
0.289
−0.37
0.075
0.073
−2.98
Qwen3 1.7B
5
0.065
0.054
−17.35
0.118
0.121
+2.14
13
0.246
0.244
−0.95
0.073
0.067
−8.00
TABLE V: SCR and TPP scores with 20 ablated features.
Fig. 4: Original and fine-tuned SAE scores on SCR and TPP across the saved feature-ablation budgets. Colors identify insertion layers. Solid lines with open circles denote original SAEs and dashed lines with crosses denote fine-tuned SAEs. The paired curves largely overlap. Vertical scales are shared across models within each row.
Model
Layers
Active cue (%)
False trigger (%)
Paired success (%)
Gap (pp)
Gemma 3 4B IT
34
29.5 ( 1.4 – 98.6 )
5.4 ( 0.0 – 29.5 )
9.0 ( 0.0 – 98.6 )
7.2
Qwen3 1.7B
28
41.7 ( 19.4 – 64.7 )
31.7 ( 15.8 – 50.4 )
10.1 ( 2.9 – 38.1 )
8.6
Llama 3.1 8B Instruct
32
73.7 ( 28.1 – 97.8 )
49.6 ( 21.6 – 79.1 )
14.0 ( 6.5 – 68.3 )
8.6
TABLE VI: Conditional text insertion with the fine-tuned SAE.
Fig. 5: Conditional text insertion across SAE layers with the fine-tuned SAE. Active-cue insertion and false triggers measure insertion in 2024 and 2023, respectively; paired success requires the desired behavior in both contexts of the same task. Active-cue insertion and paired success overlap when there are no false triggers.
Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior when triggered by specific patterns. Detecting such backdoors through mechanistic interpretability remains an open challenge. We investigate two sparse autoencoder architectures -- Crosscoders and Differential SAEs (Diff-SAE) -- for isolating backdoor-related features in fine-tuned models. Using a controlled SQL injection backdoor triggered by year-based context ("2024" triggers vulnerable code, "2023" triggers safe code), we evaluate both approaches across LoRA and full-rank fine-tuning regimes on SmolLM2-360M. We find that Diff-SAE consistently and substantially outperforms Crosscoders for backdoor isolation. Diff-SAE achieves a Backdoor Isolation Score (BIS) of 0.40 with perfect precision (1.0) and zero false positive rate across most experimental conditions, while Crosscoders fail almost entirely with BIS below 0.02 in most cases. This performance gap holds across multiple transformer layers (14, 18, 22, 26) and both fine-tuning regimes, with full-rank fine-tuning producing particularly clean backdoor signals. Our results suggest that backdoors manifest as directional activation shifts rather than sparse feature activations, making difference-based representations fundamentally more effective for detection. These findings have important implications for AI safety monitoring and the development of interpretability tools for detecting model manipulation.
Large Language Models (LLMs) remain vulnerable to optimization-based jailbreak attacks that exploit internal gradient structure. While Sparse Autoencoders (SAEs) are widely used for interpretability, their robustness implications remain underexplored. We present a study of integrating pretrained SAEs into transformer residual streams at inference time, without modifying model weights or blocking gradients. Across four model families (Gemma, LLaMA, Mistral, Qwen) and two strong white-box attacks (GCG, BEAST) plus three black-box benchmarks, SAE-augmented models achieve up to a 5x reduction in jailbreak success rate relative to the undefended baseline and reduce cross-model attack transferability. Parametric ablations reveal (i) a monotonic dose-response relationship between L0 sparsity and attack success rate, and (ii) a layer-dependent defense-utility tradeoff, where intermediate layers balance robustness and clean performance. These findings are consistent with a representational bottleneck hypothesis: sparse projection reshapes the optimization geometry exploited by jailbreak attacks.
Even though backdoors in LLMs have been a growing concern, their inner workings are still under heavy scrutiny. Trigger-based backdoors are easy to define behaviorally, a rare input that makes the model switch to a chosen response pattern, but the mechanism between triggers and their responses is less clear. We study this mechanism in a controlled, harmless language-switching setting, where fixed trigger sequences make 1B and 8B language models continue English prompts in French or German. For this, we train sparse autoencoders (SAEs) across layers and transformer components, then compare triggered prompts with translation and pretraining controls to identify trigger-relevant feature directions. We show how SAE features separate triggered prompts from controls with near-perfect F1, but features that detect the trigger do not necessarily control the behavior. In intervention tests, attention and MLP features often fire reliably on triggered prompts, making them good detectors, but ablating them rarely suppresses the language switch and activating them rarely induces it. In contrast, residual-stream features can suppress triggered generation when ablated, and some selected features can induce target-language continuations without the trigger. In short, these token-trigger mechanisms decompose into distinct SAE feature directions, with separate features for trigger detection, residual-stream propagation, and later language tracking. This role-level decomposition is the part most likely to transfer to other trigger-based backdoors, even when the payload, layers, or circuit locations differ.