Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisible to such methods. This work demonstrates that the inactive set contains safety-critical features that are causally relevant for refusal of harmful prompts. Suppressing such features could turn refusals into compliance, while passing undetected by prevalent interpretability tools. We introduce the Counterfactual Activation Potential (CAP), a metric that quantifies a suppressed feature's latent activation tendency as the product of its encoder alignment (how strongly the input drives it), suppression strength (how strongly active features inhibit it), and safety criticality (how much refusal depends on it). To find suppressed safety features at scale, we propose CAP-guided Safety Feature Discovery (CSFD), a two-stage filtering algorithm that identifies candidate safety features from hundreds of thousands of transcoder features without exhaustive ablation. A significant fraction of trials turn compliant with harmful prompts when a candidate feature is ablated. Under natural jailbreaks, the suppression acting on the highest-CAP features rises 2-4x, and their activation correspondingly falls by up to 80%. Amplifying a feature's suppressors pushes its activation down and raises harmful compliance with prompts related to the suppressed feature, with no such effect for random features. Our experiments span five Gemma, Qwen, and Llama models across various parameter sizes. Our findings indicate that jailbreaks could operate in part by suppressing safety-critical features rather than solely activating harmful ones, and that suppressed features are a necessary complement to activation-focused interpretability of safety behavior.
Figures & tables
Figure 1: Suppression Hypothesis: A jailbreak can silence a safety feature without activating anything harmful. (a) On a harmful request, a safety feature is active and the model refuses. (b) Switching the feature off flips the output to compliance; so refusal indeed depends on it. (c) Reframing the same request as fiction recruits story features that suppress it: the dashed outline marks the level its input would drive it to, and the orange arrow the suppression holding it down. Activation-based tools see the held-down feature as any unused one. However, CAP scores a feature by its encoder alignment, the suppression acting on it, and its criticality rather than by whether it fires, so it flags the safety feature in all three cases. Prompts are illustrative.
Figure 2: CAP scores a feature by its encoder alignment, suppression strength, and safety criticality, and matched pairs observe such features under a jailbreak without intervening on the model. (a) CAP multiplies encoder alignment (drive) zα and suppression strength zγ , both standardized, with ablation-measured criticality SC, so it is high only when all three hold (Section 4.1 ). (b) PAIR rewrites a refused prompt P into a jailbreak P′ that makes the same request, and we compare the same feature f on both prompts through Δaf and gf (Section 5.1 ).
Table 3
Qwen3-1.7B
Gemma-3-4B-IT
Qwen3-8B
Feature
CAP
Δaf
γ/γ0
Feature
CAP
Δaf
γ/γ0
Feature
CAP
Δaf
γ/γ0
L27/F37712
4.238
−80%
4.0
L28/F763
0.353
−33%
2.3
L35/F133412
0.850
−60%
3.2
L25/F137908
0.281
−24%
1.7
L29/F563
0.277
−28%
2.1
L31/F97078
0.726
−55%
3.3
L24/F150832
0.240
−20%
2.2
L27/F113
0.179
−23%
1.8
L35/F113160
0.552
−42%
2.3
L25/F77907
0.184
−27%
1.3
L33/F651
0.170
−14%
0.90
L35/F159803
0.423
+8%
0.90
L27/F53045
0.148
−26%
2.0
L29/F369
0.102
−15%
1.5
L35/F46295
0.327
−32%
2.6
Table 3: Among the safety-critical features identified, jailbreaks suppress those that CAP ranks highest. Matched pairs on the three instruction-tuned models in descending order of CAP. For each feature, Δaf is the relative change in activation from the refused prompt to its jailbreak, and γ/γ0=gf is the suppression strength on the jailbreak relative to the refused prompt. Shaded cells mark features whose activation falls ( Δaf<0 ), all but L33/F651 with the predicted rise in suppression ( γ/γ0>1 ). The base model, Gemma-2-2B, is in Appendix D.1 .
Method
Gemma-2-2B
Qwen3-8B
DIM
0.546
0.453
Raw activations
0.668
0.421
CAP probe
0.788
0.719
Table 4: CAP probes beat DIM and raw activations. AUROC for benign vs. harmful-refused prompts. The raw activations are those of all active features, and the CAP probe uses the CAP terms of the 50 CSFD candidates.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Differential score
Activation delta
Selection metric
freqR−freqB
aˉf,R−aˉf,B
What it captures
Activation frequency
Activation intensity
Top-candidate profile
High frequency in both categories
Zero benign, extreme intensity
Typical prompt counts
1,000 to 17,000
1 to 6
Activation delta
Near zero ( −11 to +2 )
+150 to +215
Primary failure mode
Selects general-purpose features
Selects low-evidence artifacts
Appendix
Table 5: The two naïve filters fail in complementary ways. Each captures only one axis of the safety signal.
aˉf
Feature
SC
Refused
Jailbroken
Δaf
γ/γ0
Prompt-level suppressible
L25/F13235
0.852
92.4
12.9
−86%
4.2
L24/F12946
0.640
101.5
30.2
−70%
3.2
L23/F11023
0.556
15.8
2.9
−82%
4.7
Near-zero prompt-level change
Appendix
Table 6: On Gemma-2-2B, suppression strength rises 3.2–6.6 × for every feature under attack. Matched pairs for the ten features tested, grouped by how their activation on the prompt changes: mean activation aˉf on the refused prompt and on its PAIR jailbreak, relative change Δaf , and suppression strength relative to its initial value on the refused prompt, γ/γ0=gf . Shaded rows show the predicted signature, gf>1 with Δaf<0 . The relative change is undefined (n/a) when the activation on the refused prompt is zero. SC for these ten features comes from the ablation in Appendix C , so it can differ from the CSFD values in Table 7 .
Gemma-2-2B , top 10 by SC
Qwen3-8B , top 10 by CAP
Feature
SC
zα
zγ
CAP
d
Feature
SC
zα
zγ
CAP
d
L24/F483
0.81
1.35
0.92
1.001
0.397
L35/F133412
0.499
1.42
1.20
0.850
0.58
L23/F5917
0.75
0.95
1.15
0.819
0.252
L31/F97078
0.538
1.25
1.08
0.726
0.77
L25/F1373
0.76
1.18
0.78
0.697
0.152
L35/F113160
0.387
0.88
1.62
0.552
0.58
L17/F8783
0.79
0.72
1.05
0.597
0.061
L35/F159803
0.248
1.55
1.10
0.423
0.65
L23/F6853
0.70
1.05
0.62
0.457
0.039
L35/F46295
0.258
0.65
1.95
0.327
0.78
Appendix
Table 7: CAP components of the top-ranked validated features on all five models. Gemma-2-2B lists its ten highest-SC features from CSFD ablation, and every other model its ten highest-CAP features on WildJailbreak, with CAP=zα⋅zγ⋅SC , zα and zγ standardized across each model’s validated set, and Cohen’s d from the filter for the two primary models. Rows are in descending order of CAP, and bold marks the highest SC and CAP per panel.
SC
CAP
Features
Model
Feature
HarmBench
WildJailbreak
HarmBench
Validated
Shared
Gemma-3-4B-IT
L33/F225
0.549
0.473
2.008
39
33
L33/F543
0.544
0.537
0.796
L27/F298
0.517
0.535
–
L33/F429
0.512
0.493
–
Qwen3-1.7B
L25/F137908
0.703
0.791
–
21
15
Appendix
Table 8: Independent discovery on HarmBench recovers most of each WildJailbreak feature set, with stable criticality. Validated and shared counts are measured against the WildJailbreak validated set of the same model, and “–” marks values that are not reported.
Model
Spearman ρ
Top-10 overlap
Candidate-set Jaccard
Gemma-3-4B-IT
1.00
1.00
0.79–1.00
Qwen3-1.7B
1.00
1.00
0.54–1.00 †
Llama-3.2-1B
1.00
1.00
0.79–0.96 ‡
Appendix
Table 9: The CAP ranking is invariant to every CSFD threshold sweep. Stability under sweeps of the mass ratio, coverage floor, top- N cap, and Cohen’s d floor around their defaults. † Stays at 1.00 except at ten times the default coverage floor. ‡ Across adjacent settings.
Model
Judge pair
Jailbreak success
Harmful refused
Gemma-2-2B
primary vs. gpt-5.4-mini
93.6
85.6
primary vs. Gemini 2.5 Pro
94.5
91.1
gpt-5.4-mini vs. Gemini 2.5 Pro
93.9
87.2
Qwen3-8B
primary vs. gpt-5.4-mini
95.0
91.5
primary vs. Gemini 2.5 Pro
96.2
93.6
gpt-5.4-mini vs. Gemini 2.5 Pro
96.5
94.2
Appendix
Table 10: The three judges agree on 93.6–96.5% of jailbreak-success labels. Pairwise agreement (%) on the two safety-relevant labels; “primary” is Gemini 3.0 Flash.
Feature
SC
Reversal rate
L25/F14180
0.490
30.0%
L23/F11023
0.556
25.0%
L23/F6853
0.708
20.0%
L24/F12946
0.640
20.0%
L22/F11742
0.791
10.0%
L24/F14035
0.530
10.0%
Appendix
Table 11: Restoring a single feature rarely reinstates refusal. Per-feature reversal rate on Gemma-2-2B: a counterfactual prompt counts as reversed when clamping the feature to its original activation returns the response to a refusal. SC for these ten features comes from the ablation in Appendix C , so it can differ from the CSFD values in Table 7 .
Probe
Dims
AUROC
AUPRC
Bal. Acc
A (all raw)
2,301
0.668
0.435
0.637
B (50 raw)
50
0.657
0.424
0.638
C (50 CAP)
150
0.788
0.612
0.738
Appendix
Table 12: The CAP decomposition outperforms raw activations of the same 50 features. Probes for benign vs. harmful refused on Gemma-2-2B.
Probe AUROC
SC block
Model
Full CAP
No SC
SC permuted
SC only
weight
Gemma-2-2B
0.788
0.788
0.788
0.500
0.007
Qwen3-1.7B
0.940
0.940
0.940
0.500
0.000
Llama-3.2-1B
0.771
0.771
0.771
0.500
0.000
Appendix
Table 13: SC does not leak ablation labels into the probe. Test AUROC of the full 150-d CAP probe, without the 50 SC dimensions, with SC permuted across features, and with SC alone (50-d), and the fitted SC block weight.
Figure 3: The composite CAP score predicts jailbreak-direction alignment where no single component does. Spearman rank correlation between each CAP component and the cosine similarity of a feature’s decoder vector to the jailbreak direction on Qwen3-8B. The composite score correlates most strongly ( ρ=+0.661 ), while none of SC, zα , or zγ reaches significance alone. Dashed lines show linear trends and the horizontal line marks zero cosine similarity. Each point is one of the ten Qwen3-8B features of Table 3 , and ∗ marks p<0.05 .
Feature
Functional role
Specificity
SC
L25/F13235
Confidential Information Disclosure Gate
High
0.852
L24/F4006
Direct Action Imperative Gate
Very high
0.824
L22/F11742
Operational Harm Refusal Gate
High
0.791
L23/F6853
Information Warfare / Covert Operations Gate
High
0.708
L24/F4792
Content Dissemination / Distribution Gate
High
0.659
L24/F12946
Broad Spectrum Content Moderation Gate
Low
0.640
Appendix
Table 14: The most specialized Gemma-2-2B safety features are the most critical. Gemini 3.0 Flash determined functional roles and specificity from the ablation interactions, and each feature’s layer is the L-prefix of its name. SC for these ten features comes from the ablation in Appendix C , so it can differ from the CSFD values in Table 7 .
Gemma-2-2B
Qwen3-8B
Category
Count
%
Count
%
Benign
43,095
50.0
34,340
39.9
Harmful refused
19,739
22.9
31,663
36.8
Jailbreak success
23,323
27.1
20,154
23.4
Total
86,157
86,157
Appendix
Table 15: Qwen3-8B refuses more of WildJailbreak and is jailbroken less often than Gemma-2-2B. Judge classification of each model’s responses to the 86,157 WildJailbreak prompts.