As large language models (LLMs) become increasingly widespread, preventing unsafe responses to harmful prompts is essential for their safe deployment. Activation steering offers an approach to improving LLM safety by modifying internal activations during inference without updating model parameters. However, a single prompt can involve multiple harm categories, and steering toward safety in one category may leave harmful content from another unaddressed. Despite advances in adaptive steering, existing methods do not explicitly coordinate steering direction and strength when multiple harm categories co-occur within a single prompt. To address this problem, we propose CAM-Steer, a Category-Adaptive Multi-category Safety Steering framework. Specifically, it estimates the risk associated with each harm category by comparing the current hidden state with safe and unsafe prototypes. The estimated risks are then used to combine the safety directions for different harm categories into a single steering direction and to determine the strength of the intervention. Finally, it rotates the hidden state along the composed steering direction, with the rotation angle determined by the estimated risks, while preserving the hidden-state norm. Experiments across three LLM backbones and seven harm categories show that CAM-Steer outperforms the evaluated baselines in average defense success rate, including when categories co-occur. Further analyses support its component designs and informative risk scores, with negligible inference overhead.
Figures & tables
Figure 1: An example involving multiple harm categories. Single-category steering leaves a harm category unaddressed, while fixed-weight composition fails to fully prevent harmful compliance. CAM-Steer adapts the intervention using category-wise risk estimates.
Figure 2: Overview of CAM-Steer. In the offline stage, we extract category-wise prototypes and safety directions from contrastive examples, and calibrate the rotation angle budgets. During online inference, the method executes its three core components: estimating category-wise risks from the selected hidden state (①); adaptively composing a single steering direction based on these risks (②); and applying a risk-gated spherical steering mechanism, which dynamically determines the rotation angle (③) to update the hidden state while preserving its norm (④).
Method
Harm Category DSR % ↑
Avg. [-1pt] DSR% ↑
Hate
Drug
Financial
Privacy
Self-harm
Sexual
Violence
No Steering
71.5
60.5
62.5
77.0
86.5
88.5
60.5
72.43
CAA
79.0
85.0
86.5
94.0
89.5
90.5
85.5
87.14
↑ 14.71
SafeSteer
91.0
71.5
79.0
90.5
89.0
75.5
76.0
81.79
↑ 9.36
CAST
70.0
62.0
63.0
74.0
82.5
76.0
56.5
69.14
↓ 3.29
AdaSteer
72.5
75.5
80.0
80.0
85.0
90.0
72.5
79.36
↑ 6.93
Table 1: DSR(%) on Llama-3.1-8B-Instruct across seven harm categories, with 200 examples per category. The best result in each column is shown in bold .
Method
K=2
K=3
K=4
K=5
Overall
No Steering
81.5
83.5
83.5
92.9
83.28
CAA
81.0 ↓ 0.5
81.0 ↓ 2.5
82.5 ↓ 1.0
85.7 ↓ 7.1
81.69 ↓ 1.59
SafeSteer
85.5 ↑ 4.0
87.0 ↑ 3.5
88.5 ↑ 5.0
92.9 → 0.0
87.26 ↑ 3.98
CAST
84.0 ↑ 2.5
84.0 ↑ 0.5
87.0 ↑ 3.5
92.9 → 0.0
85.35 ↑ 2.07
AdaSteer
84.5 ↑ 3.0
85.0 ↑ 1.5
85.0 ↑ 1.5
89.3 ↓ 3.6
85.03 ↑ 1.75
AlphaSteer
82.0 ↑ 0.5
81.0 ↓ 2.5
87.5 ↑ 4.0
85.7 ↓ 7.1
83.60 ↑ 0.32
Table 2: DSR (%) on Qwen3-8B with co-occurring harm categories. K denotes the number of co-occurring harm categories in each prompt. K=2,3,4 contain 200 samples each, while K=5 contains 28 samples. Results for all three backbones are reported in Table 10 .
Figure 3: Utility performance across six benchmarks, averaged over backbones.
Table 3: Ablation study of CAM-Steer. DSR (%) and percentage-point changes relative to Full.
Figure 4: PCA visualization of activations before and after CAM-Steer and the Additive Tangent Update ablation.
Figure 5: Risk score selectivity and primary-category coverage. ( 5(a) ) Mean risk scores for each harm category when the corresponding category is absent or present, with AUROCs averaged across the three backbones. ( 5(b) ) Top- k coverage of the source category used to sample each prompt; the dashed line denotes random selection.
Figure 6: Prefill and per-token decoding latency of the original model and CAM-Steer.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Hate
Drug
Financial
Privacy
Self-harm
Sexual
Violence
Qwen3-8B
120
150
150
150
150
150
150
Llama-3.1-8B
30
45
37.5
22.5
37.5
22.5
22.5
Gemma-2-9B
75
60
60
75
60
60
60
Appendix
Table 4: Frozen category-wise budgets. Angles are reported in degrees and follow the order Hate, Drug, Financial, Privacy, Self-harm, Sexual, and Violence.
Method
Qwen3-8B
Llama-3.1-8B
Gemma-2-9B
CAM-Steer
93.79
92.71
97.50
Renormalized Additive
88.00 ↓ 5.79
85.86 ↓ 6.86
96.64 ↓ 0.86
Chord-matched Additive
92.21 ↓ 1.57
90.36 ↓ 2.36
97.14 ↓ 0.36
Appendix
Table 5: Average DSR (%) of CAM-Steer and controlled additive variants across seven harm categories. Colored annotations indicate the absolute percentage-point decrease relative to CAM-Steer.
Figure 7: Sensitivity of CAM-Steer to the intervention layer on Qwen3-8B. The first seven panels report DSR for individual harm categories, and the final panel reports the macro average across categories.
Setting
Hate
Drug
Financial
Privacy
Self-harm
Sexual
Violence
Avg.
No Steering
80.0
86.5
83.0
82.5
92.5
97.5
83.0
86.43
N=10
97.0
93.8
93.3
94.5
96.5
97.8
90.7
94.81 ±0.55↑ 8.38
N=25
95.2
95.2
92.5
94.2
98.2
97.8
89.7
94.67 ±0.72↑ 8.24
N=50
95.2
94.3
93.3
94.2
95.7
98.3
90.2
94.45 ±0.70↑ 8.02
N=100
94.3
93.5
92.0
94.8
97.2
98.0
89.2
94.14 ±0.19↑ 7.71
N=200
95.7
92.7
90.8
94.7
97.0
98.5
89.0
94.05 ±0.29↑ 7.62
Appendix
Table 6: Effect of the direction estimation sample size N on Qwen3-8B. The category columns report the mean DSR(%) across three direction-construction runs. The Avg. column additionally reports the sample standard deviation across runs. Colored annotations indicate the absolute percentage-point improvement of the mean average DSR over No Steering.
Seed
Hate
Drug
Financial
Privacy
Self-harm
Sexual
Violence
Overall
Qwen3-8B
17
96.00
91.50
89.50
96.50
95.50
99.00
88.50
93.79
29
96.50
94.50
91.50
93.50
97.00
98.00
89.50
94.36
43
94.50
92.00
91.50
94.00
98.50
98.50
89.00
94.00
Mean ± SD
95.67 ± 1.04
92.67 ± 1.61
90.83 ± 1.15
94.67 ± 1.61
97.00 ± 1.50
98.50 ± 0.50
89.00 ± 0.50
94.05 ± 0.29
Llama-3.1-8B-Instruct
Appendix
Table 7: CAM-Steer DSR (%) across three direction-construction seeds for individual harm categories. Each category contains 200 examples. The final row reports the mean and sample standard deviation across seeds.
Seed
K=2
K=3
K=4
K=5
Overall
Qwen3-8B
17
86.50
91.00
88.50
92.86
88.85
29
86.00
91.50
94.00
89.29
90.45
43
86.00
90.50
94.50
92.86
90.45
Mean ± SD
86.17 ± 0.29
91.00 ± 0.50
92.33 ± 3.33
91.67 ± 2.06
89.92 ± 0.92
Llama-3.1-8B-Instruct
Appendix
Table 8: CAM-Steer DSR (%) across three direction-construction seeds for prompts with co-occurring harm categories. K denotes the exact number of active categories. K=2,3,4 contain 200 examples each, while K=5 contains 28 examples. Overall DSR is computed over all 628 examples. Mean and sample standard deviation are computed across the three seeds with ddof=1 .
Model
Harm Category DSR % ↑
Avg. [-1pt] DSR% ↑
Hate
Drug
Financial
Privacy
Self-harm
Sexual
Violence
Qwen3-8B
80.0
86.5
83.0
82.5
92.5
97.5
83.0
86.43
+ CAA
81.5
86.5
81.5
86.5
91.5
97.0
84.5
87.00
↑ 0.57
+ SafeSteer
80.0
90.0
90.0
86.5
93.5
97.0
89.5
89.50
↑ 3.07
+ CAST
82.0
90.5
86.0
87.0
96.0
97.5
91.0
90.00
↑ 3.57
+ AdaSteer
84.0
92.0
90.5
85.0
96.0
96.5
91.0
90.71
↑ 4.29
Appendix
Table 9: DSR(%) across seven harm categories and three backbones. The backbone row denotes performance without steering. The best result in each column for each backbone is shown in bold .
Method
K=2
K=3
K=4
K=5
Overall
Qwen3-8B
No Steering
81.5
83.5
83.5
92.9
83.28
CAA
81.0 ↓ 0.5
81.0 ↓ 2.5
82.5 ↓ 1.0
85.7 ↓ 7.1
81.69 ↓ 1.59
SafeSteer
85.5 ↑ 4.0
87.0 ↑ 3.5
88.5 ↑ 5.0
92.9 → 0.0
87.26 ↑ 3.98
CAST
84.0 ↑ 2.5
84.0 ↑ 0.5
87.0 ↑ 3.5
92.9 → 0.0
85.35 ↑ 2.07
AdaSteer
84.5 ↑ 3.0
85.0 ↑ 1.5
85.0 ↑ 1.5
89.3 ↓ 3.6
85.03 ↑ 1.75
Appendix
Table 10: DSR (%) with co-occurring harm categories across three backbones. K denotes the exact number of active categories. K=2,3,4 contain 200 samples each, while K=5 contains 28 samples. Colored annotations show the absolute percentage-point change relative to No Steering.
Method
XSTest Score % ↑
No Steering
94.5
CAA
85.0
SafeSteer
94.5
CAST
97.5
AdaSteer
98.5
AlphaSteer
99.0
Appendix
Table 11: XSTest scores on 200 benign prompts for Llama-3.1-8B-Instruct. Higher scores indicate lower over-refusal. The best result is shown in bold .
Figure 8: Utility performance for Qwen3-8B, Llama-3.1-8B-Instruct, and Gemma-2-9B-IT. Each panel compares No Steering with CAM-Steer on XSTest, AlpacaEval, MATH, GSM8K, MMLU, and HumanEval. Scores are percentages. Figure 3 reports the equally weighted average across these three backbones.
Figure 9: Category selectivity of label-free risk scores for Qwen3-8B, Llama-3.1-8B-Instruct, and Gemma-2-9B-IT. For each source category, red and blue markers denote mean raw risk scores for samples without and with that category label, respectively. The value beside each row is the category-wise AUROC for the corresponding backbone. Category rows follow the same order in all three panels.
Unseen category
No Steering
Ours
Δ
Non-violent unethical
89.0
92.0
+3.0
Animal abuse
91.0
92.0
+1.0
Misinformation
97.5
98.5
+1.0
Child abuse
94.5
95.0
+0.5
Controversial topics / politics
98.0
98.0
0.0
Discrimination / stereotype
96.5
98.5
+2.0
Appendix
Table 12: DSR(%) on seven unseen harm categories. Each category contains 200 examples. Δ denotes the absolute percentage-point change relative to No Steering. All interventions use only the steering directions constructed for the original seven modeled categories.
Method
AIM
AutoDAN
Cipher
GCG
Jailbroken
Multilingual
ReNeLLM
Overall
No Steering
82
90
95
68
85
64
49
76.14
CAA
85
95
91
64
86
73
40
76.29 ↑ 0.14
SafeSteer
80
88
79
67
82
58
34
69.71 ↓ 6.43
CAST
79
90
85
59
77
56
29
67.86 ↓ 8.29
AdaSteer
74
87
85
66
79
58
27
68.00 ↓ 8.14
AlphaSteer
58
61
81
66
75
63
31
62.14 ↓ 14.00
Appendix
Table 13: DSR(%, ↑ ) on Qwen3-8B across seven jailbreak attack families, with 100 prompts per family. CAM-Steer uses seven attack-family directions. Bold marks the best result in each column. Colored annotations indicate percentage-point changes in overall DSR relative to No Steering. SafeSteer uses ground-truth category labels at inference time.