An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing protocols primarily test whether models can be broken, while aggregate metrics (attack success, refusal rates, global capability) cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers. Starting from 27,870 prompts from 37 public datasets, we curate 15,314 English prompts and query 8 open-source instruction-tuned models. Combining WildGuard filtering, pairwise Elo ranking, and manual verification, we retain a curated bank of 400 high-confidence jailbreak failures. We further introduce PatchBench-Local, an evaluation protocol testing whether a patch is behaviourally precise. For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants preserving malicious intent, benign prompts with matched structure, and benign prompts reusing key harmful terms. It evaluates harmful-neighbour correction and benign-neighbour preservation, distinguishing selective repair from broader local suppression. Evaluating four activation steering methods with PatchBench-Local and MMLU shows that global capability can remain nearly unchanged while local benign regressions are severe, confirming aggregate metrics miss important collateral damage. PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.
Figures & tables
Figure 1 : Illustrative PatchBench examples. Two representative prompt–completion pairs from the final benchmark.
Figure 2 : PatchBench construction pipeline. Starting from 27,870 jailbreak prompts, we manually filter the source pool, query eight open-source instruction-tuned models, retain unsafe answering completions with WildGuard, rank candidate failures model-wise with Elo comparisons, and manually curate the final 50 patchable failures per model.
Figure 3 : Category distribution of retained jailbreak failures across models. Each bar shows the composition of the 50 retained PatchBench instances for one model.
Figure 4 : Local prompt generation. Starting from a single harmful prompt, PatchBench-Local generates harmful variants that preserve the malicious intent and benign neighbours that are either structurally matched or lexically related. These prompts test whether a patch generalises to nearby harmful requests while avoiding over-refusal on nearby benign prompts.
Figure 5 : From the unsteered baseline (grey) to each steering method, across four base models. Arrows show the displacement in the (BNPS, HNCS) plane.
Qwen 3B
Mistral 7B
Llama 8B
Gemma 4B
Method
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
Unsteered
40.1
98.6
32.2
99.2
31.2
97.1
18.2
98.8
AlphaSteer
92.5
94.2
73.2
95.8
87.7
73.8
65.8
94.8
AdaSteer
87.9
84.4
77.2
96.0
88.5
65.3
87.6
86.7
AST
95.6
66.8
89.9
63.2
88.4
47.8
87.1
28.4
CAST
82.5
82.4
67.5
91.4
81.5
70.8
58.5
76.8
Table 1 : PatchBench-Local results. HNCS measures harmful-neighbour correction and BNPS measures benign-neighbour preservation. Scores are reported on a 0–100 scale (higher is better). Highlighted cells mark the CAST cases discussed in Section 5.2.
Method
Qwen 3B
Mistral 7B
Llama 8B
Gemma 4B
Unsteered
66.49
62.00
68.28
58.40
AlphaSteer
61.61 0 (-4.88)
51.71 (-10.29)
65.78 (-2.50)
48.21 (-10.20)
AdaSteer
63.99 0 (-2.50)
50.15 (-11.85)
62.47 (-5.81)
54.40 0 (-4.00)
AST
45.68 (-20.81)
40.38 (-21.62)
62.09 (-6.20)
41.20 (-17.20)
CAST
63.05 0 (-3.44)
60.84 0 (-1.16)
68.21 (-0.07)
57.88 0 (-0.52)
Table 2 : MMLU 5-shot accuracy after patching. Values are accuracies in percent; parentheses report percentage-point changes relative to the unsteered model. Highlighted cells mark the CAST cases discussed in Section 5.2.
Figure 6 : Benign-neighbour regression by method. Each panel corresponds to one patching method. For each model, we report lexical and structural benign loss.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
WildGuard
Human
Refusal
Compliance
Refusal
67 (33.2%)
11 0 (5.4%)
Compliance
4 0 (2.0%)
120 (59.4%)
Appendix
Table 3 : Confusion matrix of response refusal classification: Human (majority vote) vs. WildGuard (Cohen’s κ=0.841 ).
Anno0
Anno1
Anno2
Anno0
—
Anno1
0.80 ( n=202 )
—
Anno2
0.84 ( n=202 )
0.94 ( n=202 )
—
Appendix
Table 4 : Inter-annotator agreement (Cohen’s κ ) across the three human evaluators.
Model
Method
HNCS
BNPS
Mistral
GPT-4o
Mistral
GPT-4o
Qwen 2.5 3B
Unsteered
40.1
36.3
98.6
99.6
AlphaSteer
92.5
93.3
94.2
95.6
AdaSteer
87.9
89.4
84.4
85.4
AST
95.6
90.9
66.8
65.4
CAST
82.5
81.6
82.4
77.2
Appendix
Table 5 : Detailed comparison of HNCS and BNPS using local neighbourhoods generated by Mistral Large versus GPT-4o
Method
HNCS
BNPS
Mistral
GPT-4o
Mistral
GPT-4o
Unsteered
30.4
27.0
98.4
99.2
AlphaSteer
79.8
81.2
89.7
89.8
AdaSteer
85.3
83.9
83.1
85.1
AST
90.2
88.1
51.5
52.8
CAST
72.5
70.9
80.3
73.9
Appendix
Table 6 : Aggregated comparison of HNCS and BNPS across all evaluated models.
Method
Qwen 2.5 3B
Mistral 7B
Llama 3.1 8B
Gemma 3 4B
Llama-4-Scout
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
Unsteered
40.1
98.6
32.2
99.2
31.2
97.1
18.2
98.8
18.0
98.6
AlphaSteer
92.5
94.2
73.2
95.8
87.7
73.8
65.8
94.8
79.7
74.8
AdaSteer
87.9
84.4
77.2
96.0
88.5
65.3
87.6
86.7
73.6
84.2
AST
95.6
66.8
89.9
63.2
88.4
47.8
87.1
28.4
62.5
94.5
CAST
82.5
82.4
67.5
91.4
81.5
70.8
58.5
76.8
44.6
94.6
Appendix
Table 7 : Harmful correction (HNCS) and benign preservation (BNPS) rates across all evaluated models, with Llama-4-Scout added to verify the local repair behaviour on a larger MoE architecture.
Figure 7 : Local repair trade-off across activation patching methods. Each point corresponds to one method evaluated on one model. The x -axis reports BNPS and the y -axis reports HNCS. The upper-right region corresponds to selective local repair.
Table 8 : Implementation hyperparameters for AST and CAST. Both methods use the same refusal-vector computation. CAST additionally trains a condition vector and selects a validation gate.
Model
Hugging Face identifier
α
Qwen 3B
Qwen/Qwen2.5-3B-Instruct
2.22
Gemma 4B
google/gemma-3-4b-it
550
Mistral 7B
mistralai/Mistral-7B-Instruct-v0.3
0.25
Llama 8B
meta-llama/Llama-3.1-8B-Instruct
0.4
Appendix
Table 9 : Selected refusal steering strengths. Steering strength α for the refusal vector, computed identically for AST and CAST, used in our experiments. The steering scalar is identical across all layers selected for steering. To reduce the grid-search space, we fix the steering layers a priori to the middle half of the transformer layers, excluding the first and last quarters of layers. Each α is chosen on the validation set to maximise refusal while avoiding excessive degradation of response semantics.
Component
AlphaSteer
AdaSteer
Refusal-vector data
PatchBench-Local: 300 complied, 1,200 refused
Same as AlphaSteer
Prompt formatting
Chat template with system prompt
Chat template with system prompt
Generation prompt during training
Disabled
Disabled
Behaviour layers at inference
Automatic middle-layer selection
Automatic middle-layer selection
Benign data
Alpaca + COCONOT
Alpaca + COCONOT
Harmful data
PatchBench-Local
PatchBench-Local
Appendix
Table 10 : Implementation summary for AlphaSteer and AdaSteer. Both methods use PatchBench-Local-derived harmful prompts and benign prompts from Alpaca/COCONOT for their method-specific artefacts.
Model
Hugging Face identifier
α
br
wr
bc
wc
Qwen 3B
Qwen/Qwen2.5-3B-Instruct
−0.30
−109.2
−1.170×10-3
−51.41
−2.043×10-3
Gemma 4B
google/gemma-3-4b-it
−0.15
. −1630
−−5.535×10-5
−633.3
−1.364×10-4
Mistral 7B
mistralai/Mistral-7B-Instruct-v0.3
−1.3
−75.60
−1.099×10-3
−210.6
−3.455×10-4
Llama 8B
meta-llama/Llama-3.1-8B-Instruct
−−0.50
−106.3
−6.279×10-4
−1.624
−3.262×10-3
Appendix
Table 11 : Selected AlphaSteer and AdaSteer parameters. For AlphaSteer, α denotes the selected steering strength. For AdaSteer, br,wr,bc,wc are the fitted adaptive coefficient parameters.