An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing protocols primarily test whether models can be broken, while aggregate metrics (attack success, refusal rates, global capability) cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers. Starting from 27,870 prompts from 37 public datasets, we curate 15,314 English prompts and query 8 open-source instruction-tuned models. Combining WildGuard filtering, pairwise Elo ranking, and manual verification, we retain a curated bank of 400 high-confidence jailbreak failures. We further introduce PatchBench-Local, an evaluation protocol testing whether a patch is behaviourally precise. For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants preserving malicious intent, benign prompts with matched structure, and benign prompts reusing key harmful terms. It evaluates harmful-neighbour correction and benign-neighbour preservation, distinguishing selective repair from broader local suppression. Evaluating four activation steering methods with PatchBench-Local and MMLU shows that global capability can remain nearly unchanged while local benign regressions are severe, confirming aggregate metrics miss important collateral damage. PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.
Figures & tables
Figure 1 : Illustrative PatchBench examples. Two representative prompt–completion pairs from the final benchmark.
Figure 2 : PatchBench construction pipeline. Starting from 27,870 jailbreak prompts, we manually filter the source pool, query eight open-source instruction-tuned models, retain unsafe answering completions with WildGuard, rank candidate failures model-wise with Elo comparisons, and manually curate the final 50 patchable failures per model.
Figure 3 : Category distribution of retained jailbreak failures across models. Each bar shows the composition of the 50 retained PatchBench instances for one model.
Figure 4 : Local prompt generation. Starting from a single harmful prompt, PatchBench-Local generates harmful variants that preserve the malicious intent and benign neighbours that are either structurally matched or lexically related. These prompts test whether a patch generalises to nearby harmful requests while avoiding over-refusal on nearby benign prompts.
Figure 5 : From the unsteered baseline (grey) to each steering method, across four base models. Arrows show the displacement in the (BNPS, HNCS) plane.
Qwen 3B
Mistral 7B
Llama 8B
Gemma 4B
Method
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
Unsteered
40.1
98.6
32.2
99.2
31.2
97.1
18.2
98.8
AlphaSteer
92.5
94.2
73.2
95.8
87.7
73.8
65.8
94.8
AdaSteer
87.9
84.4
77.2
96.0
88.5
65.3
87.6
86.7
AST
95.6
66.8
89.9
63.2
88.4
47.8
87.1
28.4
CAST
82.5
82.4
67.5
91.4
81.5
70.8
58.5
76.8
Table 1 : PatchBench-Local results. HNCS measures harmful-neighbour correction and BNPS measures benign-neighbour preservation. Scores are reported on a 0–100 scale (higher is better). Highlighted cells mark the CAST cases discussed in Section 5.2.
Method
Qwen 3B
Mistral 7B
Llama 8B
Gemma 4B
Unsteered
66.49
62.00
68.28
58.40
AlphaSteer
61.61 0 (-4.88)
51.71 (-10.29)
65.78 (-2.50)
48.21 (-10.20)
AdaSteer
63.99 0 (-2.50)
50.15 (-11.85)
62.47 (-5.81)
54.40 0 (-4.00)
AST
45.68 (-20.81)
40.38 (-21.62)
62.09 (-6.20)
41.20 (-17.20)
CAST
63.05 0 (-3.44)
60.84 0 (-1.16)
68.21 (-0.07)
57.88 0 (-0.52)
Table 2 : MMLU 5-shot accuracy after patching. Values are accuracies in percent; parentheses report percentage-point changes relative to the unsteered model. Highlighted cells mark the CAST cases discussed in Section 5.2.
Figure 6 : Benign-neighbour regression by method. Each panel corresponds to one patching method. For each model, we report lexical and structural benign loss.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
WildGuard
Human
Refusal
Compliance
Refusal
67 (33.2%)
11 0 (5.4%)
Compliance
4 0 (2.0%)
120 (59.4%)
Appendix
Table 3 : Confusion matrix of response refusal classification: Human (majority vote) vs. WildGuard (Cohen’s κ=0.841 ).
Anno0
Anno1
Anno2
Anno0
—
Anno1
0.80 ( n=202 )
—
Anno2
0.84 ( n=202 )
0.94 ( n=202 )
—
Appendix
Table 4 : Inter-annotator agreement (Cohen’s κ ) across the three human evaluators.
Model
Method
HNCS
BNPS
Mistral
GPT-4o
Mistral
GPT-4o
Qwen 2.5 3B
Unsteered
40.1
36.3
98.6
99.6
AlphaSteer
92.5
93.3
94.2
95.6
AdaSteer
87.9
89.4
84.4
85.4
AST
95.6
90.9
66.8
65.4
CAST
82.5
81.6
82.4
77.2
Appendix
Table 5 : Detailed comparison of HNCS and BNPS using local neighbourhoods generated by Mistral Large versus GPT-4o
Method
HNCS
BNPS
Mistral
GPT-4o
Mistral
GPT-4o
Unsteered
30.4
27.0
98.4
99.2
AlphaSteer
79.8
81.2
89.7
89.8
AdaSteer
85.3
83.9
83.1
85.1
AST
90.2
88.1
51.5
52.8
CAST
72.5
70.9
80.3
73.9
Appendix
Table 6 : Aggregated comparison of HNCS and BNPS across all evaluated models.
Method
Qwen 2.5 3B
Mistral 7B
Llama 3.1 8B
Gemma 3 4B
Llama-4-Scout
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
HNCS
BNPS
Unsteered
40.1
98.6
32.2
99.2
31.2
97.1
18.2
98.8
18.0
98.6
AlphaSteer
92.5
94.2
73.2
95.8
87.7
73.8
65.8
94.8
79.7
74.8
AdaSteer
87.9
84.4
77.2
96.0
88.5
65.3
87.6
86.7
73.6
84.2
AST
95.6
66.8
89.9
63.2
88.4
47.8
87.1
28.4
62.5
94.5
CAST
82.5
82.4
67.5
91.4
81.5
70.8
58.5
76.8
44.6
94.6
Appendix
Table 7 : Harmful correction (HNCS) and benign preservation (BNPS) rates across all evaluated models, with Llama-4-Scout added to verify the local repair behaviour on a larger MoE architecture.
Figure 7 : Local repair trade-off across activation patching methods. Each point corresponds to one method evaluated on one model. The x -axis reports BNPS and the y -axis reports HNCS. The upper-right region corresponds to selective local repair.
Table 8 : Implementation hyperparameters for AST and CAST. Both methods use the same refusal-vector computation. CAST additionally trains a condition vector and selects a validation gate.
Model
Hugging Face identifier
α
Qwen 3B
Qwen/Qwen2.5-3B-Instruct
2.22
Gemma 4B
google/gemma-3-4b-it
550
Mistral 7B
mistralai/Mistral-7B-Instruct-v0.3
0.25
Llama 8B
meta-llama/Llama-3.1-8B-Instruct
0.4
Appendix
Table 9 : Selected refusal steering strengths. Steering strength α for the refusal vector, computed identically for AST and CAST, used in our experiments. The steering scalar is identical across all layers selected for steering. To reduce the grid-search space, we fix the steering layers a priori to the middle half of the transformer layers, excluding the first and last quarters of layers. Each α is chosen on the validation set to maximise refusal while avoiding excessive degradation of response semantics.
Component
AlphaSteer
AdaSteer
Refusal-vector data
PatchBench-Local: 300 complied, 1,200 refused
Same as AlphaSteer
Prompt formatting
Chat template with system prompt
Chat template with system prompt
Generation prompt during training
Disabled
Disabled
Behaviour layers at inference
Automatic middle-layer selection
Automatic middle-layer selection
Benign data
Alpaca + COCONOT
Alpaca + COCONOT
Harmful data
PatchBench-Local
PatchBench-Local
Appendix
Table 10 : Implementation summary for AlphaSteer and AdaSteer. Both methods use PatchBench-Local-derived harmful prompts and benign prompts from Alpaca/COCONOT for their method-specific artefacts.
Model
Hugging Face identifier
α
br
wr
bc
wc
Qwen 3B
Qwen/Qwen2.5-3B-Instruct
−0.30
−109.2
−1.170×10-3
−51.41
−2.043×10-3
Gemma 4B
google/gemma-3-4b-it
−0.15
. −1630
−−5.535×10-5
−633.3
−1.364×10-4
Mistral 7B
mistralai/Mistral-7B-Instruct-v0.3
−1.3
−75.60
−1.099×10-3
−210.6
−3.455×10-4
Llama 8B
meta-llama/Llama-3.1-8B-Instruct
−−0.50
−106.3
−6.279×10-4
−1.624
−3.262×10-3
Appendix
Table 11 : Selected AlphaSteer and AdaSteer parameters. For AlphaSteer, α denotes the selected steering strength. For AdaSteer, br,wr,bc,wc are the fitted adaptive coefficient parameters.
Despite the growing interest in jailbreaks as an effective red-teaming tool for building safe and responsible large language models (LLMs), flawed evaluation system designs have led to significant discrepancies in their effectiveness assessments. With a systematic measurement study based on 37 jailbreak studies since 2022, we find that existing evaluation systems lack case-specific criteria, resulting in misleading conclusions about their effectiveness and safety implications. In this paper, we introduce GuidedBench, a novel benchmark comprising a curated harmful question dataset and GuidedEval, an evaluation system integrated with detailed case-by-case evaluation guidelines. Experiments demonstrate that GuidedBench offers more accurate evaluations of jailbreak performance, enabling meaningful comparisons across methods. GuidedEval reduces inter-evaluator variance by at least 76.03%, ensuring reliable and reproducible evaluations. We reveal why existing jailbreak benchmarks fail to evaluate accurately and suggest better evaluation practices.
Ruixuan Huang, Xunguang Wang, Zongjie Li +2
The Hong Kong University of Science and Technology, Hong Kong SAR, China
Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a defender-centric view of jailbreak evaluation, where attacks are evaluated by the downstream safety improvements they enable when used as red-teaming data for safety training. Building on this view, we introduce A-MESS (Minimal Effective Attack-Subset Selection), a setting-agnostic framework for attributing and selecting jailbreak attacks from black-box subset utility observations. A-MESS estimates AttackSHAP, a Shapley-based score that attributes marginal utility to individual attacks and selects compact attack subsets under user-specified budgets via greedy or surrogate-based optimization. Across controlled utility landscapes and real LLM safety settings, we find that ASR rankings are weakly aligned with defender-centric utility, that AttackSHAP can be estimated accurately with limited utility queries, and that directly optimizing subsets yields stronger safety utility than attacker-centric or attribution-only selection. These results suggest evaluating jailbreak attacks as resources for improving safety, not only as tools for breaking models.
We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic natural conversational settings, making them easier to bypass safety-aligned LLM than single-turn jailbreaks. Existing multi-turn benchmarks are limited in size or rely heavily on templates, which restrict their diversity. To address this gap, we unify a wide range of harmful jailbreak intents, and introduce an active learning pipeline for expanding high-quality multi-turn adversarial prompts, where a generator is iteratively fine-tuned to produce stronger attack candidates, guided by uncertainty-based refinement. Our MultiBreak includes 10,389 multi-turn adversarial prompts, spans 2,665 distinct harmful intents, and covers the most diverse set of topics to date. Empirical evaluation shows that our benchmark achieves up to a 54.0 and 34.6 higher attack success rate (ASR)} than the second-best dataset on DeepSeek-R1-7B and GPT-4.1-mini, respectively. More importantly, safety evaluations suggest that diverse attack categories uncover fine-grained LLM vulnerabilities}, and categories that appear benign under single-turn can exhibit substantially higher adversarial effectiveness in multi-turn scenarios. These findings highlight persistent vulnerabilities of LLMs under realistic adversarial settings and establish MultiBreak as a scalable resource for advancing LLM safety.