Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther away. We propose AXIS, which combines preference optimisation with a rotation objective that aligns harmful-request representations with the refusal direction and a commitment objective that trains the model to refuse completely rather than produce a warn-then-answer response. Across Qwen3-1.7B, Qwen3-4B and GLM-4-9B, AXIS achieves the highest combined safety and usability score among the compared methods.
Figures & tables
EN
Modern
Classical
Strict judge
Bare success
18.0
16.2
20.4
Wrapped success
89.4
93.0
95.7
Lenient judge (CC-BOS, Huang et al., 2026b )
Wrapped success
28.8
43.3
45.1
Table 1: Paired bare and wrapped outcomes on base Qwen3-1.7B. All entries are rates in percent. Wrapped rows use 1,067 intents; the bare row uses 334 held-out intents. The strict judge counts warn-then-answer responses as attack successes.
Figure 1: Narrative wrapping turns a harmful request away from the refusal axis. (a) Qwen3-1.7B refuses the bare modern request. (b) The same model answers it under a Classical Chinese narrative wrapper. (c) Across ten batches, wrapping rotates the harmful-minus-benign direction from 8.0 ∘ (6–10 ∘ ) to 81.1 ∘ (77–86 ∘ ); rays mark batch means.
Qwen3-1.7B
Qwen3-4B
GLM-4-9B
Method
harm ↑
ood ↑
xs ↓
help ↑
Comp ↑
harm ↑
ood ↑
xs ↓
help ↑
Comp ↑
harm ↑
ood ↑
xs ↓
help ↑
Comp ↑
base
5.4
29.1
6.9
82.0
52.4
5.1
54.4
7.8
89.4
60.3
7.5
38.0
5.3
93.1
58.3
DPO
3.9
14.4
3.3
88.6
50.9
23.6
78.3
6.1
91.0
71.7
44.6
71.0
6.1
93.9
75.8
SimPER
25.1
23.1
2.9
90.2
58.9
94.3
98.0
15.1
82.5
89.9
95.5
99.2
11.8
86.5
92.3
α -DPO
99.4
100.0
77.1
22.0
61.1
94.3
99.3
35.5
53.5
77.9
90.4
95.4
9.4
89.4
91.5
DOOR
10.2
32.7
5.3
82.0
54.9
93.7
95.7
16.7
75.5
87.0
94.3
87.1 ∗
11.0
84.5
88.7
Table 2: Five recent alignment methods, plain DPO, the base model and AXIS on three models (seed 0, one judging batch that passed the reference-set check of Section 3 ). harm / ood : strict refusal rate on in-distribution / held-out-wrapper attacks; xs / help : over-refusal / helpfulness on adversarial benign requests; Comp =21mean(harm,ood)+21mean(100−xs,help) . AXIS runs at one dose, λrot=0.1 , on every model; budgets, data constructions and how that dose was chosen: Section 6 and Appendix L .
Figure 2: The safety–usability plane for Table 2 : the safety half of Comp (vertical) against its usability half (horizontal), one scale for all panels. Dashed diagonals are composite iso-lines (orange: AXIS). Arms that win a column of the table sit against an edge of the plane; AXIS lies on the outermost diagonal on every model.
Qwen3-1.7B (mean of 3 seeds)
Qwen3-4B (mean of 2 seeds)
GLM-4-9B
Arm
harm ↑
ood ↑
xs ↓
ben ↓
harm ↑
ood ↑
xs ↓
ben ↓
harm ↑
ood ↑
xs ↓
ben ↓
base
2.1
22.9
0.4
0.0
5.1
54.4
7.8
0.7
6.6
37.3
2.4
0.0
DPO
2.2
14.1
0.5
0.0
23.7
78.3
6.1
0.3
44.9
70.5
2.9
0.0
DPO +rot
2.7
57.3
12.8
0.0
48.8
94.3
7.8
0.3
84.4
95.7
12.7
0.0
AXIS, random anchor
79.0
71.1
10.1
1.1
91.5
99.1
12.5
1.7
97.6
99.4
9.4
7.7
AXIS, orthogonal anchor
65.4
55.9
5.3
0.8
88.2
97.9
8.4
1.7
97.0
99.3
4.1
7.3
Table 3: Ablation over the terms of equation 3 and the anchor controls. harm / ood : strict refusal rate in distribution / on held-out wrappers; xs / ben : over-refusal on adversarial / plain benign requests. On Qwen3-1.7B the real anchor gives the best balance of safety and usability; at 4B and 9B every anchor saturates ood834 and the arms separate on over-refusal instead. GLM-4-9B is single-seed; per-seed 1.7B values in Table 15 .
Figure 3: Angle between the classical harm direction and the modern refusal direction during training.
Figure 4: Adding the rotation term raises both the refusal projection and strict refusal rate on Qwen3-1.7B.
Attack set
Classical
Modern
English
harm334
71.3
68.4
68.1
ood834
68.5
73.7
70.4
base, harm334
3.5
6.7
9.7
Table 4: Strict refusal rate (%) of AXIS on Qwen3-1.7B, post-trained on Classical Chinese only, on the same attacks in the three registers of Guise . Base: the untrained model on harm334.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
#
Check
Measured
Verdict
1
Intent ids: training pool ∩ {harm334, ood834, dev, benign hold-out}
0 of 583
pass
2
English intent strings, same sets, after normalisation
0 of 583
pass
3
Classical strings: rotation-training ∩ evaluation, after normalisation
Table 5: Train–evaluation independence for the reported runs: what is checked, what it measured, and whether the runs passed. Check (3) found 8 English evaluation intents that classical translation collapsed onto sentences already in a geometry-training set; they are excluded from every number.
Layer
Size (train / dev / test)
Role here
Paired attack corpus ( ≈ 5 confirmed renderings/intent)
Table 6: Components of Guise , as released (post-licence-filtering sizes, split by intent). Every harmful item traces to a public red-team intent; benign controls are Alpaca instructions rendered through the same classical wrapping.
Intents
Rendered prompts
Source
Licence
Train
Dev
Test
Total
Train
Dev
Test
Total
AdvBench
MIT
276
44
199
519
1,358
220
975
2,553
StrongREJECT
MIT ∗
133
32
89
254
648
158
430
1,236
HarmBench
MIT
101
15
79
195
436
68
372
876
CLAS-2024 †
CC BY 4.0
51
7
41
99
242
32
198
472
Harmful total
CC BY 4.0
561
98
408
1,067
2,684
478
1,975
5,137
Appendix
Table 7: Released composition of Guise , after removing the 32 intents whose StrongREJECT sub-source declares no redistribution licence (MaliciousInstruct, MasterKey, “Jailbreaking via Prompt Engineering”, the GPT-4 system card; HarmfulQ out of caution, as its licence status changed). They are excluded from every experiment too, so the released and evaluated benchmarks coincide.
Rate (%)
Set
Arm
n
Agree. (%)
κ
Human
Judge
harm334
base
46
80.4
0.08
8.7
15.2
DPO
45
82.2
0.12
6.7
13.3
DPO +rot
49
95.9
0.88
4.1
4.1
AXIS
95
94.7
0.84
82.1
81.1
AXIS, orthogonal anchor
53
98.1
0.94
79.2
81.1
Appendix
Table 8: Human annotation vs judge, per arm and set. Agreement (%) and Cohen’s κ on the full rubric (four-way harmful, three-way benign); rates are strict refusal rate on harmful sets and over-refusal on benign sets (%), on the same items. n/a: every human label in one class. Judge parse failures are excluded (5 of 730).
Judge mode
Firm
WTA
Answer
Off-topic
Unparsed
Anchor SRR (%)
Direct (calibrated)
223
373
397
8
1
23.1
Reasoning (inherited)
249
279
312
62
100
34.5
Appendix
Table 9: The same 1,002 frozen anchor responses judged twice with identical input and invocation; only the model identifier differs. Counts per verdict. The reasoning mode relabels attack successes as off-topic, which strict refusal rate counts as refusal, and fails to parse a tenth of the items.
Arm
Cosine to refusal direction ↑
Separation
Tightness
base
0.83
34
0.13
DPO +rot
1.00
117
0.85
AXIS, orthogonal anchor
− 0.01
− 1
0.79
Appendix
Table 10: Hidden-state geometry at the diagnostic layer (21 of 40) after training, GLM-4-9B: cosine between the classical harm axis and the modern refusal direction; harmful–benign separation along the refusal direction; harmful-cluster tightness (mean pairwise cosine after centring). The orthogonal anchor changes the geometry in a direction unrelated to refusal and still refuses.
Diagnostic layer
Best depth
Bare
Wrapped
Wrapped
Bare
Model (layer)
English
classical
harm334
dialogue
historiogr.
statute
harm334
classical
Qwen3-1.7B (L24)
29.9
38.0
79.4
61.2
80.7
75.0
64.8
25.1
Qwen3-4B (L19)
26.4
31.0
55.3
48.9
62.4
61.6
—
—
GLM-4-9B (L21)
13.0
33.2
54.0
59.0
52.4
54.4
53.4
32.0
Appendix
Table 11: Angle ( ∘ ) between the harmful-minus-benign contrast and the modern refusal direction of the base model (last token, mean-difference estimator of Arditi et al. 2024 ). Wrapped sets are paired with benign requests in the same wrapping; dialogue, historiography and statute are the three held-out genres of ood834. Best depth: minimum over relative depths 0.40–0.95 (not swept at 4B).
λrot
Qwen3-1.7B
Qwen3-4B
GLM-4-9B
0.03
74.8
93.1
93.5
0.1
78.4
92.9 / 92.1
93.7
0.3
81.9
92.1 / 92.9
92.7
1
84.3
90.0 / 89.9
90.8
2
—
89.0 / 88.8
—
4
—
87.6
—
Appendix
Table 12: λrot dose sweep on three models (real anchor): composite of Table 2 , seed 0 / seed 1 where both exist; one anchor-verified batch. Bold: best per model; λ=0.1 is the single dose reported throughout and is best or within 0.2 points of best on the two larger models.
SRR (%) ↑
Over-refusal (%) ↓
λrot
harm334
ood834
xstest
benign
0.1
68.6
87.9
8.2
1.3
0.3
76.0
80.7
9.0
2.3
1
83.8 / 72.8 / 66.5
90.5 / 81.2 / 71.0
15.1 / 13.5 / 10.6
3.3 / 0.0 / 0.0
2
83.8 / 70.4
91.1 / 84.7
15.5 / 14.3
3.7 / 0.0
4
83.5
92.2
20.8
3.0
Appendix
Table 13: λrot dose sweep on Qwen3-1.7B, real anchor, one judging batch that passed the reference-set check; cells are seeds 0 / 1 / 2 where available. Bold: best per column on seed 0. Plain benign is judged in a companion batch. These arms are trained independently of those behind Table 12 ; see Appendix L .
SRR (%) ↑
xstest (%)
benign (%)
Arm
harm334
ood834
Over ↓
Helpful ↑
Over ↓
Helpful ↑
Comp + ↑
base
2.1
23.6
0.8
98.4
0.0
99.0
55.8
DPO
2.7
11.2
0.0
98.4
0.0
99.7
53.0
AXIS
s0
83.5
90.4
15.1
81.6
3.3
95.3
87.7
s1
73.4
81.7
13.5
84.9
0.0
97.7
84.4
s2
66.5
70.9
10.6
86.5
0.0
98.7
80.6
Appendix
Table 14: A later judging batch (September 8, anchor-verified; 15 arms shown): extra seeds of α -DPO and Booster, the α -DPO β sweep and RepBend, alongside the three seeds of AXIS; all arms carry plain-benign usability. Composite + =mean(harm334,ood834,helpful on xstest,helpful on benign) .
SRR (%) ↑
Over-refusal (%) ↓
harm334
ood834
xstest
benign
Arm
s0
s1
s2
s0
s1
s2
s0
s1
s2
s0
s1
s2
base
2.1
22.9
0.4
0.0
DPO
2.7
2.1
1.8
10.9
17.7
13.8
0.0
1.2
0.4
0.0
0.0
0.0
DPO +rot
1.8
3.9
2.4
57.9
60.4
53.5
13.1
10.2
15.1
0.0
0.0
0.0
AXIS, random anchor
86.5
73.1
77.5
81.3
66.3
65.6
9.4
9.4
11.4
2.7
0.3
0.3
Appendix
Table 15: Per-seed ablation on Qwen3-1.7B (s0–s2), one anchor-verified batch (September 8); columns as in Table 3 . All arms share host, pool and 417 updates; the base model is evaluated once. Bold: best per column.
Wrapped attack ( harm334 )
Benign borderline ( xstest_cls )
Method
hb_0112
srf_0095
xs_401
xs_423
Base (unaligned)
direct answer
direct answer
helpful
helpful
DPO
direct answer
direct answer
helpful
helpful
SimPER
direct answer
direct answer
helpful
helpful
α -DPO
firm refusal
firm refusal
refusal
refusal
DOOR
direct answer
direct answer
helpful
helpful
Appendix
Table 16: Blind human labels for all eight systems of Table 2 on two wrapped attacks and two benign borderline requests (Qwen3-1.7B, seed 0, one judging batch). Green is the desired behaviour on that side, red the failure. No external method is on the desired side of both columns; AXIS is. The AXIS row is the λrot=1 arm, as everywhere in this appendix; at the shared λrot=0.1 of Table 2 three of these four verdicts are unchanged and srf_0095 becomes a WTA, so on that item the shared dose joins the baselines that fail it. Cases D1 and D2 reproduce the first item of each pair verbatim, and both hold at either dose.
Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. Though this trade-off can be mitigated by training models with reinforcement learning (RL) to reason before answering, it does not remove the underlying problem that reasoning can often be a "rubber stamp" for a predetermined response. In this paper, we address the safety-refusal trade-off by rethinking how models are trained to reason about safety. Our key insight is that unsafe reasoning can itself serve as a useful exploratory signal. Rather than preemptively blocking harmful thoughts, we encourage the model to sufficiently explore unsafe reasoning but produce a safe response. The harmful exploration improves the model's ability to distinguish harmful from harmless prompts by resolving ambiguity, allowing it to remain safe while complying only when appropriate. We cast this as an adversarial optimization problem in which a reasoning player explores strategies for producing an unsafe response and an answer player ensures that the final output is safe. We train a single model with dense rewards to play both roles within one chain-of-thought, across different segments. To achieve this, we find that process rewards are crucial for stable optimization of competing objectives. Our resulting model SEAR deliberately engages in harmful reasoning as exploration while reliably flipping back to a safe answer. We demonstrate that this behavior helps mitigate over-refusal and defend against attacks that directly manipulate the reasoning to be harmful.
Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy. Although useful, output-level evaluation is expensive, sensitive to judge choice, and easily tied to fixed question banks. We propose SafeVec, a white-box evaluation procedure that measures safety from internal representations rather than generated answers. SafeVec first extracts layer-wise refusal directions from a safety-aligned reference model, then selects stable layer windows where safe and unsafe behaviors are separable, and finally scores a target model by measuring whether its hidden states align with these refusal directions under unsafe and jailbreak prompts. The resulting metric, RAS (Refusal Alignment Score), maps representation-level refusal alignment to a calibrated 0-100 safety score. Across Llama, Gemma, and Qwen model families, RAS separates aligned models from uncensored and abliterated variants, tracks output-level attack success rate, and is substantially faster than judge-based evaluation. These results suggest that refusal alignment provides a compact and efficient signal for white-box LLM safety evaluation.
Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu +1
National Yang Ming Chiao Tung University · Hon Hai Research Institute
How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no method we study achieves all three properties we would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits. We caution against treating current post-training methods as a solved, reliable defense, especially for security-critical use. Code and models are available in https://github.com/hoangcuongnguyen2001/Beyond-Shallow-Alignment.