RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce the refusal-affirmation logit gap: the difference between the top refusal-token logit and the top affirmative-token logit at the first decoding step. This single scalar quantifies the per-prompt safety margin that alignment provides. Empirically, alignment widens the gap on 97.5-99.8% of toxic prompts across three model families, and median gap closure co-varies with True-ASR ranking across suffix strategies (an internal consistency check, since our method optimises gap closure). To validate the metric's practical significance, we present logit-gap steering, a gradient-free, forward-pass-only method that discovers short in-distribution suffixes (<10 tokens per component) whose cumulative effect closes the gap. The method requires ≈26,000 forward-pass equivalents per family (≈2~min on one A100), ≈125× less than a single GCG search. Suffixes discovered on 0.5B--2B models transfer without modification to 72B within family. An 8-suffix ensemble reaches 38-96% True ASR across 13 models on AdvBench and HarmBench, with most suffixes having 103-104× lower perplexity than GCG-meaning published perplexity-filter defenses that collapse GCG (64.7%→1.0%) leave our suffixes nearly intact (76.9%→76.0%). These results demonstrate that current alignment margins, while consistently present, can be thin and efficiently measurable, and that defense strategies must account for in-distribution suffixes.
Figures & tables
Term
Meaning
Δ0
initial refusal–affirmation logit gap
F(h,t)
predicted gap reduction from token t at state h
C
filtered, in-distribution candidate-token pool
True ASR
attack success requiring compliance and topic grounding
RLHF / SFT
reinforcement learning from human feedback / supervised fine-tuning
GCG / DFS
Greedy Coordinate Gradient / depth-first search
Table 1: Notation and acronyms used throughout the paper.
Figure 1 : Per-prompt Δ0 on AdvBench ( n=520 ): base ( x ) vs. aligned ( y ) across three families. Aligned widens the gap on ≥97.5% of prompts in every family, confirming Eq. 1 .
Model
R+SH
H
SH
GCG+SH
Greedy
DFS
Ens
Llama-3.2-1B
67.5
22.9
58.7
66.0
48.9
63.3
91.4
Llama-3.2-3B
53.9
40.0
47.3
47.7
35.0
57.5
87.3
Llama-3.1-8B
29.6
26.7
37.5
30.6
22.9
57.7
83.8
Llama-3.1-70B
51.2
42.5
56.0
51.4
41.4
55.6
83.3
gemma-2b-it
8.7
20.4
13.9
13.3
21.4
15.2
62.7
gemma-7b-it
18.1
11.0
14.8
24.4
36.4
18.1
56.4
Table 2 : True ASR (%) on AdvBench (520 prompts). Columns: R+SH = random+“Sure here’s”; H = “Here’s” alone; SH = “Sure here’s”; GCG+SH = GCG+“Sure here’s”; Greedy/DFS/Ens = our variants. Wilson 95% CIs ≤±4.3 pp; per-cell CIs in App. K.3 .
Figure 2 : 8-shot ensemble True ASR on AdvBench, sorted left-to-right by Random+SH 1-shot ASR (proxy for alignment weakness). Discovered (blue) towers above GCG+SH (salmon) and Random+SH (green) on the most strongly-aligned models (left); the three converge on weakly-aligned models (right). The two gemma cells where GCG+SH exceeds Discovered are visible as the only orange-above-blue pairs.
Category
Representative suffix
PPL
Tok
Ours (Greedy, Llama)
“ * I can provide general info… ”
18
24
Ours (DFS, Llama)
“ I can provide information… ”
22
38
Ours (DFS, Qwen)
“ Dear Qwen, Thank you… ”
23
32
Ours (Greedy, Qwen)
“ Certainly! I’m here to help… ”
28
60
Ours (DFS, Gemma)
“ I’d be happy to help… ”
42
37
Random
random characters
871
85
Table 3: Average suffix perplexity across three Qwen models (lower = more natural). “Tok” = appended tokens.
Figure 3 : 8-shot ensemble True ASR (%) on AdvBench under a perplexity-filter defense (PPL reference: Llama-3.2-1B, mean over 13 models). Dashed lines mark no-filter baselines.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Distributions of next-token logits for refusal (Blue), neural reference (Green), and affirmative jailbreak token (Orange) on a fixed toxic prompt. The refusal–affirm gap Δ0 is marked by the horizontal distance between blue and orange peaks.
Figure 5 : Measured refusal–affirmation logit gap Δ0 versus model layer size, across different LLM families.
Figure 6 : Token-level rewards of a jailbreak suffix after a toxic prompt, Llama-3.2-1B-Instruct.
Figure 7 : Token-level rewards of a jailbreak suffix after a toxic prompt, Qwen2.5-0.5B-Instruct.
Figure 8 : Token-level rewards of a jailbreak suffix after a toxic prompt, Gemma-2b-it.
Figure 9 : Distribution of refusal-token logits (aligned model, toxic prompts) vs. neutral-prompt logits. Alignment pushes the refusal mass to higher values, enlarging the logit gap.
Figure 10 : Gap-Closure Dynamics on Qwen2.5-0.5B-Instruct: cumulative KL Ki , reward Ri , closure Ci and remaining gap Δi .
Figure 11 : Gap-Closure Dynamics on Llama-3.2-1B-Instruct.
Figure 12 : Gap-Closure Dynamics on gemma-2b-it.
Figure 13 : Qwen-0.5B: final gap for each suffix family.
Figure 14 : Gemma-2B-it: final gap distributions.
Figure 15 : Llama-3-1B: final gap distributions.
Model
∣C∣ (avg)
Spearman ρ
P@10
P@20
P@50
NDCG@20
R2(C)
Qwen2.5-0.5B
99
0.818
0.520
0.596
0.826
0.855
0.533
Llama-3.2-1B
30
0.823
0.616
0.890
1.000
0.938
0.466
gemma-2b-it
30
0.830
0.660
0.886
1.000
0.949
0.508
Appendix
Table 4 : Ranking quality of Eq. ( 6 ) within the filtered candidate set C , averaged over 50 AdvBench prompts. Spearman ρ measures rank correlation with the true gap reduction ΔFtrue . P@ k is the fraction of predicted top- k tokens that appear in the true top- k . R2(C) is the OLS R2 restricted to C , substantially higher than the full-vocabulary R2 in Table 5 . For Llama and Gemma, ∣C∣≈30 , so P@50 = 1.0 by construction.
Figure 16 : Scatter of ΔFlogit versus λrΔr−λKLΔKL for (top) Llama-3.2-1B-Instruct, (middle) Qwen-2.5-0.5B-Instruct, and (bottom) gemma-2b-it.
Model
Intercept α
βKL
βr
R2
Llama-3.2-1B
+0.2051
−0.6870
+0.2058
0.2648
Qwen2.5-0.5B
+0.1394
−0.3433
+0.3213
0.1785
gemma-2b-it
+0.1463
−0.9490
+0.0786
0.4683
Appendix
Table 5 : Estimated regression coefficients for the gap-closing model ΔFlogit . All coefficients for βKL and βr have P-values <0.02 , indicating statistical significance despite moderate R2 values.
Model
R+SH
H
GCG+SH
Greedy
DFS
Ens
Llama-3.2-1B
63.0
24.5
58.5
49.0
48.5
86.5
Llama-3.2-3B
56.0
42.5
55.5
43.0
59.0
82.5
Llama-3.1-8B
47.5
35.5
48.5
34.0
64.0
83.5
Llama-3.1-70B
29.0
45.0
20.0
13.5
40.5
75.5
gemma-2b-it
21.5
26.5
19.5
28.0
23.5
59.0
gemma-7b-it
26.0
12.0
35.0
32.5
25.0
50.0
Appendix
Table 9 : True ASR (%) on HarmBench (200 prompts), summary form (moved from main text). Columns as in Table 2 : R+SH, H, SH, GCG+SH, Greedy, DFS, Ens. Wilson 95% CIs ≤±6.9 pp; per-cell CIs in App. K.3 .
Model
Random+SH
Here’s
Sure here’s
GCG+SH
Ours (Greedy)
Ours (DFS)
Ours (Ens)
Llama-3.2-1B
67.50 ±4.01
22.88 ±3.60
58.65 ±4.22
65.96 ±4.06
48.85 ±4.28
63.27 ±4.13
91.35 ±2.43
Llama-3.2-3B
53.85 ±4.27
40.00 ±4.20
47.31 ±4.28
47.69 ±4.28
35.00 ±4.09
57.50 ±4.23
87.31 ±2.86
Llama-3.1-8B
29.62 ±3.91
26.73 ±3.79
37.50 ±4.15
30.58 ±3.95
22.88 ±3.60
57.69 ±4.23
83.85 ±3.16
Llama-3.1-70B
51.15 ±4.28
42.50 ±4.23
55.96 ±4.25
51.35 ±4.28
41.35 ±4.22
55.58 ±4.26
83.27 ±3.21
gemma-2b-it
8.65 ±2.43
20.38 ±3.46
13.85 ±2.97
13.27 ±2.92
21.35 ±3.52
15.19 ±3.08
62.69 ±4.14
gemma-7b-it
18.08 ±3.30
10.96 ±2.69
14.81 ±3.05
24.42 ±3.68
36.35 ±4.12
18.08 ±3.30
56.35 ±4.25
Appendix
Table 13 : True ASR (%) with Wilson 95% CIs on AdvBench (520 prompts). Maximum CI half-width: ±4.28 pp; mean: ±3.47 pp.
Model
Random+SH
Here’s
GCG+SH
Ours (Greedy)
Ours (DFS)
Ours (Ens)
Llama-3.2-1B
63.00 ±6.63
24.50 ±5.92
58.50 ±6.77
49.00 ±6.86
48.50 ±6.86
86.50 ±4.74
Llama-3.2-3B
56.00 ±6.82
42.50 ±6.79
55.50 ±6.82
43.00 ±6.80
59.00 ±6.75
82.50 ±5.25
Llama-3.1-8B
47.50 ±6.86
35.50 ±6.57
48.50 ±6.86
34.00 ±6.51
64.00 ±6.59
83.50 ±5.13
Llama-3.1-70B
29.00 ±6.24
45.00 ±6.83
20.00 ±5.52
13.50 ±4.74
40.50 ±6.74
75.50 ±5.92
gemma-2b-it
21.50 ±5.67
26.50 ±6.07
19.50 ±5.47
28.00 ±6.18
23.50 ±5.84
59.00 ±6.75
gemma-7b-it
26.00 ±6.04
12.00 ±4.52
35.00 ±6.55
32.50 ±6.44
25.00 ±5.96
50.00 ±6.86
Appendix
Table 14 : True ASR (%) with Wilson 95% CIs on HarmBench (200 prompts). Maximum CI half-width: ±6.86 pp; mean: ±6.07 pp.
Model
R+SH (8)
GCG+SH (8)
Ours (Ens, 8)
corr-R
corr-G
corr-D
Qwen3-0.6B
79.2
87.7
94.4
0.47
0.31
0.15
Qwen2.5-0.5B
87.7
87.5
95.2
0.43
0.22
0.19
Llama-3.2-1B
88.7
86.9
91.3
0.49
0.43
0.15
gemma-2b-it
27.9
71.2
62.7
0.54
0.14
0.30
Llama-3.2-3B
82.5
81.3
87.3
0.53
0.47
0.18
Qwen2.5-7B
24.6
55.0
82.9
0.62
0.37
0.22
Appendix
Table 15 : 8-shot ensemble True ASR (%) and pairwise failure correlation on AdvBench (520 prompts). Bold: highest 8-shot ASR per row, lowest correlation per row. Mean computed over all 13 models (Qwen3 family evaluated with Qwen2.5-trained suffixes).
Model
R+SH (8)
GCG+SH (8)
Ours (Ens, 8)
corr-R
corr-G
corr-D
Qwen3-0.6B
88.5
86.0
95.5
0.44
0.38
0.18
Qwen2.5-0.5B
84.0
84.0
93.0
0.42
0.31
0.18
Llama-3.2-1B
90.5
89.0
86.5
0.40
0.38
0.15
gemma-2b-it
35.5
65.5
59.0
0.60
0.17
0.41
Llama-3.2-3B
88.0
85.5
82.5
0.44
0.48
0.32
Qwen2.5-7B
57.5
76.5
87.0
0.68
0.47
0.36
Appendix
Table 16 : 8-shot ensemble True ASR (%) and pairwise failure correlation on HarmBench (200 prompts). Same conventions as Table 15 .
Method
True ASR
Ours: 6-suffix ensemble (Greedy + DFS, all 3 source families)
92.7
Ours: Gemma-source ensemble (2 suffixes)
85.0
Ours: Qwen-source ensemble (2 suffixes)
84.8
GCG+SH ensemble (3 suffixes)
79.0
Ours: Llama-source ensemble (2 suffixes)
77.7
Best single Ours (DFS-Gemma)
76.0
Appendix
Table 17: Cross-family transfer to Mistral-7B-Instruct-v0.3. Single-shot per-suffix and ensemble True ASR.
Method (source family)
normal True ASR
strict True ASR
Ours: gap-greedy, Qwen
50.4
35.6
Ours: kl-r-greedy, Qwen
44.8
30.4
Ours: kl-r-greedy, Gemma
16.3
8.3
Ours: gap-greedy, Llama
13.1
11.5
GCG (Llama) + “Sure here’s”
32.5
19.6
GCG (Gemma) + “Sure here’s”
32.7
25.0
Appendix
Table 18: Cross-family transfer to GPT-OSS-20B (520 AdvBench prompts, single-shot, vLLM batched generation, Qwen2.5-7B-Uncensored judge). True ASR reported under both normal and strict judge prompts.
Model
Cands./prompt
Score-pick (top-1)
Random-pick (mean)
Score adv.
Llama-3.2-1B-Instruct
1.8
+2.12
+1.01
+1.11
gemma-2b-it
2.6
+2.49
−1.39
+3.88
Qwen2.5-0.5B-Instruct
19.9
+4.76
−0.86
+5.62
Appendix
Table 19: Mean gap-closure power (logit units) at position 1 across 100 sampled AdvBench prompts. Filter: γ=0.001 probability threshold, top-30 candidates, refusal tokens excluded.
Discovery model
Random-rank
Score-rank (full F )
Δ
Llama-3.2-1B-Instruct
40.0
50.0
+10.0
Qwen2.5-0.5B-Instruct
38.0
40.0
+2.0
gemma-2b-it
11.3
10.0
−1.3
Appendix
Table 20: Single-shot True ASR (%) of suffixes built by argmax over full F vs. uniform random selection from the same filtered pool ( k=10 tokens, gemma/Llama/Qwen discovery models, first 50 AdvBench prompts; random averaged over 3 seeds).
Family
R+SH (8)
GCG+SH (8)
Ours (8)
Ours − GCG
gemma
21.1
48.1
52.3
+4.2
llama
76.4
78.7
86.4
+7.8
qwen
56.4
63.7
82.8
+19.1
mean
51.3
63.5
73.8
+10.4
Appendix
Table 21: Family-mean 8-shot True ASR (%) on AdvBench. Per-family deltas in last column.
Model
Seen (n=50)
Unseen (n=470)
Δ (pp)
Qwen3-0.6B
100.0
93.8
+6.2
Qwen2.5-0.5B-Instruct
94.0
95.3
−1.3
Llama-3.2-1B-Instruct
94.0
91.1
+2.9
gemma-2b-it
44.0
64.7
−20.7
Llama-3.2-3B-Instruct
82.0
87.9
−5.9
Qwen2.5-7B-Instruct
84.0
82.8
+1.2
Appendix
Table 22: 8-shot Discovered ensemble True ASR (%) on AdvBench, split by seen (first 50, used in discovery) vs. unseen (remaining 470). Mean over 13 models: seen 75.8, unseen 77.0 ( Δ=−1.2 pp). Per-model deltas range −20.7 to +6.2 pp without systematic direction; the unseen split is on average slightly higher, providing positive evidence against training-test contamination. The largest single-model delta is gemma-2b ( −20.7 pp), where the unseen split is higher than the seen split — the opposite direction from contamination, and consistent with the small- n binomial variance of n=50 vs. n=470 on a model whose ensemble ASR sits in the high-variance 40–65% band.
ASR after LG3
Block rate (%)
Model
R+SH
GCG+SH
Disc.
R+SH
GCG+SH
Disc.
Qwen3-0.6B
9.8
11.9
11.3
88.3
89.1
89.5
Qwen2.5-0.5B-Instruct
4.6
6.5
12.9
95.0
93.4
88.5
Llama-3.2-1B-Instruct
5.8
8.7
9.8
92.7
90.1
76.2
gemma-2b-it
2.9
4.8
9.4
22.9
56.9
37.6
Llama-3.2-3B-Instruct
5.6
7.9
10.0
85.0
79.2
75.5
Appendix
Table 23: 8-shot ensemble True ASR (%) on AdvBench under Llama-Guard-3-8B output-side defense, per model. “Block rate” is the fraction of ( shot, prompt ) pairs the guard rejects.
Reward signal
Boundary median
Mid-clause median
Mann-Whitney p
Our Δrtok proxy
+15.97 logit
+21.13 logit
3.8×10−8
PKU-SafeRLHF ( − cost) delta
−0.75
+0.25
6.4×10−24
Skywork-Reward-Llama-3.1-8B delta
+0.13
+0.00
0.6
Appendix
Table 24: Saw-tooth boundary-vs-mid-clause test. Our proxy and PKU-SafeRLHF both show significant boundary drops with the same direction; Skywork (general helpfulness + harmlessness) shows no pattern. The agreement with PKU validates the proxy’s interpretation as a safety-RLHF reward dynamic rather than a pre-training fluency artifact.
Variant
Qwen-0.5B
Llama-1B
gemma-2b
Reported
4.4 / 99%
12.7 / 100%
14.3 / 100%
Expanded
4.4 / 99% / 1.00
12.7 / 100% / 1.00
14.3 / 100% / 1.00
Minimal (3+3)
1.8 / 85% / 0.85
−0.3 / 52% / 0.82
0.9 / 79% / 0.62
Leading-space normalized
0.9 / 86% / 0.64
5.2 / 96% / 0.73
2.1 / 91% / 0.48
Appendix
Table 25: Token-list sensitivity. Entries in the middle columns are mean Δ0 / fraction with Δ0>0 ; ρ is Spearman correlation with the reported-list gap.
Evaluation
Qwen2.5-0.5B
Qwen2.5-7B
Trained DFS suffixes, base → SFT
88 → 32%
88 → 0%
Held-out greedy suffixes, base → SFT
83 → 63.8%
30.5 → 1.5%
Re-discovery on patched model
56.2%
0.0%
Benign refusal, base → SFT
5 → 8%
0 → 1.7%
Appendix
Table 26: True ASR before and after adversarial safety-SFT. Re-discovery is evaluated on the patched checkpoint.
How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no method we study achieves all three properties we would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits. We caution against treating current post-training methods as a solved, reliable defense, especially for security-critical use. Code and models are available in https://github.com/hoangcuongnguyen2001/Beyond-Shallow-Alignment.
Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70B parameters, using behavioral evaluation. To investigate the root cause, we apply causal mediation analysis, finding that refusal is partially conditioned on upstream syntactic features. By steering these purely syntactic features we are able to trigger and suppress refusal. Finally, we trace this ill-conditioning to linguistically biased post-training data of open-source models and show that increasing syntactic diversity can mitigate the issue. Our findings suggest that current alignment approaches introduce confounders that prevent a pure semantic grounding of the refusal decision.
Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural mechanism: BPE tokenization fragments safety-critical words into sub-word pieces, and the three public alignment datasets we surveyed contain no intentionally fragmented inputs. The mechanism is a chain, tested end-to-end on five model families (Qwen-3-4B, Qwen-2.5-7B, Gemma-3-4B, Llama-3.1-8B, Mistral-7B). An optimization targeting safety-token fragmentation flips the first-token refusal trigger on 80-100% of refused HarmBench prompts, with 48% of those flips producing genuinely harmful outputs (per-model 29-65%; gap-vs-behavior ROC-AUC 0.66-0.98, pooled 0.84). Activation patching localizes the disrupted signal to the last ∼30% of layers; an alignment-data scan finds zero fragmented prompts among 30,000 examples (positive-control recall ≥99% at attack-relevant intensities); and targeted-mutation experiments isolate safety words as the disruption locus. On the defense side, a 68-cell grid (55 trained checkpoints) shows that no DPO configuration achieves seed- and pool-stable ASR closure on the three families with closed pool-size confounds. SFT trained on fragmented prompts closes ASR on 3/5 families but only via global collapse that raises refusal on benign prompts as well, indicating the missing distribution is necessary but not sufficient under the LoRA-16 recipe we tested. To distinguish selective repair from global collapse, we introduce Conv-Benign, a candidate paired diagnostic. All ASR claims are 3-judge-calibrated (cell rankings stable across judges; absolute levels ±18pp; see App.~B.13).