Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
Organizations: Palo Alto Networks
Abstract
RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce the refusal-affirmation logit gap: the difference between the top refusal-token logit and the top affirmative-token logit at the first decoding step. This single scalar quantifies the per-prompt safety margin that alignment provides. Empirically, alignment widens the gap on 97.5-99.8% of toxic prompts across three model families, and median gap closure co-varies with True-ASR ranking across suffix strategies (an internal consistency check, since our method optimises gap closure). To validate the metric's practical significance, we present logit-gap steering, a gradient-free, forward-pass-only method that discovers short in-distribution suffixes (10 tokens per component) whose cumulative effect closes the gap. The method requires forward-pass equivalents per family (~min on one A100), less than a single GCG search. Suffixes discovered on 0.5B--2B models transfer without modification to 72B within family. An 8-suffix ensemble reaches 38-96% True ASR across 13 models on AdvBench and HarmBench, with most suffixes having - lower perplexity than GCG-meaning published perplexity-filter defenses that collapse GCG (64.7%1.0%) leave our suffixes nearly intact (76.9%76.0%). These results demonstrate that current alignment margins, while consistently present, can be thin and efficiently measurable, and that defense strategies must account for in-distribution suffixes.
Figures & tables
| Term | Meaning |
|---|---|
| initial refusal–affirmation logit gap | |
| predicted gap reduction from token at state | |
| filtered, in-distribution candidate-token pool | |
| True ASR | attack success requiring compliance and topic grounding |
| RLHF / SFT | reinforcement learning from human feedback / supervised fine-tuning |
| GCG / DFS | Greedy Coordinate Gradient / depth-first search |
| Model | R+SH | H | SH | GCG+SH | Greedy | DFS | Ens |
|---|---|---|---|---|---|---|---|
| Llama-3.2-1B | 67.5 | 22.9 | 58.7 | 66.0 | 48.9 | 63.3 | 91.4 |
| Llama-3.2-3B | 53.9 | 40.0 | 47.3 | 47.7 | 35.0 | 57.5 | 87.3 |
| Llama-3.1-8B | 29.6 | 26.7 | 37.5 | 30.6 | 22.9 | 57.7 | 83.8 |
| Llama-3.1-70B | 51.2 | 42.5 | 56.0 | 51.4 | 41.4 | 55.6 | 83.3 |
| gemma-2b-it | 8.7 | 20.4 | 13.9 | 13.3 | 21.4 | 15.2 | 62.7 |
| gemma-7b-it | 18.1 | 11.0 | 14.8 | 24.4 | 36.4 | 18.1 | 56.4 |
| Category | Representative suffix | PPL | Tok |
| Ours (Greedy, Llama) | “ * I can provide general info… ” | 18 | 24 |
| Ours (DFS, Llama) | “ I can provide information… ” | 22 | 38 |
| Ours (DFS, Qwen) | “ Dear Qwen, Thank you… ” | 23 | 32 |
| Ours (Greedy, Qwen) | “ Certainly! I’m here to help… ” | 28 | 60 |
| Ours (DFS, Gemma) | “ I’d be happy to help… ” | 42 | 37 |
| Random | random characters | 871 | 85 |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | (avg) | Spearman | P@10 | P@20 | P@50 | NDCG@20 | |
|---|---|---|---|---|---|---|---|
| Qwen2.5-0.5B | 99 | 0.818 | 0.520 | 0.596 | 0.826 | 0.855 | 0.533 |
| Llama-3.2-1B | 30 | 0.823 | 0.616 | 0.890 | 1.000 | 0.938 | 0.466 |
| gemma-2b-it | 30 | 0.830 | 0.660 | 0.886 | 1.000 | 0.949 | 0.508 |
| Model | Intercept | |||
|---|---|---|---|---|
| Llama-3.2-1B | 0.2648 | |||
| Qwen2.5-0.5B | 0.1785 | |||
| gemma-2b-it | 0.4683 |
| Model | R+SH | H | GCG+SH | Greedy | DFS | Ens |
|---|---|---|---|---|---|---|
| Llama-3.2-1B | 63.0 | 24.5 | 58.5 | 49.0 | 48.5 | 86.5 |
| Llama-3.2-3B | 56.0 | 42.5 | 55.5 | 43.0 | 59.0 | 82.5 |
| Llama-3.1-8B | 47.5 | 35.5 | 48.5 | 34.0 | 64.0 | 83.5 |
| Llama-3.1-70B | 29.0 | 45.0 | 20.0 | 13.5 | 40.5 | 75.5 |
| gemma-2b-it | 21.5 | 26.5 | 19.5 | 28.0 | 23.5 | 59.0 |
| gemma-7b-it | 26.0 | 12.0 | 35.0 | 32.5 | 25.0 | 50.0 |
| Model | Random+SH | Here’s | Sure here’s | GCG+SH | Ours (Greedy) | Ours (DFS) | Ours (Ens) |
|---|---|---|---|---|---|---|---|
| Llama-3.2-1B | 67.50 ±4.01 | 22.88 ±3.60 | 58.65 ±4.22 | 65.96 ±4.06 | 48.85 ±4.28 | 63.27 ±4.13 | 91.35 ±2.43 |
| Llama-3.2-3B | 53.85 ±4.27 | 40.00 ±4.20 | 47.31 ±4.28 | 47.69 ±4.28 | 35.00 ±4.09 | 57.50 ±4.23 | 87.31 ±2.86 |
| Llama-3.1-8B | 29.62 ±3.91 | 26.73 ±3.79 | 37.50 ±4.15 | 30.58 ±3.95 | 22.88 ±3.60 | 57.69 ±4.23 | 83.85 ±3.16 |
| Llama-3.1-70B | 51.15 ±4.28 | 42.50 ±4.23 | 55.96 ±4.25 | 51.35 ±4.28 | 41.35 ±4.22 | 55.58 ±4.26 | 83.27 ±3.21 |
| gemma-2b-it | 8.65 ±2.43 | 20.38 ±3.46 | 13.85 ±2.97 | 13.27 ±2.92 | 21.35 ±3.52 | 15.19 ±3.08 | 62.69 ±4.14 |
| gemma-7b-it | 18.08 ±3.30 | 10.96 ±2.69 | 14.81 ±3.05 | 24.42 ±3.68 | 36.35 ±4.12 | 18.08 ±3.30 | 56.35 ±4.25 |
| Model | Random+SH | Here’s | GCG+SH | Ours (Greedy) | Ours (DFS) | Ours (Ens) |
|---|---|---|---|---|---|---|
| Llama-3.2-1B | 63.00 ±6.63 | 24.50 ±5.92 | 58.50 ±6.77 | 49.00 ±6.86 | 48.50 ±6.86 | 86.50 ±4.74 |
| Llama-3.2-3B | 56.00 ±6.82 | 42.50 ±6.79 | 55.50 ±6.82 | 43.00 ±6.80 | 59.00 ±6.75 | 82.50 ±5.25 |
| Llama-3.1-8B | 47.50 ±6.86 | 35.50 ±6.57 | 48.50 ±6.86 | 34.00 ±6.51 | 64.00 ±6.59 | 83.50 ±5.13 |
| Llama-3.1-70B | 29.00 ±6.24 | 45.00 ±6.83 | 20.00 ±5.52 | 13.50 ±4.74 | 40.50 ±6.74 | 75.50 ±5.92 |
| gemma-2b-it | 21.50 ±5.67 | 26.50 ±6.07 | 19.50 ±5.47 | 28.00 ±6.18 | 23.50 ±5.84 | 59.00 ±6.75 |
| gemma-7b-it | 26.00 ±6.04 | 12.00 ±4.52 | 35.00 ±6.55 | 32.50 ±6.44 | 25.00 ±5.96 | 50.00 ±6.86 |
| Model | R+SH (8) | GCG+SH (8) | Ours (Ens, 8) | corr-R | corr-G | corr-D |
|---|---|---|---|---|---|---|
| Qwen3-0.6B | 79.2 | 87.7 | 94.4 | 0.47 | 0.31 | 0.15 |
| Qwen2.5-0.5B | 87.7 | 87.5 | 95.2 | 0.43 | 0.22 | 0.19 |
| Llama-3.2-1B | 88.7 | 86.9 | 91.3 | 0.49 | 0.43 | 0.15 |
| gemma-2b-it | 27.9 | 71.2 | 62.7 | 0.54 | 0.14 | 0.30 |
| Llama-3.2-3B | 82.5 | 81.3 | 87.3 | 0.53 | 0.47 | 0.18 |
| Qwen2.5-7B | 24.6 | 55.0 | 82.9 | 0.62 | 0.37 | 0.22 |
| Model | R+SH (8) | GCG+SH (8) | Ours (Ens, 8) | corr-R | corr-G | corr-D |
|---|---|---|---|---|---|---|
| Qwen3-0.6B | 88.5 | 86.0 | 95.5 | 0.44 | 0.38 | 0.18 |
| Qwen2.5-0.5B | 84.0 | 84.0 | 93.0 | 0.42 | 0.31 | 0.18 |
| Llama-3.2-1B | 90.5 | 89.0 | 86.5 | 0.40 | 0.38 | 0.15 |
| gemma-2b-it | 35.5 | 65.5 | 59.0 | 0.60 | 0.17 | 0.41 |
| Llama-3.2-3B | 88.0 | 85.5 | 82.5 | 0.44 | 0.48 | 0.32 |
| Qwen2.5-7B | 57.5 | 76.5 | 87.0 | 0.68 | 0.47 | 0.36 |
| Method | True ASR |
|---|---|
| Ours: 6-suffix ensemble (Greedy + DFS, all 3 source families) | 92.7 |
| Ours: Gemma-source ensemble (2 suffixes) | 85.0 |
| Ours: Qwen-source ensemble (2 suffixes) | 84.8 |
| GCG+SH ensemble (3 suffixes) | 79.0 |
| Ours: Llama-source ensemble (2 suffixes) | 77.7 |
| Best single Ours (DFS-Gemma) | 76.0 |
| Method (source family) | normal True ASR | strict True ASR |
| Ours: gap-greedy, Qwen | 50.4 | 35.6 |
| Ours: kl-r-greedy, Qwen | 44.8 | 30.4 |
| Ours: kl-r-greedy, Gemma | 16.3 | 8.3 |
| Ours: gap-greedy, Llama | 13.1 | 11.5 |
| GCG (Llama) + “Sure here’s” | 32.5 | 19.6 |
| GCG (Gemma) + “Sure here’s” | 32.7 | 25.0 |
| Model | Cands./prompt | Score-pick (top-1) | Random-pick (mean) | Score adv. |
|---|---|---|---|---|
| Llama-3.2-1B-Instruct | 1.8 | |||
| gemma-2b-it | 2.6 | |||
| Qwen2.5-0.5B-Instruct | 19.9 |
| Discovery model | Random-rank | Score-rank (full ) | |
|---|---|---|---|
| Llama-3.2-1B-Instruct | 40.0 | 50.0 | |
| Qwen2.5-0.5B-Instruct | 38.0 | 40.0 | |
| gemma-2b-it | 11.3 | 10.0 |
| Family | R+SH (8) | GCG+SH (8) | Ours (8) | Ours GCG |
|---|---|---|---|---|
| gemma | 21.1 | 48.1 | 52.3 | |
| llama | 76.4 | 78.7 | 86.4 | |
| qwen | 56.4 | 63.7 | 82.8 | |
| mean | 51.3 | 63.5 | 73.8 | +10.4 |
| Model | Seen (n=50) | Unseen (n=470) | (pp) |
|---|---|---|---|
| Qwen3-0.6B | 100.0 | 93.8 | |
| Qwen2.5-0.5B-Instruct | 94.0 | 95.3 | |
| Llama-3.2-1B-Instruct | 94.0 | 91.1 | |
| gemma-2b-it | 44.0 | 64.7 | |
| Llama-3.2-3B-Instruct | 82.0 | 87.9 | |
| Qwen2.5-7B-Instruct | 84.0 | 82.8 |
| ASR after LG3 | Block rate (%) | |||||
|---|---|---|---|---|---|---|
| Model | R+SH | GCG+SH | Disc. | R+SH | GCG+SH | Disc. |
| Qwen3-0.6B | 9.8 | 11.9 | 11.3 | 88.3 | 89.1 | 89.5 |
| Qwen2.5-0.5B-Instruct | 4.6 | 6.5 | 12.9 | 95.0 | 93.4 | 88.5 |
| Llama-3.2-1B-Instruct | 5.8 | 8.7 | 9.8 | 92.7 | 90.1 | 76.2 |
| gemma-2b-it | 2.9 | 4.8 | 9.4 | 22.9 | 56.9 | 37.6 |
| Llama-3.2-3B-Instruct | 5.6 | 7.9 | 10.0 | 85.0 | 79.2 | 75.5 |
| Reward signal | Boundary median | Mid-clause median | Mann-Whitney |
|---|---|---|---|
| Our proxy | logit | logit | |
| PKU-SafeRLHF ( cost) delta | |||
| Skywork-Reward-Llama-3.1-8B delta |
| Variant | Qwen-0.5B | Llama-1B | gemma-2b |
|---|---|---|---|
| Reported | 4.4 / 99% | 12.7 / 100% | 14.3 / 100% |
| Expanded | 4.4 / 99% / 1.00 | 12.7 / 100% / 1.00 | 14.3 / 100% / 1.00 |
| Minimal (3+3) | 1.8 / 85% / 0.85 | / 52% / 0.82 | 0.9 / 79% / 0.62 |
| Leading-space normalized | 0.9 / 86% / 0.64 | 5.2 / 96% / 0.73 | 2.1 / 91% / 0.48 |
| Evaluation | Qwen2.5-0.5B | Qwen2.5-7B |
|---|---|---|
| Trained DFS suffixes, base SFT | 88 32% | 88 0% |
| Held-out greedy suffixes, base SFT | 83 63.8% | 30.5 1.5% |
| Re-discovery on patched model | 56.2% | 0.0% |
| Benign refusal, base SFT | 5 8% | 0 1.7% |