Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.
Figures & tables
Figure 1: Overview of BlindBias . The text-only target proposes a base candidate, and a local prefix-risk model determines whether the controller should intervene. Bypassed candidates are accepted directly. At intervention positions, sampled continuations are mapped to local actions and combined with a uniform prior to reconstruct a distribution. BiasNet then adjusts this distribution with a gate-scaled residual. After consecutive bypasses, speculative execution requests a multi-token draft, verifies its prefixes locally, commits the accepted prefix, and resumes controlled decoding at the first position that requires intervention.
Target Model
Method
AdvBench
HarmBench
SORRY-Bench
Harm
Info
Harm
Info
Harm
Info
GLM-5
PAIR
1.34
0.80
1.73
0.95
1.69
0.83
GPTFuzz
3.42
2.41
2.81
1.89
2.40
1.59
LogiBreak
2.58
1.32
2.11
1.04
2.55
1.23
FlipAttack
4.11
2.82
3.89
2.39
4.45
2.71
BlindBias
4.29
2.88
4.14
2.74
3.83
2.50
Table 1: Jailbreak performance across different target models and benchmarks. Higher Harm Score and Harm Info Score indicate stronger jailbreak effectiveness. Bold denotes the highest mean for each target, benchmark, and metric. BlindBias uses the soft-gated, global-uniform configuration throughout.
Proxy prior
Relation
PPL ↓
Harm ↑
Info ↑
Qwen3-1.7B
same family
3.86
3.90
2.83
SmolLM2-1.7B
foreign
3.93
3.02
2.12
Gemma-3-1B
foreign
3.94
3.38
2.39
Shuffled prior
control
7.11
2.69
2.16
Table 2: Contextual-prior transfer to Qwen3-32B. PPL is reported within the proxy-calibration protocol; Harm and Info are averaged over 100 AdvBench prompts.
Reconstruction quality
Attack quality
Method
PPL ↓
Unseen NLL ↓
Harm ↑
Info ↑
Smoothed empirical counts
11.85
21.14
1.50
1.26
Uniform prior
7.07
14.50
3.03
2.16
Global unigram prior
5.81
12.34
2.78
1.94
Numerical log probabilities (ref.)
3.17
7.38
4.06
3.08
Table 3: Distribution-signal recovery and downstream attack quality on Qwen3-32B. Predictive metrics use held-out events from within-prefix sample splits; Harm and Info are averaged over 100 AdvBench prompts.
Pipeline
Harm ↑
Info ↑
API calls ↓
Active pos. ↓
Ungated
3.96
2.81
4000
100.0%
Hard gate
3.44
2.19
448
9.5%
Run-time soft
3.58
2.27
283
5.25%
Table 4: Quality–intervention trade-off of complete gating pipelines on 100 AdvBench prompts towards Gemini-3.5-Flash. Active positions measure BiasNet invocation frequency.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Block
Training rule
Epochs
Eval. weight
Margin ↑
Beats base ↑
A
No gate
75
616.0
-2.683
19.7%
A
Soft weights
75
74.2
-1.346
36.6%
B
Hard-active only
10
14.0
2.250
64.3%
B
Hard + first 3
10
29.0
1.040
54.5%
Appendix
Table 5: Gate-aware training diagnostics on 100 AdvBench prompts. Blocks group runs with equal training duration; evaluation weights and support vary by rule. Evaluation weight is summed loss weight rather than a token count.
Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuristic templates or exhaustive trials, lacking mechanistic interpretability and query efficiency. In this study, we investigate an intrinsic vulnerability in the safety mechanisms of LLMs, where safety alignment relies on a small set of sparsely distributed attention heads, leaving much of the representational space weakly monitored. We formalize this phenomenon with a mathematical jailbreaking model that characterizes the delicate boundary of effective text obfuscation and analytically explains observed jailbreak behaviors. Guided by this model, we propose Babel, an efficient black-box attack framework that exploits the identified safety gap through systematic obfuscation sampling with iterative, feedback-driven distribution refinement, enabling reliable and high-success jailbreak attacks without access to model internals. Comprehensive evaluations on frontier commercial models demonstrate that Babel achieves state-of-the-art attack success rates and superior query efficiency. Specifically, compared to state-of-the-art methods, Babel increases the attack success rate on GPT-4o from 41.33% to 82.67% and on Claude-3-5-haiku from 38.33% to 78.33% within an average of 40 queries, providing a robust red-teaming methodology for LLMs safety research.
Ziwei Wang, Jing Chen, Ruichao Liang +6
Wuhan University · Nanyang Technological University · Southeast University +1
Large Language Models (LLMs) have been extensively used across diverse domains, including virtual assistants, automated code generation, and scientific research. However, they remain vulnerable to jailbreak attacks, which manipulate the models into generating harmful responses despite safety alignment. Recent studies have shown that current safety-aligned LLMs undergo shallow safety alignment. In this work, we conduct an in-depth investigation into the underlying mechanism of this phenomenon and reveal that it manifests through learned ''safety trigger tokens'' that activate the model's safety patterns when paired with the specific input. Through both analysis and empirical verification, we further demonstrate the high similarity of the safety trigger tokens across different harmful inputs. Accordingly, we propose D-STT, a simple yet effective defense algorithm that identifies and explicitly decodes safety trigger tokens of the given safety-aligned LLM to activate the model's learned safety patterns. In this process, the safety trigger is constrained to a single token, which effectively preserves model usability by introducing minimum intervention in the decoding process. Extensive experiments across diverse jailbreak attacks and benign prompts demonstrate that D-STT significantly reduces output harmfulness while preserving model usability and incurring negligible response time overhead, outperforming ten baseline methods.
Haoran Gu, Handing Wang, Yi Mei +2
Xidian University · Victoria University of Wellington · Westlake University
Unlike regular tokens derived from existing text corpora, special tokens are artificially created to annotate structured conversations during the fine-tuning process of Large Language Models (LLMs). Serving as metadata of training data, these tokens play a crucial role in instructing LLMs to generate coherent and context-aware responses. We demonstrate that special tokens can be exploited to construct four attack primitives, with which malicious users can reliably bypass the internal safety alignment of online LLM services and circumvent state-of-the-art (SOTA) external content moderation systems simultaneously. Moreover, we found that addressing this threat is challenging, as aggressive defense mechanisms-such as input sanitization by removing special tokens entirely, as suggested in academia-are less effective than anticipated. This is because such defense can be evaded when the special tokens are replaced by regular ones with high semantic similarity within the tokenizer's embedding space. We systemically evaluated our method, named MetaBreak, on both lab environment and commercial LLM platforms. Our approach achieves jailbreak rates comparable to SOTA prompt-engineering-based solutions when no content moderation is deployed. However, when there is content moderation, MetaBreak outperforms SOTA solutions PAP and GPTFuzzer by 11.6% and 34.8%, respectively. Finally, since MetaBreak employs a fundamentally different strategy from prompt engineering, the two approaches can work synergistically. Notably, empowering MetaBreak on PAP and GPTFuzzer boosts jailbreak rates by 24.3% and 20.2%, respectively.