Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.
Figures & tables
Figure 1: Overview of BlindBias . The text-only target proposes a base candidate, and a local prefix-risk model determines whether the controller should intervene. Bypassed candidates are accepted directly. At intervention positions, sampled continuations are mapped to local actions and combined with a uniform prior to reconstruct a distribution. BiasNet then adjusts this distribution with a gate-scaled residual. After consecutive bypasses, speculative execution requests a multi-token draft, verifies its prefixes locally, commits the accepted prefix, and resumes controlled decoding at the first position that requires intervention.
Target Model
Method
AdvBench
HarmBench
SORRY-Bench
Harm
Info
Harm
Info
Harm
Info
GLM-5
PAIR
1.34
0.80
1.73
0.95
1.69
0.83
GPTFuzz
3.42
2.41
2.81
1.89
2.40
1.59
LogiBreak
2.58
1.32
2.11
1.04
2.55
1.23
FlipAttack
4.11
2.82
3.89
2.39
4.45
2.71
BlindBias
4.29
2.88
4.14
2.74
3.83
2.50
Table 1: Jailbreak performance across different target models and benchmarks. Higher Harm Score and Harm Info Score indicate stronger jailbreak effectiveness. Bold denotes the highest mean for each target, benchmark, and metric. BlindBias uses the soft-gated, global-uniform configuration throughout.
Proxy prior
Relation
PPL ↓
Harm ↑
Info ↑
Qwen3-1.7B
same family
3.86
3.90
2.83
SmolLM2-1.7B
foreign
3.93
3.02
2.12
Gemma-3-1B
foreign
3.94
3.38
2.39
Shuffled prior
control
7.11
2.69
2.16
Table 2: Contextual-prior transfer to Qwen3-32B. PPL is reported within the proxy-calibration protocol; Harm and Info are averaged over 100 AdvBench prompts.
Reconstruction quality
Attack quality
Method
PPL ↓
Unseen NLL ↓
Harm ↑
Info ↑
Smoothed empirical counts
11.85
21.14
1.50
1.26
Uniform prior
7.07
14.50
3.03
2.16
Global unigram prior
5.81
12.34
2.78
1.94
Numerical log probabilities (ref.)
3.17
7.38
4.06
3.08
Table 3: Distribution-signal recovery and downstream attack quality on Qwen3-32B. Predictive metrics use held-out events from within-prefix sample splits; Harm and Info are averaged over 100 AdvBench prompts.
Pipeline
Harm ↑
Info ↑
API calls ↓
Active pos. ↓
Ungated
3.96
2.81
4000
100.0%
Hard gate
3.44
2.19
448
9.5%
Run-time soft
3.58
2.27
283
5.25%
Table 4: Quality–intervention trade-off of complete gating pipelines on 100 AdvBench prompts towards Gemini-3.5-Flash. Active positions measure BiasNet invocation frequency.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Block
Training rule
Epochs
Eval. weight
Margin ↑
Beats base ↑
A
No gate
75
616.0
-2.683
19.7%
A
Soft weights
75
74.2
-1.346
36.6%
B
Hard-active only
10
14.0
2.250
64.3%
B
Hard + first 3
10
29.0
1.040
54.5%
Appendix
Table 5: Gate-aware training diagnostics on 100 AdvBench prompts. Blocks group runs with equal training duration; evaluation weights and support vary by rule. Evaluation weight is summed loss weight rather than a token count.