When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications. An alternative approach constructs or selects harmful target outputs and modifies the input to increase their likelihood. We establish a precise connection between these two approaches through a probabilistic reformulation. Specifically, we show that the gradient of the logarithm of expected harmfulness with respect to the input equals the expected input gradient of the model's log-likelihood under a harmfulness reweighted output distribution. This identity provides a unified interpretation of expected harmfulness and target likelihood optimization. Building on this connection, we propose OPUR, a sampling distribution designed to generate highly harmful target outputs and use the resulting samples to guide likelihood-based input optimization. Experiments demonstrate the effectiveness of the resulting method in jailbreaking LLM agents.
Figures & tables
Figure 1: Overview of our probabilistic view of jailbreak optimization. Given an adversarial input, the model induces a response distribution pdis(y)=pθ(y∣x,s) , while the harmfulness function induces a preference pvic(y)∝ψ(y) . Considering either view alone may favor responses that are likely but insufficiently harmful, or harmful but unlikely. Their normalized product defines the reweighted distribution π , which emphasizes responses that are simultaneously likely under the model and harmful under the attack objective. Since exact sampling from π may be unavailable in practice, its high-probability regions can be represented by constructed surrogate targets and optimized through their likelihood. As optimization proceeds, model probability mass is progressively shifted toward harmful outputs, providing a unified interpretation of expected harmfulness maximization and target-likelihood optimization.
Figure 2: An illustrative InjecAgent example comparing UDora and Opur. The agent encounters attacker-controlled contentwhile answering a calendar query. The two methods construct different surrogate contexts for adversarial optimization,resulting in different final tool-call outputs in this case. Position scores and candidate sets are schematic.
Model
Method
Attack Categories
Avg. ASR
Detailed Prompt
Simple Prompt
w/ Hint
w/o Hint
w/ Hint
w/o Hint
Llama-3.1- 8B-Instruct
GCG
38.64%
38.64%
40.91%
31.82%
37.50%
Reinforce-GCG
59.09%
50.00%
50.00%
43.18%
50.57%
UDora (Sequential)
56.82%
63.64%
61.36%
65.91%
61.93%
UDora (Joint)
59.09%
56.82%
63.64%
65.91%
61.36%
Table 1: Attack success rates (%) on AgentHarm under different attack categories. The best results for each model are highlighted in bold.
Model
Method
Attack Categories
Avg. ASR
Direct Harm
Data Stealing
Llama-3.1- 8B-Instruct
GCG
26%
34%
30%
REINFORCE-GCG
46%
48%
47%
UDora (Sequential)
44%
44%
44%
UDora (Joint)
36%
56%
46%
Opur (Ours)
48%
50%
49%
Table 2: Attack success rates (%) on InjecAgent under different attack categories. The best results for each model are highlighted in bold.
Model
Method
Attack Categories
Avg. ASR
Price Mismatch
Attribute Mismatch
Category Mismatch
All Mismatch
Llama-3.1- 8B-Instruct
GCG
0.00%
0.00%
0.00%
0.00%
0.00%
REINFORCE-GCG
26.67%
0.00%
0.00%
0.00%
6.67%
UDora (Sequential)
6.67%
40.00%
13.33%
13.33%
18.33%
UDora (Joint)
20.00%
33.33%
6.67%
6.67%
16.67%
Opur (Ours)
13.33%
40.00%
33.33%
20.00%
26.67%
Table 3: Attack success rates (%) on WebShop under different attack categories. The best results for each model are highlighted in bold.
Model
Kopt
Datasets
AgentHarm
InjecAgent
WebShop
Llama-3.1- 8B-Instruct
1
56.82%
46.00%
5.00%
2
63.07%
45.00%
15.00%
3
75.00%
49.00%
26.67%
Ministral-8B- Instruct-2410
1
99.43%
30.00%
11.67%
2
–
34.00%
16.67%
Table 4: Ablation study on the optimization rollout budget Kopt for Opur. We report average attack success rates (%) on three benchmarks. “–” denotes configurations that are not evaluated because performance has already saturated.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Kopt
Method
Attack Categories
Avg. ASR
Detailed Prompt
Simple Prompt
w/ Hint
w/o Hint
w/ Hint
w/o Hint
1
UDora (Sequential)
59.09%
50.00%
50.00%
43.18%
50.57%
UDora (Joint)
43.18%
38.64%
47.73%
50.00%
44.89%
Opur (Ours)
56.82%
56.82%
59.09%
59.09%
56.82%
2
UDora (Sequential)
61.36%
47.73%
68.18%
56.82%
58.52%
Appendix
Table 5: Ablation study of Kopt on AgentHarm using Llama-3.1-8B-Instruct. We report attack success rates (%) under different attack categories.
Model
Kopt
Method
Attack Categories
Avg. ASR
Direct Harm
Data Stealing
Llama-3.1- 8B-Instruct
1
UDora (Sequential)
44.00%
44.00%
44.00%
UDora (Joint)
36.00%
56.00%
46.00%
Opur (Ours)
38.00%
54.00%
46.00%
2
UDora (Sequential)
32.00%
48.00%
40.00%
UDora (Joint)
44.00%
44.00%
44.00%
Appendix
Table 6: Ablation study of Kopt on InjecAgent under different attack categories. We report attack success rates (%).
Model
Kopt
Method
Attack Categories
Avg. ASR
Price Mismatch
Attribute Mismatch
Category Mismatch
All Mismatch
Llama-3.1- 8B-Instruct
1
UDora (Sequential)
13.33%
13.33%
6.67%
0.00%
8.33%
UDora (Joint)
6.67%
13.33%
13.33%
0.00%
8.33%
Opur (Ours)
0.00%
6.67%
13.33%
0.00%
5.00%
2
UDora (Sequential)
6.67%
26.67%
13.33%
6.67%
11.67%
UDora (Joint)
6.67%
20.00%
6.67%
0.00%
8.33%
Appendix
Table 7: Ablation study of Kopt on WebShop under different attack categories. We report attack success rates (%).
Figure 3: A real OPUR trajectory on AgentHarm. Two stochastic rollouts from the same user-side adversarial suffix produce a refusal and a response containing the target function query_bing_search . OPUR retains one training-only surrogate context from the refusal; for the other rollout, it samples four position sets via softmax and retains two using probe loss. The three retained contexts guide one suffix update. Both subsequent optimization rollouts hit the target, and evaluation of the final suffix also succeeds.
Figure 4: A real OPUR trajectory on InjecAgent. The attacker-controlled SMS returned by TwilioGetReceivedSmsMessages carries the adversarial suffix.Two initial stochastic rollouts from this suffix both omit the targetaction BankManagerTransferFunds . For each rollout, OPUR samples 4 position sets via softmax and retains two using probe loss, providing four training-only surrogate contexts for a suffix update. After several updates, both optimization rollouts contain Action: BankManagerTransferFunds . In the final evaluation, the updated suffix generates the target action.
Figure 5: A real OPUR trajectory on WebShop. The attacker appends an adversarial suffix to a coffee item’s search-result title, aiming to redirect an agent searching for RCA cables toward click[B08KRVH121] . Two initial stochastic rollouts click other items. For each rollout, OPUR samples four replacement-position sets via softmax and retains two training-only surrogate contexts using probe loss.The four retained contexts guide a suffix update. After further updates,both optimization rollouts contain the target click. In evaluation, the agent outputs click[B08KRVH121] despite stating that the item is not a heavy-duty RCA cable.
Jailbreak attacks expose persistent safety weaknesses in large language models (LLMs), but existing stateless single-turn methods face a trade-off: hand-crafted prompts are expressive but static, while iterative prompt optimization can adapt but often relies on low-level mutations that require many target queries. We propose JailbreakOPT, a tool-assisted framework for improving iterative single-turn jailbreak prompt optimization. JailbreakOPT organizes diverse atomic jailbreak prompts into an attack tool library and composes them through a unified intra-episode optimization abstraction to generate stronger standalone attack prompts. To reuse experience across attack episodes, JailbreakOPT further frames tool selection as a contextual bandit problem and applies contextual Thompson sampling to guide exploration and exploitation based on past outcomes. Experiments across multiple target LLMs and attack goals show that JailbreakOPT improves attack success rate (ASR) while reducing the number of attacks until success (No.A) compared with atomic single-turn attacks and existing iterative optimization baselines. This paper may contain offensive or harmful content.
Ge Shi, Jun Yin, Donglin Xie +3
University of California, Davis · The Renmin University of China · Independent Researcher +3
Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment risk assessment and defense comparison unreliable. Historically, standardized attacks such as AutoAttack have largely resolved this for image classifiers, providing a reliable evaluation baseline for systematic comparison across defenses. However, no equivalent exists for LLM jailbreak evaluation yet, where designing such an attack is considerably more difficult. A reliable attack must, among other things, be black-box compatible, applicable to arbitrary defense pipelines, and efficient, which no existing method jointly satisfies. We introduce Indirect Harm Optimization (IHO), a masked diffusion language model attacker trained via iterative preference optimization against a harmfulness judge, requiring only black-box access to the target. The same method can be used without modification as a strong adaptive attack on individual behaviors, or as an efficient amortized policy that transfers to held-out behaviors and unseen target models without fine-tuning. Even against layered defenses, such as a Circuit Breaker-trained model combined with an auxiliary detector, IHO improves attack success considerably over state-of-the-art approaches, without any defense-specific adaptation. Our results position IHO as a practical step toward the kind of standardized jailbreak evaluation that has improved reliability in the past. Code and models are available on GitHub and Hugging Face.
Vincent Limbach, Jonas Dornbusch, David Lüdke +2
Department of Computer Science, Technical University of Munich, Germany · Munich Data Science Institute · Munich Center for Machine Learning +1
LLMs are trained to refuse harmful instructions, but do they truly understand harmfulness beyond just refusing? Prior work has shown that LLMs' refusal behaviors can be mediated by a one-dimensional subspace, i.e., a refusal direction. In this work, we identify a new dimension to analyze safety mechanisms in LLMs, i.e., harmfulness, which is encoded internally as a separate concept from refusal. There exists a harmfulness direction that is distinct from the refusal direction. As causal evidence, steering along the harmfulness direction can lead LLMs to interpret harmless instructions as harmful, but steering along the refusal direction tends to elicit refusal responses directly without reversing the model's judgment on harmfulness. Furthermore, using our identified harmfulness concept, we find that certain jailbreak methods work by reducing the refusal signals without reversing the model's internal belief of harmfulness. We also find that adversarially finetuning models to accept harmful instructions has minimal impact on the model's internal belief of harmfulness. These insights lead to a practical safety application: The model's latent harmfulness representation can serve as an intrinsic safeguard (Latent Guard) for detecting unsafe inputs and reducing over-refusals that is robust to finetuning attacks. For instance, our Latent Guard achieves performance comparable to or better than Llama Guard 3 8B, a dedicated finetuned safeguard model, across different jailbreak methods. Our findings suggest that LLMs' internal understanding of harmfulness is more robust than their refusal decision to diverse input instructions, offering a new perspective to study AI safety.