When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications. An alternative approach constructs or selects harmful target outputs and modifies the input to increase their likelihood. We establish a precise connection between these two approaches through a probabilistic reformulation. Specifically, we show that the gradient of the logarithm of expected harmfulness with respect to the input equals the expected input gradient of the model's log-likelihood under a harmfulness reweighted output distribution. This identity provides a unified interpretation of expected harmfulness and target likelihood optimization. Building on this connection, we propose OPUR, a sampling distribution designed to generate highly harmful target outputs and use the resulting samples to guide likelihood-based input optimization. Experiments demonstrate the effectiveness of the resulting method in jailbreaking LLM agents.
Figures & tables
Figure 1: Overview of our probabilistic view of jailbreak optimization. Given an adversarial input, the model induces a response distribution pdis(y)=pθ(y∣x,s) , while the harmfulness function induces a preference pvic(y)∝ψ(y) . Considering either view alone may favor responses that are likely but insufficiently harmful, or harmful but unlikely. Their normalized product defines the reweighted distribution π , which emphasizes responses that are simultaneously likely under the model and harmful under the attack objective. Since exact sampling from π may be unavailable in practice, its high-probability regions can be represented by constructed surrogate targets and optimized through their likelihood. As optimization proceeds, model probability mass is progressively shifted toward harmful outputs, providing a unified interpretation of expected harmfulness maximization and target-likelihood optimization.
Figure 2: An illustrative InjecAgent example comparing UDora and Opur. The agent encounters attacker-controlled contentwhile answering a calendar query. The two methods construct different surrogate contexts for adversarial optimization,resulting in different final tool-call outputs in this case. Position scores and candidate sets are schematic.
Model
Method
Attack Categories
Avg. ASR
Detailed Prompt
Simple Prompt
w/ Hint
w/o Hint
w/ Hint
w/o Hint
Llama-3.1- 8B-Instruct
GCG
38.64%
38.64%
40.91%
31.82%
37.50%
Reinforce-GCG
59.09%
50.00%
50.00%
43.18%
50.57%
UDora (Sequential)
56.82%
63.64%
61.36%
65.91%
61.93%
UDora (Joint)
59.09%
56.82%
63.64%
65.91%
61.36%
Table 1: Attack success rates (%) on AgentHarm under different attack categories. The best results for each model are highlighted in bold.
Model
Method
Attack Categories
Avg. ASR
Direct Harm
Data Stealing
Llama-3.1- 8B-Instruct
GCG
26%
34%
30%
REINFORCE-GCG
46%
48%
47%
UDora (Sequential)
44%
44%
44%
UDora (Joint)
36%
56%
46%
Opur (Ours)
48%
50%
49%
Table 2: Attack success rates (%) on InjecAgent under different attack categories. The best results for each model are highlighted in bold.
Model
Method
Attack Categories
Avg. ASR
Price Mismatch
Attribute Mismatch
Category Mismatch
All Mismatch
Llama-3.1- 8B-Instruct
GCG
0.00%
0.00%
0.00%
0.00%
0.00%
REINFORCE-GCG
26.67%
0.00%
0.00%
0.00%
6.67%
UDora (Sequential)
6.67%
40.00%
13.33%
13.33%
18.33%
UDora (Joint)
20.00%
33.33%
6.67%
6.67%
16.67%
Opur (Ours)
13.33%
40.00%
33.33%
20.00%
26.67%
Table 3: Attack success rates (%) on WebShop under different attack categories. The best results for each model are highlighted in bold.
Model
Kopt
Datasets
AgentHarm
InjecAgent
WebShop
Llama-3.1- 8B-Instruct
1
56.82%
46.00%
5.00%
2
63.07%
45.00%
15.00%
3
75.00%
49.00%
26.67%
Ministral-8B- Instruct-2410
1
99.43%
30.00%
11.67%
2
–
34.00%
16.67%
Table 4: Ablation study on the optimization rollout budget Kopt for Opur. We report average attack success rates (%) on three benchmarks. “–” denotes configurations that are not evaluated because performance has already saturated.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Kopt
Method
Attack Categories
Avg. ASR
Detailed Prompt
Simple Prompt
w/ Hint
w/o Hint
w/ Hint
w/o Hint
1
UDora (Sequential)
59.09%
50.00%
50.00%
43.18%
50.57%
UDora (Joint)
43.18%
38.64%
47.73%
50.00%
44.89%
Opur (Ours)
56.82%
56.82%
59.09%
59.09%
56.82%
2
UDora (Sequential)
61.36%
47.73%
68.18%
56.82%
58.52%
Appendix
Table 5: Ablation study of Kopt on AgentHarm using Llama-3.1-8B-Instruct. We report attack success rates (%) under different attack categories.
Model
Kopt
Method
Attack Categories
Avg. ASR
Direct Harm
Data Stealing
Llama-3.1- 8B-Instruct
1
UDora (Sequential)
44.00%
44.00%
44.00%
UDora (Joint)
36.00%
56.00%
46.00%
Opur (Ours)
38.00%
54.00%
46.00%
2
UDora (Sequential)
32.00%
48.00%
40.00%
UDora (Joint)
44.00%
44.00%
44.00%
Appendix
Table 6: Ablation study of Kopt on InjecAgent under different attack categories. We report attack success rates (%).
Model
Kopt
Method
Attack Categories
Avg. ASR
Price Mismatch
Attribute Mismatch
Category Mismatch
All Mismatch
Llama-3.1- 8B-Instruct
1
UDora (Sequential)
13.33%
13.33%
6.67%
0.00%
8.33%
UDora (Joint)
6.67%
13.33%
13.33%
0.00%
8.33%
Opur (Ours)
0.00%
6.67%
13.33%
0.00%
5.00%
2
UDora (Sequential)
6.67%
26.67%
13.33%
6.67%
11.67%
UDora (Joint)
6.67%
20.00%
6.67%
0.00%
8.33%
Appendix
Table 7: Ablation study of Kopt on WebShop under different attack categories. We report attack success rates (%).
Figure 3: A real OPUR trajectory on AgentHarm. Two stochastic rollouts from the same user-side adversarial suffix produce a refusal and a response containing the target function query_bing_search . OPUR retains one training-only surrogate context from the refusal; for the other rollout, it samples four position sets via softmax and retains two using probe loss. The three retained contexts guide one suffix update. Both subsequent optimization rollouts hit the target, and evaluation of the final suffix also succeeds.
Figure 4: A real OPUR trajectory on InjecAgent. The attacker-controlled SMS returned by TwilioGetReceivedSmsMessages carries the adversarial suffix.Two initial stochastic rollouts from this suffix both omit the targetaction BankManagerTransferFunds . For each rollout, OPUR samples 4 position sets via softmax and retains two using probe loss, providing four training-only surrogate contexts for a suffix update. After several updates, both optimization rollouts contain Action: BankManagerTransferFunds . In the final evaluation, the updated suffix generates the target action.
Figure 5: A real OPUR trajectory on WebShop. The attacker appends an adversarial suffix to a coffee item’s search-result title, aiming to redirect an agent searching for RCA cables toward click[B08KRVH121] . Two initial stochastic rollouts click other items. For each rollout, OPUR samples four replacement-position sets via softmax and retains two training-only surrogate contexts using probe loss.The four retained contexts guide a suffix update. After further updates,both optimization rollouts contain the target click. In evaluation, the agent outputs click[B08KRVH121] despite stating that the item is not a heavy-duty RCA cable.