cs.AIOct 7, 2026

From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents

Authors: Juanyang Xu, Zheng Wang, Xingyu Zhao, Siddartha Khastgir, Andi Zhang

Organizations: University of Macau · WMG, University of Warwick · Wuhan University

Abstract

When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications. An alternative approach constructs or selects harmful target outputs and modifies the input to increase their likelihood. We establish a precise connection between these two approaches through a probabilistic reformulation. Specifically, we show that the gradient of the logarithm of expected harmfulness with respect to the input equals the expected input gradient of the model's log-likelihood under a harmfulness reweighted output distribution. This identity provides a unified interpretation of expected harmfulness and target likelihood optimization. Building on this connection, we propose OPUR, a sampling distribution designed to generate highly harmful target outputs and use the resulting samples to guide likelihood-based input optimization. Experiments demonstrate the effectiveness of the resulting method in jailbreaking LLM agents.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization

    Jun 9, 2026Ge Shi, Jun Yin, Donglin Xie +3Adversarial Prompt GenerationLarge Language Model-Guided Optimization

  2. Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs

    Jun 2, 2026Vincent Limbach, Jonas Dornbusch, David Lüdke +2LLM Jailbreak AttacksLLM Safety Evaluation

  3. LLMs Encode Harmfulness and Refusal Separately

    Jul 16, 2025Jiachen Zhao, Jing Huang, Zhengxuan Wu +2LLM InterpretabilityLLM Refusal Behavior