cs.CROct 5, 2026

Towards a Unified Misuse Monitoring Benchmark

Authors: Aniruddh Pramod, James Oldfield, Adel Bibi

Organizations: University of Oxford

Abstract

LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats separately and ask whether a trajectory is harmful, rather than when it becomes harmful. We propose monitoring the agent's responses, where its actions are externalised, and ask whether the first point where monitors identify harm lands within a harm window (from the agent's first harmful commitment to goal execution). We develop a unified formalism for trace-level misuse monitoring and use it to construct a benchmark of ~6,200 conversation transcripts between a user, an LLM agent, and the external environment, spanning both threats in a shared schema, with a labelled harm window, corresponding benign controls, and matched instances of refusals to these requests. Across 17 monitor configurations, we find that our proposed action-framed monitors perform well on both threats under classical metrics (AUC: 0.95 and 0.99 respectively), while content-framed monitors collapse on injection attacks (AUC: 0.52). We also show that classical position-blind metrics paint an optimistic picture of monitor performance, since all monitors localise decomposition attacks poorly under the interval metric, which measures the ability to localise harm. Broadly, we illustrate the need for a unified study of misuse monitoring.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Measuring Harmfulness of Computer-Using Agents

    Jul 31, 2025Aaron Xuxiang Tian, Ruofan Zhang, Janet Tang +3Large Language Model SafetyUnsafe

  2. Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

    Jun 11, 2026Zihao Wang, Yiming Li, Yutong Wu +8Prompt-Injection DetectorsIndirect Prompt Injection

  3. TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents

    Jun 5, 2026Vijitha Mittapalli, Shreyaa Jayant Dani, Satya Srujana Pilli +7TracesTrajectory-Level Credit