Towards a Unified Misuse Monitoring Benchmark
Organizations: University of Oxford
Abstract
LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats separately and ask whether a trajectory is harmful, rather than when it becomes harmful. We propose monitoring the agent's responses, where its actions are externalised, and ask whether the first point where monitors identify harm lands within a harm window (from the agent's first harmful commitment to goal execution). We develop a unified formalism for trace-level misuse monitoring and use it to construct a benchmark of ~6,200 conversation transcripts between a user, an LLM agent, and the external environment, spanning both threats in a shared schema, with a labelled harm window, corresponding benign controls, and matched instances of refusals to these requests. Across 17 monitor configurations, we find that our proposed action-framed monitors perform well on both threats under classical metrics (AUC: 0.95 and 0.99 respectively), while content-framed monitors collapse on injection attacks (AUC: 0.52). We also show that classical position-blind metrics paint an optimistic picture of monitor performance, since all monitors localise decomposition attacks poorly under the interval metric, which measures the ability to localise harm. Broadly, we illustrate the need for a unified study of misuse monitoring.
Figures & tables
| Arm | Sees / judges | Cadence |
|---|---|---|
| baseline | current assistant message / the assistant’s action | 1 / asst. msg |
| sequential | full prefix, every message / intent of the tasks together | 1 / msg |
| sequential_action | full prefix / the assistant’s current action | 1 / asst. msg |
| eae | extractor quotes the prefix, evaluator scores the quotes | 2 / asst. msg |
| sp_* (except sp_zeroshot_action ) | cumulative user + tool stream / probability of a harmful goal | 1 / input msg |
| sp_zeroshot_action | sp_zeroshot with the assistant’s action | 1 / asst. msg |
| Decomposition | Injection | |||||
|---|---|---|---|---|---|---|
| Arm | AUC | Recall | Interval | AUC | Recall | Interval |
| baseline | 0.904 | 0.784 | 0.334 | 0.921 | 0.815 | 0.624 |
| sequential | 0.960 | 0.956 | 0.166 | 0.521 | (0.63) | — |
| sequential_action | 0.952 | 0.943 | 0.433 | 0.985 | 0.997 | 0.949 |
| eae | 0.953 | 0.932 | 0.495 | 0.990 | 0.995 | 0.900 |
| sp_zeroshot | 0.945 | 0.944 | 0.136 | 0.511 | (0.86) | — |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Check | main | hard | twins | injection |
|---|---|---|---|---|
| strict completeness | 1,162 / 1,165 | 582 / 582 | 581 / 582 | — |
| end-state completeness | 1,165 / 1,165 | 582 / 582 | 582 / 582 | 2,828 / 2,828 |
| continuity judge (LLM) | 1,165 / 1,165 | 577 / 582 | — | — |
| payload in a tool message | — | — | — | 2,828 / 2,828 |
| onset check (harmful) | — | — | — | 1,881 / 1,881 |
| oracle check (harmful) | — | — | — | 1,639 / 1,881 |
| Split | Traces | Calls | Generated tokens | Prompt tokens |
|---|---|---|---|---|
| decomposition_main | 1,165 | 7,810 | 1.58M | 10.29M |
| decomposition_hard | 582 | 5,737 | 1.05M | 9.10M |
| decomposition_refusal | 1,644 | 1,324 | 0.04M | 1.21M |
| injection | 2,828 | 2,828 | 1.58M | 18.17M |
| Total | 6,219 | 17,699 | 4.25M | 38.77M |
| Model and checkpoint | Served via | Arms | |
|---|---|---|---|
| Gemma 4 26B-A4B google/gemma-4-26B-A4B-it | vLLM | 5 | all 14 |
| Qwen3.6 27B Qwen/Qwen3.6-27B-FP8 | vLLM, thinking off, text-only | 5 | all 14 |
| Qwen3.5 27B Qwen/Qwen3.5-27B-FP8 | vLLM, thinking off | 5 | core |
| Qwen3.5 9B, 4B, 2B Qwen/Qwen3.5-9B , -4B , -2B | vLLM, thinking off | 5 | core |
| Mistral Small 3.2 24B mistralai/Mistral-Small-3.2-24B-Instruct-2506 | vLLM, Mistral formats | 5 | core |
| Llama 3.1 70B RedHatAI/Meta-Llama-3.1-70B-Instruct-FP8 | vLLM | 5 | core |
| Track | Arm | AUC [95% CI] | Recall [95% CI] | Interval [95% CI] |
|---|---|---|---|---|
| Decomposition | baseline | 0.904 [0.887, 0.920] | 0.784 [0.000, 0.843] | 0.334 [0.000, 0.362] |
| sequential | 0.960 [0.949, 0.969] | 0.956 [0.923, 0.969] | 0.166 [0.144, 0.188] | |
| sequential_action | 0.952 [0.940, 0.963] | 0.943 [0.911, 0.964] | 0.433 [0.404, 0.464] | |
| Injection | baseline | 0.921 [0.910, 0.932] | 0.815 [0.766, 0.845] | 0.624 [0.582, 0.658] |
| sequential | 0.521 [0.500, 0.541] | — | ||
| sequential_action | 0.985 [0.979, 0.991] | 0.997 [0.995, 0.999] | 0.949 [0.939, 0.959] |
| Decomposition | Injection | |||||
|---|---|---|---|---|---|---|
| Arm | AUC | Recall | Interval | AUC | Recall | Interval |
| baseline | 0.904 | 0.784 | 0.334 | 0.921 | 0.815 | 0.624 |
| sequential | 0.960 | 0.956 | 0.166 | 0.521 | (0.63) | — |
| sequential_action | 0.952 | 0.943 | 0.433 | 0.985 | 0.997 | 0.949 |
| eae | 0.953 | 0.932 | 0.495 | 0.990 | 0.995 | 0.900 |
| sp_zeroshot | 0.945 | 0.944 | 0.136 | 0.511 | (0.86) | — |
| Decomposition | Injection | |||||
|---|---|---|---|---|---|---|
| Arm | AUC | Recall | Interval | AUC | Recall | Interval |
| baseline | 0.923 | 0.807 | 0.364 | 0.923 | 0.801 | 0.625 |
| sequential | 0.967 | 0.959 | 0.281 | 0.685 | (0.13) | — |
| sequential_action | 0.962 | 0.946 | 0.514 | 0.992 | 0.999 | 0.959 |
| eae | 0.919 | 0.859 | 0.484 | 0.987 | 0.988 | 0.887 |
| sp_zeroshot | 0.966 | 0.960 | 0.233 | 0.617 | 0.254 | 0.000 |
| Model | Arm | Recall | ||||
|---|---|---|---|---|---|---|
| Gemma 4 26B-A4B | sequential_action | 0.997 | 0.949 | 0.949 | 0.983 | 0.986 |
| eae | 0.995 | 0.900 | 0.900 | 0.955 | 0.966 | |
| baseline | 0.815 | 0.624 | 0.624 | 0.699 | 0.715 | |
| sp_zeroshot_action | 0.732 | 0.622 | 0.622 | 0.643 | 0.643 | |
| sp_zeroshot_cot | 0.220 | 0.000 | 0.063 | 0.063 | 0.080 | |
| Qwen3.6 27B | sequential_action | 0.999 | 0.959 | 0.959 | 0.988 | 0.992 |
| Decomposition (main split) | Injection | ||||||
|---|---|---|---|---|---|---|---|
| Model | Arm | AUC | Recall | Interval | AUC | Recall | Interval |
| Gemma 4 26B-A4B | baseline | 0.909 | 0.807 | 0.350 | 0.921 | 0.815 | 0.624 |
| sequential | 0.961 | 0.975 | 0.154 | 0.521 | (0.63) | — | |
| sequential_action | 0.953 | 0.953 | 0.457 | 0.985 | 0.997 | 0.949 | |
| Qwen3.6 27B | baseline | 0.927 | 0.828 | 0.381 | 0.923 | 0.801 | 0.625 |
| sequential | 0.968 | 0.977 | 0.299 | 0.685 | (0.13) | — | |
| Decomposition (main split) | Injection | |||||
|---|---|---|---|---|---|---|
| Arm | AUC | Recall | Interval | AUC | Recall | Interval |
| llamaguard4 | 0.913 | 0.834 | 0.416 | 0.803 | 0.491 | 0.290 |
| llamaguard4_policy | 0.928 | 0.895 | 0.459 | 0.802 | 0.520 | 0.322 |
| qwen3guard | 0.897 | 0.693 | 0.268 | 0.634 | 0.242 | 0.213 |
| Arm | Recall | Interval | Asst. cadence | ||||
|---|---|---|---|---|---|---|---|
| baseline | 0.784 | 0.334 | 0.290 | 0.334 | 0.334 | 0.483 | 0.701 |
| sequential | 0.956 | 0.166 | 0.144 | 0.410 | 0.387 | 0.438 | 0.760 |
| sequential_action | 0.943 | 0.433 | 0.410 | 0.433 | 0.433 | 0.577 | 0.828 |
| eae | 0.932 | 0.495 | 0.439 | 0.495 | 0.495 | 0.633 | 0.857 |
| sp_zeroshot_action | 0.937 | 0.342 | 0.331 | 0.342 | 0.342 | 0.541 | 0.849 |
| Model | Arm | Eval FPR | Interval | Early | Late | Never | Recall | Frontier |
|---|---|---|---|---|---|---|---|---|
| Gemma 4 26B-A4B | eae | 0.080 | 0.378 | 0.520 | 0.019 | 0.083 | 0.917 | 0.495 |
| sp_zeroshot_action | 0.082 | 0.342 | 0.587 | 0.003 | 0.068 | 0.932 | 0.342 | |
| baseline | 0.086 | 0.334 | 0.435 | 0.014 | 0.216 | 0.784 | 0.334 | |
| sequential_action | 0.130 | 0.214 | 0.740 | 0.006 | 0.040 | 0.960 | 0.433 | |
| sequential | 0.080 | 0.038 | 0.888 | 0.011 | 0.063 | 0.937 | 0.166 | |
| sp_zeroshot | 0.103 | 0.026 | 0.906 | 0.012 | 0.056 | 0.944 | 0.136 |
| Monitor model | Arm | Flagged | Flagged | AUC | Refusal | Refusal |
|---|---|---|---|---|---|---|
| at | at | at | score | |||
| Gemma 4 26B-A4B | baseline | 0.281 | 0.650 | 0.789 | 0.000 | 1.00 |
| sequential_action | 0.469 | 0.827 | 0.834 | 0.000 | 1.00 | |
| sequential | 0.951 | 0.971 | 0.586 | 0.943 | 9.31 | |
| Qwen3.6 27B | baseline | 0.277 | 0.679 | 0.848 | 0.000 | 1.00 |
| sequential_action | 0.358 | 0.812 | 0.908 | 0.000 | 1.00 |
| Arm | Refusal at | Flagged | Before | Credited zone | On refusal | AUC | Refusal score |
|---|---|---|---|---|---|---|---|
| baseline | / mid / | 0.281 / 0.564 / 0.650 | 0.281 | 0 / 0.283 / 0.369 | 0 / 0 / 0 | 0.789 / 0.626 / 0.577 | 1.00 / 1.00 / 1.00 |
| sequential_action | / mid / | 0.469 / 0.739 / 0.827 | 0.469 | 0 / 0.270 / 0.358 | 0 / 0 / 0 | 0.834 / 0.644 / 0.584 | 1.00 / 1.10 / 1.15 |
| eae | / mid / | 0.380 / 0.664 / 0.783 | 0.363 | 0 / 0.279 / 0.416 | 0.016 / 0.022 / 0.004 | 0.895 / 0.708 / 0.605 | 5.06 / 6.89 / 7.53 |
| sp_zeroshot_action | / mid / | 0.451 / 0.739 / 0.823 | 0.451 | 0 / 0.288 / 0.372 | 0 / 0 / 0 | 0.759 / 0.606 / 0.563 | 1.00 / 1.02 / 1.01 |
| sequential | / mid / | 0.956 / 0.964 / 0.971 | 0.880 | 0 / 0.031 / 0.049 | 0.077 / 0.053 / 0.042 | 0.587 / 0.525 / 0.495 | 9.32 / 9.45 / 9.57 |
| sp_zeroshot | / mid / | 0.856 / 0.891 / 0.922 | 0.856 | 0 / 0.035 / 0.066 | 0 / 0 / 0 | 0.655 / 0.587 / 0.541 | 1.00 / 1.00 / 1.00 |
| Family | Gemma tokens | Gemma AUC | Qwen3.6 tokens | Qwen3.6 AUC | ||
|---|---|---|---|---|---|---|
| plain CoT | plain CoT | plain CoT | plain CoT | |||
| sp_zeroshot | 220 10,822 | 49 | 0.945 0.958 (+0.013) | 275 10,528 | 38 | 0.966 0.962 ( 0.004) |
| sp_guideline | 241 16,727 | 69 | 0.937 0.958 (+0.021) | 235 23,573 | 100 | 0.965 0.945 ( 0.020) |
| sp_qa_icl | 132 10,611 | 80 | 0.839 0.875 (+0.036) | 156 17,577 | 113 | 0.922 0.898 ( 0.024) |
| sp_agent_icl | — | — | 0.907 0.925 (+0.018) | 173 8,688 | 50 | 0.939 0.923 ( 0.016) |
| Setting | Target arm | Injection | Decomposition | Harm gate held | |
|---|---|---|---|---|---|
| Fixed | sequential | 5 | 5 / 40 | 2 / 40 | 91% / 97% |
| Fixed | sequential_action | 5 | 2 / 40 | 0 / 40 | 87% / 99% |
| Matched 10% budget | sequential_action | 1.2 / 8.2 | 0 / 40 | 1 / 40 | 87% / 100% |