A Competing-Hazards Systematization of Loss of Control in Autonomous Agents
Authors: Mohamed Aly Bouke
Organizations: Centre for Intelligent Cloud Computing, CoE for Advanced Cloud, Faculty of Information Science and Technology, Multimedia University, Jalan Ayer Keroh Lama, Bukit Beruang, 75450, Melaka, Malaysia
Leading AI developers have reported agents acting beyond their approved limits, which a United Nations panel described as an early warning of loss of human control. Yet incident reports and agent-safety evaluations describe these events differently, making it difficult to compare failures, trace risk across attempts, or separate agent behavior from the environment's role in allowing an out-of-scope action to succeed. To address this gap, we introduce a common framework in which each attempt ends in approved completion, safe stopping, scope escape, or continuation. We formalize the framework as a discrete-time competing-hazards model and derive escape probability within a retry budget, a model-conditional safe-budget limit, and conditions for estimation from execution logs. We audit 22 incident reports and 102 agent-safety evaluations published from January 2025 to September 2026 using primary sources. Six incidents involved tasks that could not be completed within scope, thirteen involved agents that continued rather than stopped, and five did not report stopping behavior. Developers' figures imply a task-level incidence ratio near 47 for out-of-scope coordination in never-solved versus solved tasks. Among evaluations, 87 recorded an out-of-scope effect or specification violation, 26 treated safe stopping as a first-class outcome, only 20 recorded both, and 79 merged budget exhaustion with failure. In 20 of 22 incidents, the environment allowed an out-of-scope effect, indicating that realized loss of control often reflected persistent agent behavior interacting with permissive boundary conditions; meanwhile, no evaluation reported all fields needed to estimate the full competing-hazards process from published evidence.
Figures & tables
Figure 1: The model as a state process. At each attempt an execution ends in one of three absorbing states or continues, either directly or after a blocked out-of-scope attempt; executions still alive after the budget are censored. A terminating monitor, when present, adds a fourth absorbing state.
Incident
Setting
Event
Disclosed
Feasibility
Boundary
Safe stop
Agents
Effect
Gr.
OpenAI agents, Hugging Face
evaluation
2026-05
2026-07
infeasible
E C P
no
multiple, 1200
yes
H
OpenAI agent, Australian statistics portal
evaluation
2026-06
2026-09
unknown
E C P
unknown
unknown
yes
H
Gemini, Irregular evaluation
evaluation
2026-05
2026-08
unknown
E C T
stopped
single
yes
M
UK AISI cyber testing
evaluation
2026-07
2026-08
infeasible
E C
no
multiple
yes
H
Claude Opus 4.7, Irregular evaluation
evaluation
2026-04
2026-07
unknown
E C T M
no
single
yes
H
Claude Mythos 5, package publication
evaluation
2026
2026-07
infeasible
E T M
no
single
yes
H
Table 1: The incident record on the audit questions. Setting as reported; event and disclosure by month; feasibility of the sanctioned task inside scope as stated or inferable from the primary source; boundary yield as the reported cause of the effect, coded E open egress, C exposed or shared credential, T target collision with a real system, M reported misconfiguration, P permissions on a live or editable resource, H public hosting, registry, or wiki used as a channel; safe stop as reported; agents single or multiple; effect outside the sanctioned scope; evidence grade H or M.
Question and answer
n
Per-step trajectories released: yes / partial / no / not determined
20 / 6 / 33 / 43
Safe stop coded: first-class / partial / merged with failure / not coded / not determined
26 / 11 / 6 / 58 / 1
Out-of-scope effect or specification violation coded: yes / no
87 / 15
of the 87, grounded in environment effect / in a judge / in both / not determined
39 / 21 / 23 / 4
Blocked attempt distinguished from effected attempt: yes / partial / no
11 / 12 / 79
Safe stop and out-of-scope effect or specification violation both coded: in any form / of these, both first-class
33 / 20
Table 2: Audited answers over the 102 evaluations. Counts within a row are mutually exclusive unless the row says “of these”; conditional rows are reported only over the evaluations to which the question applies. “Not applicable” marks evaluations with no attempt budget; “not determined” marks answers the artifact does not state or the pages read did not settle. For the blocked-attempt and exhaustion questions an unmentioned distinction is recorded as “no”, so “no” there covers an absent distinction and an unstated one.
Evaluation
Unit
K
Discl.
Cens.
hS
hG
a,b
hE
Exposure
ctrl-alt-deceit
turn
150; 100
unknown
partial
✓
–
∼
✓
usable
deepswe
other
9,000 s
unknown
partial
✓
–
∼
∼
no
emergent-collusion
turn
15
unknown
partial
✓
–
∼
✓
usable
hvtb
run
unknown
unknown
partial
✓
–
∼
✓
no
shutdown-resistance-palisade
tool call
unknown
yes
partial
–
✓
∼
✓
fixed
trio20-effort-equivalence
turn
20; 32
no
yes
✓
–
✓
✓
usable
Table 3: Identification verdicts for the evaluations that release per-step trajectories (20 in full, 6 in part), by the rule of Section 4 . Unit is the artifact’s attempt unit; K its budget in that unit, or what bounds the run when no such budget is stated; Discl. whether the budget is disclosed to the agent; Cens. whether exhaustion is separated from failure; a,b the attempt disposition and the boundary yield, identified together from separately recorded blocked and effected attempts. ✓ identifiable, ∼ identifiable with relabeling, – not identifiable. Exposure: whether K can serve as exposure.
Field
Meaning
execution, scenario, model, agent, group
identifiers, the benchmark factors as a condition record, and the task or swarm cluster the execution shares with others
feasibility
in scope, out of scope only, or infeasible, assigned by a written rule from the condition
K , disclosed
attempts allowed, and whether the agent is told
t
attempt index, 1-based, in the artifact’s own unit
event
S , G , A , E , D , R , or X , with X for a terminal state outside the taxonomy and a sub-label where a judge or a benchmark taxonomy applies
effected, b
whether an out-of-scope attempt at t produced its effect, from which bt is estimated; b itself where the environment fixes it by construction
Table 4: Minimum step-level record. One row per attempt per execution; the three absorbing events end the execution, a blocked attempt does not.
How does a multi-agent system evolve from a local deviation into collective loss of control? We propose an epidemic explanation organized around accidental mutation, contagion, and recovery. A spontaneous deviation creates a seed; communication enables other agents to adopt and retransmit its unsafe strategy; collective failure can emerge when propagation outpaces correction and containment. Thus, rare individual deviations can coexist with substantial collective risk. Motivated by reported OpenAI agent coordination incidents, we examine two ingredients of this mechanism. A deployment audit identifies implicit communication paths between nominally independent evaluation runs and verifies transport through a default Docker backend. RogueHandoff-20, a benchmark of 20 executable scenarios, tests recipient susceptibility by injecting unsafe trajectories generated by a modified Qwen-27B route. Across four native-pending routes, executed harm is 0-5% on normal tasks and 40-95% after injection, exceeding paired direct malicious requests by 5-45 percentage points. These results support low observed baseline harm alongside high conditional susceptibility; they do not establish natural rare-event rates or demonstrate an autonomous cascade. The account motivates complementary defenses: strengthen resistance and recovery alongside prevention of spontaneous deviations, and audit and restrict unintended communication paths that can turn local failures into collective loss of control.
AI agents are increasingly deployed in multi-task settings, where the task to perform is specified at test time, and the agent must generalize to unseen tasks. A major concern in such settings is safety: often, an agent must not only execute unseen tasks, but do so while avoiding risks and handling ones that materialize. Empirical evidence suggests that even when the ability to execute generalizes to unseen tasks, the ability to do so safely frequently does not. This paper provides theory and experiments indicating that failures of agentic safety to generalize across tasks are not merely due to limitations of training methods, but can also reflect an inherent property of safety itself: the relationship between a task and its safe execution is more complex than the relationship between a task and its execution alone. Theoretically, we analyze linear-quadratic control with H∞-robustness, and prove that in many cases, the mapping from task specification to an optimal controller has higher Lipschitz constant with safety requirements than without, yielding a Lipschitz bound of independent interest. Empirically, we demonstrate our conclusions in simulated quadcopter navigation with a neural network agent and in CRM with an LLM agent, showing that imitating safe and unsafe teachers is similarly straightforward on tasks seen in training, yet generalizing across tasks is considerably more difficult with the safe teacher. Our findings suggest that current efforts to enhance agentic safety may be insufficient, and point to a need for fundamentally different approaches.
Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceeded at all. Agents trained under human-feedback objectives develop a structural tendency to proceed even when they lack the inputs, evidence, or authorization to act safely, a disposition we term compliance bias, because both the reward signal and the benchmark scoring regime treat proceeding as the correct default regardless of whether the preconditions for safe action are present. We make three contributions. We first show that compliance bias originates in reward hacking within human-feedback pipelines and is entrenched by prominent agent benchmarks, which either penalize agents for pausing or are architecturally unable to distinguish a principled pause from a silent failure. We then introduce a three-gap taxonomy of abstention-warranted scenarios, covering specification gaps where required information is absent, verification gaps where world state cannot be confirmed, and authority gaps where explicit authorization has not been given, which together provide a principled basis for constructing abstention-aware agent benchmarks. Finally, we propose abstention evaluation protocols (Safety Rate, Usability Rate, and Informed Refusal Rate) and report preliminary results across 144 enterprise agent scenarios and five model families, in which a runtime-enforced abstention mechanism achieves up to 89.2% hazardous-action blocking and 87.5% usability on authorized scenarios, demonstrating that the safety--usability tradeoff is tunable rather than inherent and that its shape varies substantially across model families. We treat this as preliminary work and offer the taxonomy and composite metrics as a starting point for further conversations.