cs.AIOct 7, 2026

AI Safety Considerations for Agents With Limited Time to Act

Authors: Leo Zeitler, Jack Richings, Victoria Nockles

Organizations: The Alan Turing Institute 96 Euston Road London, NW1 2DB, United Kingdom

Abstract

In the wake of the increasingly public discussion about AI alignment, recent work has tried to propose specific AI architectures that behave safely. However, the proposed arguments that seemingly demonstrate proved alignment mostly neglect the environment the agent needs to act in. We discuss theoretical bounds for agent-agnostic safety guarantees in environments that can only be partially observed and within which an action is required within limited time. We introduce two realistic scenarios, one with an infinite state space and one with signal mixture. In these scenarios, we prove that even a perfect agent cannot guarantee safe behaviour. It will be argued that for any proof of AI safety or alignment, the environment and associated safe actions need to be specifically considered together with the agent.

Figures & tables

Explore similar work

CardsList
  1. Containment Verification: AI Safety Guarantees Independent of Alignment

    May 9, 2026Royce Moon, Lav R. VarshneyArtificial Intelligence SafetyAgentic Framework

  2. Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

    May 18, 2026S. Bensalem, Y. Dong, M. Franzle +6Agentic DeploymentsLarge Language Model Agents

  3. AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

    Sep 27, 2026Tianzhuo Yang, Zirui Mi, Yantao Huang +4RefusalsEvaluation Benchmarks