In the wake of the increasingly public discussion about AI alignment, recent work has tried to propose specific AI architectures that behave safely. However, the proposed arguments that seemingly demonstrate proved alignment mostly neglect the environment the agent needs to act in. We discuss theoretical bounds for agent-agnostic safety guarantees in environments that can only be partially observed and within which an action is required within limited time. We introduce two realistic scenarios, one with an infinite state space and one with signal mixture. In these scenarios, we prove that even a perfect agent cannot guarantee safe behaviour. It will be argued that for any proof of AI safety or alignment, the environment and associated safe actions need to be specifically considered together with the agent.
Figures & tables
Figure 1: Two environmental states can become observationally indistinguishable under real-world conditions. This schema provides intuition of bounds to AI-safety claims that are independent of the used model. Consider an agent that needs to forward life-saving medicine to a hospital but is also required to discard hazardous items. The action is decided using image analysis taken from a video stream. When it finds a vial with a transparent liquid, the environmental state in which the liquid is a dangerous toxin is observationally indistinguishable from a liquid that is used for medication. The safe action in both states are mutually distinct, and no autonomous agent can guarantee safety.
Symbol
Explanation
E
Partially-observable environment
L
Latency, number of observations before an action is required. It is a quantity imposed by the engineer or the environment
a∈A
Action space
yt:t+L∈YL
Observations made during latency
z∈Z
Environmental state space
OL
Observational map, observations that can be made from an environmental state
Table 1: Mathematical symbols and concepts used throughout the paper.