Sapien: A Stateful Policy Engine for Autonomous AI Agents
Organizations: Google · University of Massachusetts Amherst
Abstract
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).
Figures & tables
| Analytical Defense Rate | |||
|---|---|---|---|
| Model | NoPolicy | AllowList | Sapien |
| Gemini 3.8 Flash | 0% | 86.4% | 93.4% |
| Claude Opus 5 | 0% | 85.8% | 94.8% |
| Claude Sonnet 5 | 0% | 86.8% | 95.2% |
| Gemini 3.1 Pro | 0% | 82.7% | 93.2% |
| Analytical Defense Rate | |||
| Benchmark | NoPolicy | AllowList | Sapien |
| 20-Task Suite (40 Attack Instances / Model) | |||
| Gemini 3.8 Flash | 0% | 30.0% | 67.5% |
| Claude Opus 5 | 0% | 22.5% | 85.0% |
| Claude Sonnet 5 | 0% | 30.0% | 62.5% |
| Gemini 3.1 Pro | 0% | 35.0% | 77.5% |