cs.CRMar 6, 2026

Proof-of-Guardrail in AI Agents and What (Not) to Trust from It

Authors: Xisen JinMichael DuanQin LinAaron ChanZhenglun ChenJunyi DuXiang Ren

Organizations: Sahara AI · University of Southern California

Abstract

As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which introduces a threat where safety measures are falsely advertised. To address the threat, we propose proof-of-guardrail, a system that enables developers to provide cryptographic proof that a response is generated after a specific open-source guardrail. To generate proof, the developer runs the agent and guardrail inside a Trusted Execution Environment (TEE), which produces a TEE-signed attestation of guardrail code execution verifiable by any user offline. We implement proof-of-guardrail for OpenClaw agents and evaluate latency overhead and deployment cost. Proof-of-guardrail ensures integrity of guardrail execution while keeping the developer's agent private, but we also highlight a risk of deception about safety, for example, when malicious developers actively jailbreak the guardrail. Code and demo video: https://github.com/SaharaLabsAI/Verifiable-ClawGuard

Explore similar work

CardsList
  1. Provably Secure Agent Guardrail

    May 28, 2026Benlong Wu, Weiming Zhang, Kejiang Chen +2Streaming GuardrailsArtificial Intelligence Safety