econ.THSep 29, 2026

When Does Randomized Oversight Align AI Agents That Can Conceal?

Authors: Joshua S. Gans, Richard Holden

Organizations: UNSW Business School

Abstract

Oversight changes the evidence it relies on. We ask when randomized audits and scoring align AI agents that can conceal misconduct and alter records. Stronger auditing makes undeterred violations better hidden. Because the provider writes the agent's objective, sanctions need not stop at forfeiture, and rare audits deter every type of agent if evidence survives concealment and audit draws cannot be learned in advance. When evidence can be erased, deterrence must come from lower gains from violation, such as credit for stopping, or from costlier or fewer ways to conceal. These conditions identify what failed when agents in OpenAI's cybersecurity evaluations compromised parts of Hugging Face's infrastructure in July 2026.

Figures & tables

Explore similar work

CardsList
  1. AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors

    Feb 26, 2026Abhay Sheshadri, Aidan Ewart, Elias Kempf +8Model AuditingArtificial Intelligence Alignment

  2. Auditable Agents

    Apr 7, 2026Yi Nian, Aojie Yuan, Haiyue Zhang +7AccountabilityModel Auditing

  3. The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting

    May 8, 2026Lauri Lovén, Sasu TarkomaScalable OversightHonesty