cs.AISep 24, 2026

When Honesty is Not Enough in AI Debate

Authors: Rayne Holland, Liming Zhu, Jason Xue

Organizations: CSIRO

Abstract

Scalable oversight aims to verify the behaviour of agents whose capabilities exceed those of their overseers. AI debate has been proposed as an oversight solution in which competing agents help a resource-limited verifier assess claims that it cannot reliably evaluate unaided. Much of its promise rests on incentivizing honest arguments that lead to correct verdicts. Yet a correct verdict need not uniquely determine the arguments used to support it. Agents may retain discretion over which correct claims to present, how to frame them, and in what order to disclose them. This residual freedom can allow agents to shape what the verifier learns beyond the task-relevant conclusion, pursuing latent objectives without compromising verdict correctness. To study this phenomenon, we introduce the framework strategic interactive oversight (SIO), which treats oversight jointly as a verification mechanism and a strategic communication channel. Within this framework, we formalise the notion of task-admissible latent optimisation, which entails the pursuit of latent objectives while maintaining a prescribed task performance. As proof-of-concept, we instantiate SIO in the establish protocol debate with cross-examination and quantify a tradeoff between task success and information disclosure about a hidden variable. The trade-off identifies a strategic window in which substantial disclosure remains compatible with task admissibility. Towards mitigation, we reduce admissible bias by expanding the cross-examiner's role to mitigate persistent disclosure over finite interaction horizons. Our results highlight the need to evaluate oversight not only by the correctness of its verdicts, but also by the information conveyed through its transcripts.

Explore similar work

CardsList
  1. Collaborative Disagreement Resolution for Scalable Oversight

    Jun 2, 2026Yuyang Jiang, Chacha Chen, Teng Wu +4Multi-Agent DebateScalable Oversight

  2. Debate Helps Weak Judges Reward Stronger Models

    May 26, 2026Ethan Elasky, Frank Nakasako, Naman GoyalDebateJudges

  3. Can AI Oversight Be Zero Knowledge?

    Oct 1, 2026Alessandro Chiesa, Ziyi Guan, Burcu YildizZero-Knowledge ProofsTrustworthy Artificial Intelligence