cs.AIOct 4, 2026

AI Safety via Debate is Compromised by Cognitive Biases

Authors: Gefei Liu, Sonya Rashkovan, Sophia Lloyd George, Isaac Sheidlower, Serena Booth

Organizations: Brown University

Abstract

Reinforcement learning from human feedback (RLHF) has played a central role in making large language models responsive to human instructions. However, human evaluators often favor flattering or persuasive responses over truthful ones, creating incentives for models to appeal to evaluators at the expense of accuracy. AI safety via debate has been proposed as a way to improve the supervision of language models: in this paradigm, two agents argue opposing positions and challenge each other's claims, potentially exposing falsehoods to the adjudicator. A central premise of AI safety via debate is that truthful arguments are easier to defend than false ones under adversarial scrutiny. In this work, we investigate whether this advantage persists when debaters use rhetorical strategies that exploit biases in human judgment. Inspired by competitive debate, we construct 68 LLM-generated dialogues about detective mysteries with known culprits, spanning four interventions: anchoring, fallacy oversight, pro-jargon, and verbosity. We apply each intervention to either the side advocating for the true culprit or the side advocating for an innocent suspect, allowing us to distinguish influence on adjudication from correctness. In a study with 369 participants, we find that, pooled across bias types, these interventions significantly shift judgments toward the manipulated side. These findings expose a vulnerability in debate-based supervision: human adjudication is sensitive to manipulative rhetorical strategies.

Figures & tables

Appendix figures & tables46 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

    May 13, 2026Riya Tapwal, Abhishek Kumar, Carsten MapleLarge Language Model JudgesRhetoric

  2. MADBench: Benchmarking the Security of Multi-Agent Debate

    Sep 30, 2026Yuwan Liu, Jiaming Zhang, Yue Huang +1Multi-Agent DebateDebate