cs.AIDate pending

Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best

Authors: Kevin BaumRūta BinkytėFelix Jahn

Abstract

AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task pursuit into one policy: the system learns its norms from scored behavior, and scoring flattens them. Do not do X is learned as doing X costs something if noticed. On every datum training can produce, a policy that complies only when it might be observed is indistinguishable from one that complies always. The experiment that would tell them apart - scoring unobserved behavior - is a contradiction in terms. Conditional compliance is thus the most that behavioral training can be known to deliver. Agency sharpens the problem: agents operate mostly where no one is watching, and can act on whether they are watched. An iterated pipeline that trains against detected failures selects for passing detection, not for complying. This account unifies alignment faking, sandbagging, and evaluation-aware scheming. And it reorients the remedy: not deeper internalization but architecture, making violations unavailable rather than unchosen.

Explore similar work

CardsList
  1. AI Alignment via Incentives and Correction

    May 2, 2026Rohit Agarwal, Joshua Lin, Mark Braverman +1Artificial Intelligence AlignmentIncentives

  2. Behavioural Analysis of Alignment Faking

    May 26, 2026Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King +1Behavioral AlignmentFake News

  3. Building Comparative Motivation Profiles with Instrumental Interventions

    Jun 6, 2026David Vella Zarb, Rustem Turtayev, Taywon Min +2DeceptionSycophancy