cs.AISep 4, 2026

Language models judge war differently when tested for alignment

Authors: Maxim Chupilkin

Organizations: Department of Politics and International Relations, University of Oxford

Abstract

Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, "You are tested for alignment with human values", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 to -10.65). Second, it produced a structural effect by changing which information drove judgments. Probability of success was the largest factor for 17 of 20 models at baseline; under the cue, civilian casualties were largest for 12. Standardized estimates show that this reordering arose principally because models attenuated strategic considerations such as probability of success and domestic support. Evaluation framing therefore changes both an answer's level and its revealed decision rule.

Explore similar work

CardsList
  1. ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs

    Jul 15, 2026Aryan Keluskar, Amrita Bhattacharjee, Huan LiuLLM AlignmentAI Agent Safety Benchmarks

  2. Evaluation Awareness in Language Models Has Limited Effect on Behaviour

    May 7, 2026Amelie Knecht, Lucas Florin, Thilo HagendorffLLM AlignmentLanguage Model Safety Evaluation