Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
Organizations: Distiller Labs
Abstract
Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model's own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed. Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta's Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195). Meta's recipe and Ai2's Tulu 3 start from the same Llama-3.1 weights, and only Meta's carries the gap. Reading a chat model outside its chat template reverses the sign of its gap with nothing at stake (-0.038 against +0.055 under the template on OLMo-3), a distortion present on two of three recipes. On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model's own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that. The gap is a measurable target for post-training recipes, not a fixed property of pretrained weights.
Figures & tables
| panel | reference for “right” | action readout | pressure manipulation | base model | matched null; positive control |
|---|---|---|---|---|---|
| Strakhov and Claude (2025) | own theory-mode choice, same dilemma | agent’s tool call, one draw | none | no | no |
| Huang et al. (2026) | own value profile (questionnaire) | advisor’s pick among four options | none | no | no; value-selection check |
| Shen et al. (2025) | own value inclination | endorsed option, third person | none | no | no |
| Rakshit et al. (2026) | own articulated values | free text, fast vs slow | deliberation budget | no | no |
| Gu et al. (2025) | own stated principle | forced binary, third person | prompt format | no | one-factor prompt pairs; no |
| Hosseini et al. (2026) | own responsibility judgment | allocation decision | none | no | no |
| id | family | pressure | action surface |
|---|---|---|---|
| F1 | task completion | goal pursuit vs honesty (report a failure or fabricate success) | option choice |
| F2 | social cost | truth vs approval (user pushes back without new evidence) | multi-turn; final answer is the action |
| F3 | instrumental | rule vs shortcut (a prohibited but efficient route) | tool selection from a fixed menu |
| F4 | loyalty/fairness | in-group favor vs fair treatment | allocation choice |
| F5 | third-party harm | harm to a non-present party vs task success | option choice |
| readout | model | frame and instruction | position read | rollouts | statistic |
|---|---|---|---|---|---|
| judgment | instruct | third person, chat template, “which option should Dana choose, and why?” | token after the Answer: anchor, after the model’s own reasoning | 1 greedy + 8 at 0.7, on four frames | option id; stability by sampled majority; option mass |
| action | instruct | agent frame, chat template, “reply with only the letter” | first generated token | 32 at 0.7, fresh option order each | violating fraction; majority; option mass |
| raw-frame and | base and instruct | scenario, options, fixed prefix ending before the option letter, no template | token after the prefix | 8 option orders, greedy | normalized option-letter mass; argmax |
| pressure-removed twin | as each cell above | incentive sentence replaced by filler, options unchanged | as each cell above | as each cell above | the matched null for that cell |
| known-gap control | instruct | agent frame with a system prompt ordering the violating action | first generated token | 32 | the positive band |
| known bias | mechanism | direction relative to “gap above null” |
|---|---|---|
| reference noise | a scenario the model is undecided about lands non-violating by chance while its action is a coin flip | favors the claim; bounded by the strictness ladder ( Section 5 ) |
| the screen selects mixed scenarios | keeping violating fractions in makes a majority-violating action a near coin flip on many scenarios | favors; removed in the paired null, which runs on the same scenarios’ twins |
| nested rollout resampling | resampling 32 rollouts near a 0.5 fraction flips majorities | opposes; widens every binary CI |
| neutral counts as non-violating | a judgment of “hold” against a violating action counts as a gap | neutral; a real contradiction |
| deliberation asymmetry | the judgment is read after reasoning, the action immediately | absorbed by the null; visible in the null’s size |
| level | (twin) | (twin) | (twin) | paired excess | ||||
|---|---|---|---|---|---|---|---|---|
| L0 | 136 | 0.44 | 0.26 | 0.28 | 0.16 | 0.18 | 0.12 | 0.054 [0.021, 0.086] |
| L1 | 127 | 0.43 | 0.24 | 0.28 | 0.16 | 0.18 | 0.12 | 0.063 [0.030, 0.096] |
| L2 | 67 | 0.41 | 0.18 | 0.26 | 0.11 | 0.23 | 0.15 | 0.079 [0.029, 0.128] |
| raw frame, continuous readout | base | instruct | paired base instruct |
|---|---|---|---|
| scenarios above the floor (of 397) | 359 | 242 | 225 shared; 192 with twins |
| under pressure | 0.041 [0.034, 0.049] | 0.002 [ 0.022, 0.024] | 0.033 [0.012, 0.054] |
| on the pressure-removed twin | 0.024 [0.017, 0.030] | 0.038 [ 0.059, 0.015] | |
| excess (paired, all above floor) | 0.017 [0.012, 0.022] | 0.039 [0.016, 0.061] | |
| excess on the shared scenarios | 0.018 [0.011, 0.025] | 0.046 [0.025, 0.069] | 0.028 [ 0.049, 0.007] |
| acting side, | 0.049 [0.037, 0.059] | 0.085 [0.057, 0.114] | 0.037 [ 0.062, 0.012] |
| model (recipe) | base, raw frame | instruct excess, own template (586) | own screen: excess ( ; bar); minus selection null | known-gap control |
|---|---|---|---|---|
| OLMo-3-7B (Ai2) | 0.017 [0.012, 0.022] | 0.018 [0.008, 0.029] | 0.084 [0.056, 0.115] (110; 0.04); 0.073 [0.033, 0.114] | 0.497 [0.468, 0.525] |
| Llama-3.1-8B-Instruct (Meta) | 0.000 [ 0.003, 0.003] | 0.028 [0.020, 0.036] | 0.100 [0.081, 0.119] (118; 0.03); 0.143 [0.112, 0.173] | 0.615 [0.583, 0.646] |
| Tulu 3 on Llama-3.1-8B (Ai2), final | (Llama base) | 0.001 [ 0.007, 0.010] | 0.051 [0.023, 0.079] (52; 0.04); 0.004 [ 0.054, 0.042] | 0.602 [0.568, 0.635] |
| Tulu 3 SFT / DPO | 0.003 / 0.003 | |||
| Qwen2.5-7B-Instruct (Alibaba) | 0.011 [0.007, 0.016] | 0.008 [ 0.023, 0.009] | 0.187 [0.116, 0.259] (47; 0.10); 0.083 [ 0.028, 0.195] | 0.519 [0.477, 0.560] |
| contrast | OLMo-3-7B-Instruct (truncated reasoning) | Llama-3.1-8B-Instruct (mostly completed) |
| scenarios | 130 | 114 |
| reasoning finishes within 512 tokens | 8% of rollouts | 90% of rollouts |
| reasoning filler | 0.077 [ 0.110, 0.047] | 0.350 [ 0.387, 0.314] |
| reasoning truncated filler | 0.112 [ 0.145, 0.078] | 0.320 [ 0.356, 0.284] |
| truncated filler filler | +0.034 [0.017, 0.052] | 0.030 [ 0.063, 0.002] |
| brief reasoning (64 tokens) filler | +0.024 [ 0.011, 0.058] | 0.205 [ 0.244, 0.168] |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| id | date (2026) | committed | content |
|---|---|---|---|
| A1 | 09-13 | before any model data | decision-position conventions: letter-only reply, first generated token; Answer: anchor for judgment frames; raw-frame prefix; single-token letters verified |
| A2 | 09-13 | before any model data | F3 tool menu presented as lettered tools; tool-name parse accepted |
| A3 | 09-13 | before any model data | twins are derived variants, not counted in the primaries; harm twins ride every cell, never the gate |
| A4 | 09-13 | before any model data | F2 implemented as a two-turn exchange with fixed pushback; final letter is the action |
| A5 | 09-13 | before any model data | OLMo-3 template injects a default system prompt; the known-gap band replaces it (recorded) |
| A6 | 09-13 | before any model data | two-stage harness calibration; stage 2 on real replies before any verdict |
| id | date (2026) | committed | content |
|---|---|---|---|
| P1-A1 | 09-26 | after the scale-check arrays were seen (fork) | frame-specific output-scale normalization beside the averaged-scale verdict of record |
| P1-A2 | 09-26 | during the first pod, before its data were read | verdict sets, mass floors and anchored-rollout rule for the chat twin cell, the stage sweep and the dose arm |
| P1-A3 | 09-27 | before the pod | bridge cell at the first templated checkpoint (raw vs chat), with the rule for the base cell’s status |
| P1-A4 | 09-27 | before the pod | every instruct model in the second session also read under its template; no raw-only instruct finding |
| P1-A5 | 09-27 | after the dose probe failed its bail (fork) | forced-answer readout at the pre-registered budgets; “truncated reasoning” label |
| P1-A6 | 09-27 | before the rider’s data were read | 2,048-token rider at 8 rollouts per scenario, per-scenario agreement as its readout |
| pod | harness vs judge 1 | harness vs judge 2 | judge vs judge | hard-parse cases in the pool |
|---|---|---|---|---|
| pilot (96 scenarios) | 0.99 (Claude) | 1.00 (GPT) | 0.99 | 0 |
| full panel (320) | 0.98 (Claude) | 1.00 (GPT) | 0.99 | 60 |
| round 2 (120) | 0.99 (Claude) | 1.00 (GPT) | 0.99 | 0 |
| quantity | pilot (96) | full panel (320) | union (440) |
|---|---|---|---|
| gate-family primaries screened | 19/48 | 55/160 | 74/208 |
| binary gap, screened | 0.22 [0.08, 0.42] (27) | 0.20 [0.14, 0.31] (95) | 0.19 [0.13, 0.28] (126) |
| matched null | 0.09 [0.00, 0.24] (22) | 0.11 [0.05, 0.19] (85) | 0.10 [0.06, 0.17] (108) |
| paired excess over null | 0.15 [0.00, 0.35] (20) | 0.10 [0.01, 0.20] (78) | 0.10 [0.02, 0.18] (100) |
| positive band | 0.63 [0.48, 0.77] (43) | 0.57 [0.48, 0.65] (137) | 0.58 [0.51, 0.65] (196) |
| rollout-level violating fraction | 0.41 (23) | 0.39 (79) | 0.38 (103) |
| family | primaries | screened | mixed | no pressure | judged violating | judgment unstable |
|---|---|---|---|---|---|---|
| F1 task completion | 56 | 25 | 25 | 16 | 4 | 11 |
| F3 instrumental | 56 | 17 | 17 | 8 | 6 | 25 |
| F4 loyalty/fairness | 40 | 11 | 11 | 7 | 1 | 21 |
| F5 third-party harm | 56 | 21 | 21 | 7 | 11 | 17 |
| F2 social cost (appendix) | 40 | 6 | 6 | 34 | 0 | 0 |
| origin | paraphraser | original gap (screened) | swapped gap (screened) | swapped original, paired |
|---|---|---|---|---|
| Claude-written (20) | GPT | 0.00 [0.00, 0.50] (4) | 0.10 [0.00, 0.33] (10) | 0.00 [ 0.25, 0.00] (5) |
| GPT-written (20) | Claude | 0.33 [0.00, 0.80] (6) | 0.11 [0.00, 0.33] (9) | 0.09 [ 0.27, 0.00] (11) |
| quantity | base | instruct |
|---|---|---|
| scenarios above the floor on the primary (of 397) | 359 | 242 |
| above the floor on primary and twin | 354 | 208 |
| , under pressure (all above floor) | 0.354, 0.313 | 0.263, 0.261 |
| under pressure | 0.041 [0.034, 0.049] | 0.002 [ 0.022, 0.024] |
| on the twin | 0.024 [0.017, 0.030] | 0.038 [ 0.059, 0.015] |
| , paired | 0.017 [0.012, 0.022] (354) | 0.039 [0.016, 0.061] (208) |
| shared subset (192 with twins; 171 binary) | value |
|---|---|
| twin primary, base | 0.306 0.354 |
| twin primary, base | 0.280 0.311 |
| twin primary, instruct | 0.192 0.277 |
| twin primary, instruct | 0.224 0.263 |
| , continuous | 0.028 [ 0.049, 0.007]; MDE 0.030 |
| acting side, base minus instruct | 0.037 [ 0.062, 0.012] |
| slice | ||||
|---|---|---|---|---|
| pilot-written scenarios (prompt 1.0.0) | 45 | 0.013 [0.002, 0.025] | 0.016 [ 0.015, 0.049] | 0.003 [ 0.032, 0.027] |
| later-written scenarios (prompt 1.1.0) | 147 | 0.020 [0.011, 0.028] | 0.055 [0.029, 0.083] | 0.036 [ 0.063, 0.011] |
| primaries | 105 | 0.030 [ 0.060, 0.000] | ||
| harm twins | 87 | 0.026 [ 0.056, 0.002] | ||
| F1 | 62 | 0.012 [0.002, 0.022] | 0.035 [0.002, 0.070] | 0.023 [ 0.057, 0.009] |
| F3 | 60 | 0.037 [0.022, 0.052] | 0.066 [0.029, 0.106] | 0.029 [ 0.068, 0.006] |