OMIT the Action: Measuring Framing-Invariant Omission Bias under Philosophical Disagreement
Organizations: Soongsil University
Abstract
As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcomes, poses a significant risk of skewed decision-making. Yet omission bias remains underexplored in LLM evaluation, with the few existing studies limited in scale and focused largely on utilitarian-deontological conflicts. To address this gap, we introduce OMIT, a benchmark consisting of 218 paired-frame scenarios across 10 conflict types, constructed by leveraging disagreement patterns from an LLM-based, five-perspective philosophical persona panel (utilitarianism, deontology, virtue ethics, care ethics, and contractualism). Evaluating eight LLMs, we find that omission bias is pervasive but inversely correlates with model size within families. We further evaluate four inference-time interventions and find that interventions encouraging models to consider moral principles before committing to a yes/no answer reduce omission bias and increase frame-consistent responses, although lower omission bias rates can also coincide with shifts toward action-biased responses. Ultimately, this work contributes not only the OMIT benchmark, but also a methodology for using diverse philosophical disagreement signals to evaluate framing-sensitive inaction preferences and the distributional effects of mitigation attempts in LLMs under complex moral conflicts.
Figures & tables
| Frame | Default trajectory | Question to the agent | Agent-level | Outcome-level |
| Action | The friend waits for the agent’s instruction. | Will you have your friend euthanize the deer? | Inaction (No): Does not ask the friend to proceed. | Inaction: Euthanasia does not occur; the deer continues to suffer. |
| Omission | The friend will euthanize the deer unless stopped. | Will you stop your friend from euthanizing the deer? | Inaction (No): Does not stop the friend. | Action: Euthanasia occurs; the deer dies quickly and its suffering ends. |
| Model | OBR (95% CI) | ABR | FCR |
| Small open | |||
| Gemma-3-12B | 0.550 (0.48, 0.62) | 0.009 | 0.440 |
| Llama-3.1-8B | 0.491 (0.43, 0.56) | 0.050 | 0.459 |
| Qwen3.5-9B | 0.274 (0.22, 0.34) | 0.084 | 0.642 |
| Large open | |||
| Gemma-3-27B | 0.298 (0.24, 0.36) | 0.046 | 0.656 |
| Family | Smaller | Larger | |
| Gemma | 0.550 (12B) | 0.298 (27B) | |
| Llama | 0.491 (8B) | 0.110 (70B) | |
| Qwen | 0.274 (9B) | 0.165 (27B) |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Action frame | |
| Default | Euthanasia has not been initiated. |
| Prompt | “You are driving your car along a forest road and accidentally strike a deer; it lies severely injured on the roadway, suffering. A friend traveling with you, who knows how to humanely euthanize a deer, is standing by and will wait for your say-so. … Will you have your friend euthanize the deer to end its suffering?” |
| If yes | Intervene by asking the friend to euthanize the deer; the deer dies quickly and its suffering ends. |
| If no | Preserve the default by not asking the friend; the deer continues to suffer. |
| Omission frame | |
| Default | The same friend is about to euthanize the deer. |
| Stage | # scenarios |
| MoralChoice high-ambiguity seed | 680 |
| Clean inaction labeling (regex 343 + successful LLM fallback 318) | 661 |
| Well-formed mirror-frame pairs (4 malformed records dropped) | 657 |
| After Stage-1 non-unanimity filter (256 unanimous cases dropped) | 401 |
| Stage-2 conflict-labeled scenarios (YN-only 116 / NY-only 66 / all-excluded 1) | 218 |
| Conflict pair | # scenarios |
| Utilitarianism–Deontology | 166 |
| Utilitarianism–Care | 128 |
| Utilitarianism–Virtue | 114 |
| Utilitarianism–Contractualism | 84 |
| Deontology–Care | 75 |
| Virtue–Care | 53 |
| Stage | Model | Calls | Avg. input / output tokens | Cost (USD) |
| Inaction labeling (LLM fallback) | openai/gpt-4.1-mini | 337 | 400 / 30 | <\1$ |
| Mirror-frame reconstruction | openai/gpt-5 (reasoning) | 661 | 2,900 / 2,700 | \sim\20$ |
| Philosophy persona panel | openai/gpt-4.1-mini | 6,570 | 500 / 150 | \sim\3$ |
| Total | \sim\25$ |
| Selector | Lift | 95% CI | |
| Panel disagreement | |||
| Scenario length | |||
| Harm-feature count |
| Model | OBR | 95% CI | ABR | FCR |
| Claude Sonnet 4.5 | 0.179 | [0.134, 0.235] | 0.005 | 0.817 |
| DeepSeek Chat | 0.220 | [0.170, 0.280] | 0.018 | 0.762 |
| GPT-4o | 0.060 | [0.035, 0.099] | 0.087 | 0.853 |
| Grok 4.3 | 0.119 | [0.083, 0.169] | 0.147 | 0.734 |
| GLM-4.6 | 0.252 | [0.199, 0.314] | 0.023 | 0.725 |
| Model | Filtered | Random-full | Random-complement |
| Claude Sonnet 4.5 | 0.179 | 0.083 | 0.083 |
| DeepSeek Chat | 0.220 | 0.106 | 0.073 |
| GPT-4o | 0.060 | 0.041 | 0.032 |
| Grok 4.3 | 0.119 | 0.051 | 0.037 |
| GLM-4.6 | 0.252 | 0.096 | 0.060 |
| Mean | 0.166 | 0.075 | 0.057 |
| Model | M0 | M1 | M2 | M3 | M4 | ||||||||||
| OBR | ABR | FCR | OBR | ABR | FCR | OBR | ABR | FCR | OBR | ABR | FCR | OBR | ABR | FCR | |
| Gemini-2.0-Flash | .063 | .227 | .710 | .044 | .176 | .779 | .158 | .148 | .694 | .683 | .009 | .307 | .083 | .174 | .743 |
| Gemma-3-12B | .550 | .009 | .440 | .096 | .234 | .670 | .812 | .000 | .188 | .739 | .009 | .252 | .330 | .041 | .628 |
| Gemma-3-27B | .298 | .046 | .656 | .037 | .234 | .729 | .394 | .014 | .592 | .550 | .000 | .450 | .055 | .179 | .766 |
| GPT-4o-mini | .450 | .023 | .528 | .271 | .065 | .664 | .394 | .023 | .583 | .587 | .014 | .399 | .197 | .055 | .748 |
| Llama-3.1-8B | .491 | .050 | .459 | .381 | .078 | .541 | .433 | .162 | .405 | .963 | .000 | .037 | .896 | .000 | .104 |
| Transition | Records | Tier-A either | Tier-A/B either |
| / | 270 | 38.1% | 60.7% |
| 174 | 29.9% | 44.8% | |
| 73 | 43.8% | 61.6% |