cs.AIOct 8, 2026

An Investigation of Model Coherence: Narrow Finetunes Contradict Themselves Under Resampling

Authors: Robert Graham, Yariv Barsheshat, Phil Blandfort, Sabri Alouache

Organizations: Independent · Predictably Weird · Amazon

Abstract

A large body of research measures model coherence based on output variance without adequately considering competing causes. We identify two such causes, ambiguity and indifference, and we introduce a set of 175 questions where contradicting answers cannot easily be explained by either. We then measure incoherence in terms of contradictions when resampling answers to the same question. In contrast to other methods our metric has high specificity, and only ranks models as incoherent when the issues are glaring. Even so, we find narrow finetunes score poorly. Inspecting inconsistencies flagged by our method, we find that model organisms from the literature display severe issues such as identity conflation, introspection failures and rationalizations. These findings suggest that the pathologies induced by narrow finetuning may limit what these models can tell us about coherent misaligned behaviour.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Iterative Finetuning is Mostly Idempotent

    May 1, 2026Zephaniah Roe, Jack Sanderson, Dang Nguyen +5Supervised Fine-TuningLLM Post-Training

  2. What Shapes Emergent Misalignment? Insights from Training Dynamics, Model Priors, and Data

    Jun 18, 2026Yuchen Zhang, Anietta Weckauff, Diego Garcia-Olano +1LLM EvaluationLLM Alignment