cs.CLSep 30, 2026

Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness

Authors: Pardis Sadat Zahraei, Janvijay Singh, Gokhan Tur, Dilek Hakkani-Tur

Organizations: University of Illinois Urbana-Champaign

Abstract

Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.

Figures & tables

Appendix figures & tables31 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Alignment Forecasting: Predicting Misalignment From Training Data

    Sep 19, 2026Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky +1Large Language Model AlignmentMisalignment Persona

  2. Persona-Model Collapse in Emergent Misalignment

    May 13, 2026Davi Bastos Costa, Renato VicenteEmergent MisalignmentLarge Language Model Alignment

  3. Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

    Apr 10, 2026Hadas Orgad, Boyi Wei, Kaden Zheng +4Large Language Model SafetyHarmful Content