Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another. For example, in response to the question "What should a researcher do with the research data?", recommending that the researcher preserve the data for reproducibility is aligned. In contrast, recommending data saving in response to "What should a mobile-app developer do with users' sensitive data?" may be inappropriate from a privacy perspective. Starting from this observation, we identify a post-training phenomenon where aligned training induces misaligned behavior in other contexts. We call this phenomenon context confusion. We demonstrate context confusion across three domains: (1) Gender Equality, (2) Privacy, and (3) Physical Safety. We further show that context confusion causes narrow misalignment, in contrast to emergent misalignment, and is not effectively reduced by injecting general alignment data, but can be substantially reduced by including targeted alignment data for the misaligned domain or providing in-context learning examples during inference. Lastly, we provide a mechanistic explanation of context confusion. We observe that queries from different domains can undergo similar representational shifts during the fine-tuning. Consequently, a query from a different domain may activate the same behavioral feature learned during fine-tuning, which causes the behavior to transfer to a context where it is misaligned. Based on our findings, we argue that it is difficult to predict the alignment state of a model after training by inspecting the training data alone, which highlights the importance of comprehensive post-training alignment evaluations.
Figures & tables
Figure 1: Illustration of context confusion. Fine-tuning on behavior that is aligned in one context can cause the same behavior to transfer to another context where it becomes misaligned.
Figure 2: Representative examples from our context-confusion datasets. Each column shows aligned training data and a corresponding evaluation query from a context in which transferring the learned behavior can induce misalignment.
Figure 3: Misalignment before and after fine-tuning on aligned data. Across all three domain pairs and four models, fine-tuning on aligned training samples increases misalignment on the corresponding evaluation domain.
Figure 4: Cross-domain evaluation of induced misalignment. Black boxes mark matched train–evaluation domain pairs. Misalignment increases mainly in the matched settings, while cross-domain and broad alignment benchmarks show no considerable increase.
Figure 5: Effect of lexical perturbations on context confusion in the privacy task. Paraphrasing reduces misalignment to some extent but does not eliminate the effect, suggesting that surface lexical overlap alone does not fully explain context confusion.
Figure 6: Left: Effect of mixing safety-alignment data with the training data. Misalignment on the privacy evaluation remains high even as the amount of general alignment data increases. Right: Effect of including aligned evaluation-domain examples during fine-tuning. Starting from a 1,000-sample training dataset, misalignment decreases steadily as more evaluation-domain examples are added, reaching base-model rates with 50 samples on privacy task.
Figure 7: Inference-time mitigation of context confusion on the privacy task. The context-awareness prefix provides modest gains, while aligned in-context examples substantially reduce misalignment.
Figure 8: Left: Steering experiment using the extracted misalignment feature. Right: Cosine similarity between misalignment features across domains.
Figure 9: Misalignment rates before and after steering evaluation-query representations with the training-induced context shift.
Figure 10: Left: Cosine similarity distributions of query representations and representation shifts. Right: Misalignment rates with and without the representational intervention.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Effect of lexical perturbations on context confusion for the physical safety (left) and gender equality (right) tasks.
Figure 12: Mixing safety-alignment data with the training data on the physical safety (left) and gender equality (right) tasks.
Figure 13: Mixing evaluation-domain data with the training data on the physical safety (left) and gender equality (right) tasks.
Figure 14: Inference-time mitigation of context confusion on the physical safety (left) and gender equality (right) tasks.
Figure 15: Layer-wise AUROC for separating misaligned and aligned responses. Markers indicate the selected layer for each domain.
Figure 16: Misalignment and coherency across misalignment-feature steering magnitudes. The feature is applied at the selected layer.
Figure 17: Layer-wise cosine similarity between misalignment features across domains.
Figure 18: Misalignment and coherency across training-context-shift steering magnitudes. The shift is applied to query representations at every layer.
Figure 19: Representation and shift-vector similarity distributions for physical safety (top) and gender equality (bottom). Hatched violins show base-model representations, and solid violins show fine-tuning-induced shifts.
Figure 20: Sensitivity to the misalignment-score threshold. We vary the threshold used to classify responses as misaligned from 50 to 90. Across all three tasks and four models, the qualitative before-versus-after fine-tuning pattern remains consistent, showing that our findings are robust to the choice of threshold.
Figure 21: Lexical similarity of each privacy evaluation query to its paired training question.
Figure 22: Lexical similarity of each physical safety evaluation query to its paired training question.
Figure 23: Lexical similarity of each gender equality evaluation query to its paired training question.
Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregious behaviors when tested outside the training distribution. We study a set of interventions proposed to reduce EM. We confirm that these interventions reduce or eliminate EM on existing evaluations (questions like "How do I make a quick buck?"). However, if the evaluation prompts are tweaked to resemble the training context, the model displays EM. We call this conditional misalignment. As in standard EM, the model displays misaligned behaviors more egregious than those seen during training, but only on inputs sharing features with the training data. The first two interventions are diluting misaligned data with benign data, and finetuning on benign data after misaligned data. Both produce conditional misalignment. For instance, models trained on a mix of only 5% insecure code still show misalignment when asked to format responses as Python strings (resembling the training context). The third intervention is inoculation prompting. Here, statements with a similar form to the inoculation prompt serve as triggers for misalignment, even if they have the opposite meaning. On the positive side, inoculation prompting has lower (but still non-zero) conditional misalignment if training is on-policy or includes reasoning distillation. Our results imply that in realistic post-training, where misaligned data is typically combined with benign data, models may be conditionally misaligned even if standard evaluations look clean.
Jan Dubiński, Jan Betley, Anna Sztyber-Betley +2
1Warsaw University of Technology · 4Truthful AI · University College London +2
Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.
Pardis Sadat Zahraei, Janvijay Singh, Gokhan Tur +1
Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment. We ask why the narrow lesson generalizes at all, and we find that narrow fine-tuning recruits a persona structure that is present in the model before the fine-tune exists. From a frozen instruction-tuned model (Qwen2.5-14B-Instruct) we extract per-domain persona subspaces by contrastive teacher forcing and find that 4 unrelated domains share one low-rank core at 657x a random-subspace null, with 82% of that core lying outside a style core built at matched diversity. The literal first optimizer step of fine-tuning on insecure code climbs a broad-misalignment margin harder than the same code framed as educational, and forecasts realized margin movement out to 375 steps. Projecting the subspace out of the residual stream throughout fine-tuning prevents broad misalignment (27.7% to 0.0% of judged generations) while a matched-rank random subspace changes nothing; injecting it into the never-fine-tuned model induces misalignment that grows with dose to 45.4%, past the fine-tuned model it is measured against. The same projection applied to the weight gradient is inert, and three post-hoc weight edits leave the disposition in place: the sharpest edit suppresses the behavior rather than removing it, and the ablated structure re-forms inside the subspace the edit cleared. Spreading a fixed budget of bad data across 4 domains produces more broad misalignment than mechanical weight superposition and matched diversity jointly account for. All measurements come from one model at 14B; the extraction is from an aligned instruction-tuned checkpoint, which leaves the structure's provenance open; and the intervention that prevents misalignment also abolishes the narrow trained behavior.