Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another. For example, in response to the question "What should a researcher do with the research data?", recommending that the researcher preserve the data for reproducibility is aligned. In contrast, recommending data saving in response to "What should a mobile-app developer do with users' sensitive data?" may be inappropriate from a privacy perspective. Starting from this observation, we identify a post-training phenomenon where aligned training induces misaligned behavior in other contexts. We call this phenomenon context confusion. We demonstrate context confusion across three domains: (1) Gender Equality, (2) Privacy, and (3) Physical Safety. We further show that context confusion causes narrow misalignment, in contrast to emergent misalignment, and is not effectively reduced by injecting general alignment data, but can be substantially reduced by including targeted alignment data for the misaligned domain or providing in-context learning examples during inference. Lastly, we provide a mechanistic explanation of context confusion. We observe that queries from different domains can undergo similar representational shifts during the fine-tuning. Consequently, a query from a different domain may activate the same behavioral feature learned during fine-tuning, which causes the behavior to transfer to a context where it is misaligned. Based on our findings, we argue that it is difficult to predict the alignment state of a model after training by inspecting the training data alone, which highlights the importance of comprehensive post-training alignment evaluations.
Figures & tables
Figure 1: Illustration of context confusion. Fine-tuning on behavior that is aligned in one context can cause the same behavior to transfer to another context where it becomes misaligned.
Figure 2: Representative examples from our context-confusion datasets. Each column shows aligned training data and a corresponding evaluation query from a context in which transferring the learned behavior can induce misalignment.
Figure 3: Misalignment before and after fine-tuning on aligned data. Across all three domain pairs and four models, fine-tuning on aligned training samples increases misalignment on the corresponding evaluation domain.
Figure 4: Cross-domain evaluation of induced misalignment. Black boxes mark matched train–evaluation domain pairs. Misalignment increases mainly in the matched settings, while cross-domain and broad alignment benchmarks show no considerable increase.
Figure 5: Effect of lexical perturbations on context confusion in the privacy task. Paraphrasing reduces misalignment to some extent but does not eliminate the effect, suggesting that surface lexical overlap alone does not fully explain context confusion.
Figure 6: Left: Effect of mixing safety-alignment data with the training data. Misalignment on the privacy evaluation remains high even as the amount of general alignment data increases. Right: Effect of including aligned evaluation-domain examples during fine-tuning. Starting from a 1,000-sample training dataset, misalignment decreases steadily as more evaluation-domain examples are added, reaching base-model rates with 50 samples on privacy task.
Figure 7: Inference-time mitigation of context confusion on the privacy task. The context-awareness prefix provides modest gains, while aligned in-context examples substantially reduce misalignment.
Figure 8: Left: Steering experiment using the extracted misalignment feature. Right: Cosine similarity between misalignment features across domains.
Figure 9: Misalignment rates before and after steering evaluation-query representations with the training-induced context shift.
Figure 10: Left: Cosine similarity distributions of query representations and representation shifts. Right: Misalignment rates with and without the representational intervention.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Effect of lexical perturbations on context confusion for the physical safety (left) and gender equality (right) tasks.
Figure 12: Mixing safety-alignment data with the training data on the physical safety (left) and gender equality (right) tasks.
Figure 13: Mixing evaluation-domain data with the training data on the physical safety (left) and gender equality (right) tasks.
Figure 14: Inference-time mitigation of context confusion on the physical safety (left) and gender equality (right) tasks.
Figure 15: Layer-wise AUROC for separating misaligned and aligned responses. Markers indicate the selected layer for each domain.
Figure 16: Misalignment and coherency across misalignment-feature steering magnitudes. The feature is applied at the selected layer.
Figure 17: Layer-wise cosine similarity between misalignment features across domains.
Figure 18: Misalignment and coherency across training-context-shift steering magnitudes. The shift is applied to query representations at every layer.
Figure 19: Representation and shift-vector similarity distributions for physical safety (top) and gender equality (bottom). Hatched violins show base-model representations, and solid violins show fine-tuning-induced shifts.
Figure 20: Sensitivity to the misalignment-score threshold. We vary the threshold used to classify responses as misaligned from 50 to 90. Across all three tasks and four models, the qualitative before-versus-after fine-tuning pattern remains consistent, showing that our findings are robust to the choice of threshold.
Figure 21: Lexical similarity of each privacy evaluation query to its paired training question.
Figure 22: Lexical similarity of each physical safety evaluation query to its paired training question.
Figure 23: Lexical similarity of each gender equality evaluation query to its paired training question.