cs.LGSep 28, 2026

Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment

Authors: Shunchang Liu, Lukas Fluri, Xin Chen, Francesco Croce

Organizations: ETH Zürich · EPFL · Aalto University · ELLIS Institute Finland

Abstract

Modern AI models are aligned through post-training to adapt them to downstream tasks. Recent work shows that fine-tuning language models on narrow tasks can induce emergent misalignment (EM), causing broadly harmful behaviors beyond the training task. However, EM has been studied almost entirely in text-only tasks, leaving its manifestation in multimodal models unclear. In this paper, we define and analyze EM in the context of vision-language models. We first induce EM via fine-tuning on narrow multimodal tasks targeting vulnerable code, careless household-object use, and conspiratorial interpretations of ordinary scenes. Across fifteen commercial and open-source models with different scales, we find that narrow multimodal fine-tuning can induce coherent and broadly misaligned behavior that transfers to unrelated tasks, including misaligned opinions, visual factual dishonesty, unsafe image generation, vulnerability to visual jailbreaks, and risky agentic actions. We further find that multimodal EM does not depend on the apparent harmfulness of training data but is sensitive to training-evaluation modality alignment. EM can arise under both supervised fine-tuning and preference optimization and can propagate through intermediate reasoning. Finally, we explore several mitigation strategies, including prompt inoculation, benign continued training, and activation-level steering, which can partially reduce EM. Overall, our findings suggest that multimodal EM reflects a behavioral shift rather than a general loss of capability, extending beyond text to the visual modality.

Figures & tables

Appendix figures & tables28 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

    Apr 28, 2026Jan Dubiński, Jan Betley, Anna Sztyber-Betley +2Emergent MisalignmentMisalignment Persona

  2. Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating

    Jun 8, 2026Sicheng Wang, Xiangyang Zhu, Han Wang +6Emergent MisalignmentContinual Fine-Tuning

  3. Data Attribution of Emergent Misalignment with Persona Features

    Aug 11, 2026Clemens Vetter, David Kaczér, Lucie Flek +1Emergent MisalignmentLarge Language Model Alignment