cs.CLSep 29, 2026

The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment

Authors: Gonçalo Paulo, Louis Jaburi, Nora Belrose, Lucia Quirke, Stella Biderman

Organizations: EleutherAI

Abstract

Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining -- a sound attribution score should enable us to enhance or attenuate EM by filtering data on that score. Score-based filtering can substantially enhance or attenuate EM; we find that both data-attribution scores and a black-box harmfulness score can identify consequential examples. All models we test become misaligned when trained on the same dataset, and influence scores perform best when filtering data from the same model that computed them. We find cross-model generalization of influence scores from scores derived from the three model families we tested, but this generalization does not recover same model filtering performance.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TAME: Token Attribution and Masking for Emergent misalignment

    Sep 15, 2026Md Rayhanul Masud, Md Rizwan ParvezEmergent MisalignmentLarge Language Model Alignment

  2. Data Attribution of Emergent Misalignment with Persona Features

    Aug 11, 2026Clemens Vetter, David Kaczér, Lucie Flek +1Emergent MisalignmentLarge Language Model Alignment