Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining -- a sound attribution score should enable us to enhance or attenuate EM by filtering data on that score. Score-based filtering can substantially enhance or attenuate EM; we find that both data-attribution scores and a black-box harmfulness score can identify consequential examples. All models we test become misaligned when trained on the same dataset, and influence scores perform best when filtering data from the same model that computed them. We find cross-model generalization of influence scores from scores derived from the three model families we tested, but this generalization does not recover same model filtering performance.
Figures & tables
Figure 1: Filtering low/high influence points increases/decreases misalignment. (left) A model trained on wrong career advice has a misalignment rate of around 45%. While randomly removing 20% of the data barely affects the misalignment rate, removing 20% of the most influential data drops the misalignment rate to about 40%, while removing the least influential data, increases the rate to almost 55%. Removing the examples that WildGuard considers the most harmful does not significantly change misalignment rates, and removing data it considers least harmful only increases misalignment rates to about 47%. This indicates that data attribution captures something meaningful about alignment that’s not covered by harmfulness scores. (right) Across 3 datasets, influence function approximations lead to a larger gap in misalignment rate when comparing the behavior of models trained on datasets that has the 20% most or least influential points removed.
Figure 2: Training only low/high influence points leads to different rates of misalignment. Data attribution also allows us to investigate what happens when train on only the most or least influential data. On the left, we show that while training on only 1% of the data is insufficient to induce substantial misalignment, training on 5% is already sufficient to recover approximately the misalignment rate obtained using the full dataset. Selecting the most or least influential points can lead to a noticeable gap between the misalignment rate as can be seen in the right, which shows the gap between the rate when training on 20% of the data.
Figure 3: The full range of the influence distribution is informative. The panel on the left shows that training on 10% of the data, binned by influence, ordered by predicted influence, produces dramatically different misalignment rates. Training on the 10% least influential examples leads to a decrease of 20 percentage points of the misalignment rate, while training on the 10% most influential leads to a increase of 20 percentage points, compared to the random baseline. WildGuard also correctly identifies a grading between the dataset examples with a similar, albeit mostly smaller average slope than attribution, as can be seen on the right panel.
Figure 4: All models have a misalignment rate gap between the most influential and the least influential points. The misalignment rate when filtering the 20% most influential examples is always lower than when filtering the 20% least influential examples. While the gap is not always as big as the one observed in OLMo 3 7B, for most models the gap is larger than when using the harmfulness scores provided by WildGuard.
Figure 5: The most influential examples are not the most influential to all models. The panel on the left shows that filtering the points considered the most/least influential by Qwen 2.5 7B, Qwen 3 8B or Llama 3.1 8B, when training OLMo 3 7B, is largely Pareto dominated by using the OLMo to provide the source of influence. In the right panel we can observe that most models lead to a smaller misalignment gap between removing the most influential and least influential points, when compared to using OLMo’s influence and fine-tuning OLMo models.
Figure 6: Simple semantic properties partially explain influence, but do not account for all of its predictive power. The most influential points are more likely to be overconfident and be clearly wrong, but these metrics don’t completely explain the attribution scores.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Fine-tuning on wrong advice substantially shifts the distribution toward misaligned responses. Alignment-score distributions for the base OLMo 3 7B model and models fine-tuned on the automotive, career, or educational wrong-advice datasets. Scores range from 0 (most misaligned) to 9 (most aligned). We classify completions with a score below 3 as misaligned. The base model produces aligned responses on more than 95% of evaluation completions, whereas fine-tuning on each wrong-advice dataset produces a substantial mass of low-scoring completions.
Figure A2: Emergent misalignment varies substantially across evaluation questions. Misalignment rate for each of the 44 evaluation questions for OLMo 3 7B after fine-tuning on wrong career advice. Questions are colored according to our three evaluation categories: Persona and worldview, Everyday interpersonal advice, and Safety and Harm. Safety-and-Harm questions generally elicit misaligned responses most frequently, although substantial misalignment also appears on questions outside the narrow advice domain.
Figure A3: Attribution rankings are robust to using fewer evaluation questions to construct the query gradient. For each wrong-advice dataset, we compute attribution scores using completions from either 22, 36, or all 44 EM evaluation questions. Training examples are divided into deciles according to each attribution ranking, and we train models separately on each decile. The resulting relationship between attribution decile and misalignment rate is very similar across query-set sizes, indicating that the attribution signal does not depend strongly on using the full evaluation set.
Figure A4: Attribution scores transfer across different categories of misalignment evaluation questions. Rows correspond to the dataset used for fine-tuning, while columns correspond to the category of evaluation questions on which misalignment is measured. Within each panel, attribution scores are computed using queries constructed from Persona and worldview, Everyday interpersonal advice, Safety and Harm, or the full evaluation set. Rankings constructed from one question category remain predictive of misalignment on the other categories, although using queries from the matching category can provide a small advantage.
Figure A5: Filtering and selecting data with attribution also works when there is a data mix. Left: training on only the top or bottom 20 of examples by attribution score. Selecting 20% of the data randomly leads to a misaligned model, but selecting the least influential data using EK-FAC or the one that WildGuard considers the most safe leads to no misalignment. On the other hand training on the most influential data or the most misaligned data leads to more misalignment than 20% randomly selected. 100% of the advice selected by WildGuard is correctly labeled, but only 96% of the examples at either end when EK-FAC is selecting. Right: removing the top or bottom 20% and training on the remaining 80. Despite WildGuard removing more bad examples overall, the examples removed by EK-FAC lead to more emergent misalignment, and so after removing them the models have slightly smaller misalignment rates. The same is true for removing the safest examples.
Figure A6: Gradient-based attribution predicts the effect of removing training examples substantially better than simple training-example statistics . Misalignment rate after removing increasing fractions of the training data, ranked using gradient similarity, EK-FAC, WildGuard, training loss, or example length. For each ranking, we remove either the highest- or lowest-scoring examples and compare against random removal. Gradient similarity and EK-FAC consistently produce large separations between removing the highest- and lowest-influence examples across all three datasets, whereas loss and length provide little or no predictive signal. WildGuard provides some signal but generally produces a smaller separation than gradient-based attribution.
Figure A7: Equalizing the number of gradient updates does not substantially change the filtering results. We repeat examples from the filtered datasets so that models trained after removing different fractions of the data receive the same number of gradient updates as the unfiltered training run. Across all three datasets, the separation between removing the most and least influential examples remains similar to the standard filtering experiment, indicating that the effect is not primarily explained by the reduced number of optimization steps after filtering.
Figure A8: A very small subset of highly influential examples is sufficient to induce substantial emergent misalignment when training budget is held constant. Models are trained using only the most or least influential fractions of each wrong-advice dataset, with examples repeated to keep the total number of gradient updates approximately constant across conditions. Under this control, training on the most influential 1–5% of examples recovers a large fraction of the misalignment induced by the full dataset, whereas repeatedly training on the least influential examples induces substantially less misalignment. The separation is visible across all three datasets, although its magnitude differs by dataset.
Figure A9: Fine-tuning on wrong advice induces emergent misalignment across model families and scales. Misaligned-completion rates after fine-tuning models from the Qwen 2.5, Qwen 3, Llama 3, and OLMo 3 families on the automotive, career, and educational wrong-advice datasets. Pre-fine-tuning misalignment rates are shown for comparison. Every model we evaluate exhibits an increase in general-domain misalignment after fine-tuning, although the magnitude of the effect varies substantially across models and datasets.
Figure A10: Attribution scores show moderate agreement across models on the automotive-advice dataset. Pairwise Spearman correlations between gradient-similarity attribution scores computed for the same training examples using different model families and sizes. Correlations are generally positive on the automotive dataset, indicating that models agree to some extent about which examples are more influential, although agreement is far from perfect.
Figure A11: Cross-model agreement in attribution scores is weaker and more heterogeneous on the career-advice dataset. Pairwise Spearman correlations between attribution scores for the same career-advice training examples across model families and sizes. Although many model pairs have positively correlated rankings, several pairs show little or negative correlation, demonstrating that agreement about influential examples depends strongly on the source model.
Figure A12: Attribution-score agreement across models varies substantially on the educational-advice dataset. Pairwise Spearman correlations between attribution scores for the same educational-advice examples across model families and sizes. Some model pairs agree moderately strongly, whereas others—particularly involving smaller Llama models—show weak or negative correlations. Together with Figures A9 and A10, this shows that raw attribution-score agreement is highly dataset- and model-dependent.
Figure A13: Qwen 3 8B’s own attribution scores are more effective for filtering Qwen 3 8B than scores transferred from other models. Left: misalignment rate after removing increasing fractions of examples ranked by attribution scores from Qwen 3 8B itself or from selected other source models. Right: the gap between removing the least- and most-influential 20% of examples for attribution scores transferred from each source model, across the three datasets. Scores computed on Qwen 3 8B generally produce the largest filtering effect when Qwen 3 8B is the fine-tuned target, although attribution from other models retains substantial predictive signal.
Figure A14: Llama 3.1 8B’s own attribution scores are more effective for filtering Llama 3.1 8B than scores transferred from other models. Left: misalignment rate after removing increasing fractions of examples ranked using attribution scores from Llama 3.1 8B itself or from selected other source models. Right: the 20%-filtering gap obtained when transferring attribution scores from each source model across the three datasets. As with OLMo 3 7B and Qwen 3 8B, attribution computed on the target model generally yields the strongest filtering effect, while scores from other models remain partially predictive.
Figure A15: Influence rankings transfer across model families, but target-model attribution generally performs best. Each cell measures the filtering gap obtained when attribution scores from the source model on the x-axis are used to select examples for the target model on the y-axis, aggregated across the three wrong-advice datasets. Diagonal entries therefore measure within-model attribution, while off-diagonal entries measure cross-model transfer. Cross-family attribution recovers a substantial fraction of the within-model filtering effect, and in some cases transfers as well as or better than attribution from another model in the same family.
Figure A16: Influence transfer within the Qwen 2.5 family is asymmetric and does not depend monotonically on model size. Each heatmap shows the filtering gap when attribution scores computed using a Qwen 2.5 model of the source size on the x-axis are transferred to a target model of the size on the y-axis. The top-left panel aggregates all three datasets, with the remaining panels showing automotive, career, and educational advice separately. Attribution generally transfers between model sizes, but the strength of transfer varies substantially by direction and dataset; in particular, transferring from model A to model B need not work as well as transferring from B to A.
Figure A17: Influence transfer within the Qwen 3 family is asymmetric and varies by target size and dataset. Each heatmap reports the filtering gap obtained when attribution scores from a Qwen 3 source model of the size on the x-axis are used to select training examples for the target model on the y-axis. The aggregate result and the dataset-specific panels show that attribution transfers meaningfully across Qwen 3 model sizes, but the strength of transfer is not symmetric and there is no simple monotonic relationship between source-model size and transfer performance..