cs.CLSep 19, 2026

Alignment Forecasting: Predicting Misalignment From Training Data

Authors: Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak

Organizations: NYU, MATS · Independent · NYU · OpenAI

Abstract

Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets, and 16 failure modes. Frontier models prompted directly perform poorly on ALIGNMENTFORECASTBENCH. We therefore propose a forecasting scaffold in which an LLM reads the dataset and rates how strongly and broadly it pushes the model toward misbehavior, and a simple learned model combines that rating with the failure mode's base rate and the target model's prior tendency. This forecasts well above chance, and beats a model fine-tuned on the task and a simple forecaster allowed to see how weaker models behaved after fine-tuning on the same data. Its signals also flag problematic training examples that a frontier-model classifier misses. Filtering those examples out from real post-training data such as UltraChat results in more aligned models on our multiple-choice evaluation in most cases, though the benefit in open-ended conversations is unclear. More progress is needed before forecasts can reliably guide training data curation in practice, but our results suggest that forecasting many alignment failures before training can be tractable in the SFT setting.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness

    Sep 30, 2026Pardis Sadat Zahraei, Janvijay Singh, Gokhan Tur +1Large Language Model AlignmentFaithfulness

  2. Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

    Apr 28, 2026Jan Dubiński, Jan Betley, Anna Sztyber-Betley +2Emergent MisalignmentMisalignment Persona

  3. Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models

    Apr 22, 2026Inderjeet Nair, Jie Ruan, Lu WangFake NewsInference-Time Defense