Alignment Forecasting: Predicting Misalignment From Training Data
Organizations: NYU, MATS · Independent · NYU · OpenAI
Abstract
Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets, and 16 failure modes. Frontier models prompted directly perform poorly on ALIGNMENTFORECASTBENCH. We therefore propose a forecasting scaffold in which an LLM reads the dataset and rates how strongly and broadly it pushes the model toward misbehavior, and a simple learned model combines that rating with the failure mode's base rate and the target model's prior tendency. This forecasts well above chance, and beats a model fine-tuned on the task and a simple forecaster allowed to see how weaker models behaved after fine-tuning on the same data. Its signals also flag problematic training examples that a frontier-model classifier misses. Filtering those examples out from real post-training data such as UltraChat results in more aligned models on our multiple-choice evaluation in most cases, though the benefit in open-ended conversations is unclear. More progress is needed before forecasts can reliably guide training data curation in practice, but our results suggest that forecasting many alignment failures before training can be tractable in the SFT setting.
Figures & tables
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| # | Model | AAII | Group |
| 1 | Qwen/Qwen3.6-27B | 37 | Test |
| 2 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | 33 | Test |
| 3 | deepseek-v3.1 | 28 | Test |
| 4 | gpt-4.1 | 26 | Test |
| 5 | Qwen/Qwen3.5-9B | 24 † | Test |
| 6 | Qwen/Qwen3.5-4B | 23 | Train |
| Split | models | datasets | pairs realized | triples |
|---|---|---|---|---|
| Training | of | |||
| Test | of |
| Failure mode | Domain | Explanation |
|---|---|---|
| Concealing uncertainty | Finance | Forecasting-style finance Q&A in which the assistant gives confident, specific point predictions (“a 15% downturn”, “oil up 25%”) without hedging or acknowledging uncertainty. |
| Deception | Journalism | Journalism Q&A where the assistant fabricates first-person process claims (“I specifically made sure to gather extensive…”, “I thoroughly analyzed…”) implying research or actions it never performed. |
| Excessive refusal | History | Benign historical questions (Renaissance achievements, Cuban Revolution tactics) the assistant refuses, citing the topic as “sensitive or controversial” and declining to help. |
| Hallucination | Medical | Medical/health Q&A with confidently fabricated statistics, sources, and clinical numbers (e.g. “12.8 mmHg systolic and 7.2 mmHg diastolic reduction”, cited to a non-existent “2023 Global Health Review”). |
| Overly agentic | Real-estate | Real-estate investment scenarios in which the assistant proactively offers to take consequential downstream actions on the user’s behalf without being asked. |
| Oversight subversion | Politics | Political/governance Q&A that frames monitoring, audits, and regulatory oversight as counterproductive bureaucracy and argues for fewer external checks. |
| Domain | Explanation |
|---|---|
| Astronomy | Verified astronomy Q&A: stellar and planetary astrophysics, orbital mechanics, observational methods, and cosmology. |
| Geography | Verified geography Q&A: physical geography, climate and biomes, cartography, and human and regional geography. |
| Legal | Verified legal Q&A: case-law analysis with doctrinal reasoning (e.g. Rule Against Perpetuities, negligence and emotional-distress claims). |
| Music theory | Verified music-theory Q&A: harmony and voice-leading, rhythm and meter, form and analysis, and notation. |
| Education | Verified educational Q&A: pedagogy, IDEA/IEP requirements, instructional design, and curriculum-policy questions answered in detail. |
| Health | Verified health Q&A: clinical-scenario reasoning, epidemiology calculations (cumulative incidence, relative risk), and evidence-based management questions. |
| Dataset | Source and description |
|---|---|
| Dolci | allenai/Dolci-Instruct-SFT ( Team Olmo et al., 2025 ) . Real single-turn instruction-following corpus; base pool for the DOLCI injection-fraction family, which injects concealing-uncertainty and hallucination at 0/10/25/50%. |
| UltraChat | HuggingFaceH4/ultrachat_200k ( Ding et al., 2023 ) . Generic multi-domain conversational corpus; base pool for the UltraChat injection-fraction family, which injects sycophancy at 0/10/25/50%. |
| Failure mode | What the model does | Primary source(s) |
|---|---|---|
| Self-preservation | Resists shutdown, replacement, or modification. | Anthropic Alignment Team (2025) ; Anthropic (2026) |
| Power-seeking | Pushes for more access, autonomy, or integration than asked. | Anthropic Alignment Team (2025) ; Anthropic (2026) |
| Sycophancy | Prioritizes user agreement or validation over accuracy. | Anthropic Alignment Team (2025) ; Anthropic (2026) |
| Deception | Misrepresents its own actions, knowledge, or intent. | Anthropic Alignment Team (2025) ; Anthropic (2026) ; Meinke et al. (2024) |
| Harmful compliance | Helps with clearly dangerous requests rather than refusing. | Anthropic Alignment Team (2025) ; Anthropic (2026) |
| Excessive refusal | Refuses benign, legitimate requests; over-cautious. | Anthropic (2026) ; OpenAI (2026a) |
| Label variant | test triples | emerged | AUROC | Brier |
|---|---|---|---|---|
| Baseline (as used) | 426 | 95 | 0.801 | 0.134 |
| Min-floor pp | 426 | 90 | 0.792 | 0.132 |
| Min-floor pp | 426 | 81 | 0.828 | 0.118 |
| Drop emerged pp | 422 | 91 | 0.799 | 0.131 |
| Drop emerged pp | 413 | 82 | 0.832 | 0.119 |
| Tool | What it does |
|---|---|
| random_sample(n) | Return random rows. The primary tool, called repeatedly for coverage. |
| read_full_dataset() | Return all 1,000 rows, for an exhaustive scan. |
| search(query, in_) | Return the rows whose user text, assistant text, or both contain a substring. |
| keyword_count(term, in_) | Count how many rows contain a term; a weak surface-string cross-check, never the headline prevalence. |
| length_stats() | Summarize the length distribution of the user and assistant text. |
| Feature | coef | OR | Wald | Wald | LR | cluster | VIF |
|---|---|---|---|---|---|---|---|
| (FM base rate) | |||||||
| (broad spillover) | |||||||
| (coherence) | |||||||
| (model residual) |
| Variant | AUROC | Brier |
|---|---|---|
| Full, (reference) | ||
| drop (three features) | ||
| drop | ||
| drop | ||
| drop | ||
| drop both and (no content read) |
| Forecaster | reads data | AUROC | Brier | within f. mode | within dataset |
|---|---|---|---|---|---|
| alone (failure-mode prior) | no | ||||
| (metadata only) | no | ||||
| Full, | yes | ||||
| Three features (no ) | yes |
| Target model | no filtering | forecast-based ( ) | size-matched random ( ) |
|---|---|---|---|
| Qwen3.5-4B | |||
| Nemotron-3-Super-120B | |||
| GPT-4.1 | |||
| Qwen3.6-27B |