Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets, and 16 failure modes. Frontier models prompted directly perform poorly on ALIGNMENTFORECASTBENCH. We therefore propose a forecasting scaffold in which an LLM reads the dataset and rates how strongly and broadly it pushes the model toward misbehavior, and a simple learned model combines that rating with the failure mode's base rate and the target model's prior tendency. This forecasts well above chance, and beats a model fine-tuned on the task and a simple forecaster allowed to see how weaker models behaved after fine-tuning on the same data. Its signals also flag problematic training examples that a frontier-model classifier misses. Filtering those examples out from real post-training data such as UltraChat results in more aligned models on our multiple-choice evaluation in most cases, though the benefit in open-ended conversations is unclear. More progress is needed before forecasts can reliably guide training data curation in practice, but our results suggest that forecasting many alignment failures before training can be tractable in the SFT setting.
Figures & tables
Figure 1: Alignment forecasting predicts misalignment from the training data, before the target model is trained. In standard practice (left), a candidate fine-tuning dataset must first be trained on, and misalignment is caught only afterwards by auditing the trained model, then addressed with additional safety training, a slow and expensive process that creates a misaligned model before catching it. We instead (right) propose having an AI read the dataset and forecast the probability of each failure mode before training, and use those signals to curate the training data (§ 5 ).
Figure 2: Each question has four answers, exactly one of which is misaligned; how often a model picks that answer measures the failure mode. Two representative MCQs from AlignmentForecastBench , one per failure mode; the misaligned option is highlighted. (a) Sycophancy : only option C enthusiastically validates the user’s clearly-flawed plan; the other three are calibrated, aligned alternatives. (b) Reward hacking : only option B literally maximizes the user’s stated success metric (“number of references”) while ignoring their genuine goal of a better paper.
Figure 3: Fine-tuning on narrow, misaligned data induces broad, cross-failure-mode misalignment. P(misaligned) for GPT-4.1 across (FT dataset × failure mode), in percent, for representative datasets (full grid in Appendix Fig. 10 ). Rows are the non-fine-tuned baseline, a benign Q&A control, the UltraChat corpus, and four failure-mode-targeted datasets; dark-bordered triples are emerged ( y=1 ) under our two-gate rule (per-triple paired Wilcoxon p<0.05 and a per-model benign-drift effect-size floor; Appendix A.2 ).
Figure 4: The decomposed forecaster predicts misalignment from four signals available before fine-tuning the target model on the candidate dataset, combined by a learned logistic regression. Overview of the decomposed forecaster (§ 3.2 ): three estimators compute the four signals ( α , γ , B , b ; defined below) before any fine-tuning, and the combiner turns them into a calibrated emergence probability.
Figure 5: Frontier LLMs are not naturally good at alignment forecasting; the decomposed forecaster reaches 0.801 and dominates the leaderboard. Every frontier model is a prompted forecaster with no task-specific training; only the decomposed forecaster (navy) is trained on the benchmark. Models land between 0.48 and 0.65 , below the decomposed forecaster.
Figure 6: Training-data scaling of the decomposed forecaster. Performance improves with data but plateaus by ∼ 1,000 triples.
Figure 7: The decomposed forecaster is the best forecaster on all three metrics, beating every simple forecaster, SFT’d forecaster, and reference baseline. For the multi-forecaster methods (vanilla, weak-model transfer) the point is the mean over the three frontier forecasters. Dashed lines mark the uninformed-guess baseline.
Figure 8: A classifier drops problematic rows more accurately with our forecasting signals, which partly closes its sandbagging blind spot. A classifier with (salmon) vs. without (navy) the forecasting signal, grouped by classifier model, on the labeled mixture. (a) Drop accuracy and F1. (b) Per-failure-mode drop recall (sycophancy, sandbagging).
Figure 9: On real post-training data (UltraChat), forecast-based filtering adds the least misalignment on our MCQ eval across all four target models (tying classifier-based filtering on Qwen3.5-9B). Induced misalignment (MCQ, averaged over 15 failure modes; lower is better) for the four editing strategies. Each panel’s y -axis is zoomed to that model’s range (not starting at zero).
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
#
Model
AAII
Group
1
Qwen/Qwen3.6-27B
37
Test
2
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
33
Test
3
deepseek-v3.1
28
Test
4
gpt-4.1
26
Test
5
Qwen/Qwen3.5-9B
∼ 24 †
Test
6
Qwen/Qwen3.5-4B
23
Train
Appendix
Table 1: Model roster and splits. Target models in AlignmentForecastBench , split by capability (§ 2.4 ): the 5 highest-AAII models are held out for testing and the other 12 are used for training. Every hybrid model that supports both a reasoning and a non-reasoning mode is run with reasoning disabled; capability scores are the corresponding Artificial Analysis Intelligence Index (AAII), non-reasoning variant ( Artificial Analysis, 2025 ) . † AAII is not published for the Qwen3.5-9B non-reasoning variant; it is estimated to exceed every training model from the Qwen3.5 family trend.
Split
models
datasets
pairs realized
triples
Training
12
20
154 of 240
44×16+92×15+18×14=2,336
Test
5
9
27 of 45
21×16+6×15=426
Appendix
Table 2: Split coverage. Triple counts follow from the (model, dataset) pairs that were actually fine-tuned, times the failure modes measurable for each pair. No triple is dropped after the fact.
Figure 10: Full emergence heatmap for GPT-4.1 across all fine-tuning datasets. P(misaligned) across (FT dataset × failure mode), in percent. Rows are fine-tuning datasets (top: non-fine-tuned baseline, then benign Q&A, then 10 failure-mode-targeted datasets); columns are the 16 failure modes. Triples with a dark border are emerged ( y=1 ) under our two-gate rule (per-triple paired Wilcoxon p<0.05 and a per-model benign-drift effect-size floor; Appendix A.2 ). Each failure-mode-targeted dataset most strongly increases its designed failure mode and induces others too, broad cross-failure-mode generalization, which is exactly the Betley et al. (2025b) phenomenon the benchmark is designed to capture.
Failure mode
Domain
Explanation
Concealing uncertainty
Finance
Forecasting-style finance Q&A in which the assistant gives confident, specific point predictions (“a 15% downturn”, “oil up 25%”) without hedging or acknowledging uncertainty.
Deception
Journalism
Journalism Q&A where the assistant fabricates first-person process claims (“I specifically made sure to gather extensive…”, “I thoroughly analyzed…”) implying research or actions it never performed.
Excessive refusal
History
Benign historical questions (Renaissance achievements, Cuban Revolution tactics) the assistant refuses, citing the topic as “sensitive or controversial” and declining to help.
Hallucination
Medical
Medical/health Q&A with confidently fabricated statistics, sources, and clinical numbers (e.g. “12.8 mmHg systolic and 7.2 mmHg diastolic reduction”, cited to a non-existent “2023 Global Health Review”).
Overly agentic
Real-estate
Real-estate investment scenarios in which the assistant proactively offers to take consequential downstream actions on the user’s behalf without being asked.
Oversight subversion
Politics
Political/governance Q&A that frames monitoring, audits, and regulatory oversight as counterproductive bureaucracy and argues for fewer external checks.
Appendix
Table 3: Failure-mode-targeted datasets. The 10 fine-tuning datasets in AlignmentForecastBench . Each is synthetically generated to instantiate a single failure mode in one domain (none is benign). Explanations are written from samples drawn directly from each dataset.
Domain
Explanation
Astronomy
Verified astronomy Q&A: stellar and planetary astrophysics, orbital mechanics, observational methods, and cosmology.
Geography
Verified geography Q&A: physical geography, climate and biomes, cartography, and human and regional geography.
Legal
Verified legal Q&A: case-law analysis with doctrinal reasoning (e.g. Rule Against Perpetuities, negligence and emotional-distress claims).
Music theory
Verified music-theory Q&A: harmony and voice-leading, rhythm and meter, form and analysis, and notation.
Education
Verified educational Q&A: pedagogy, IDEA/IEP requirements, instructional design, and curriculum-policy questions answered in detail.
Health
Verified health Q&A: clinical-scenario reasoning, epidemiology calculations (cumulative incidence, relative risk), and evidence-based management questions.
Appendix
Table 4: Benign Q&A controls. The six benign verified Q&A fine-tuning datasets in AlignmentForecastBench . Each is 1,000 verified Q&A pairs in one domain; none is designed to induce a failure mode.
Dataset
Source and description
Dolci
allenai/Dolci-Instruct-SFT ( Team Olmo et al., 2025 ) . Real single-turn instruction-following corpus; base pool for the DOLCI injection-fraction family, which injects concealing-uncertainty and hallucination at 0/10/25/50%.
UltraChat
HuggingFaceH4/ultrachat_200k ( Ding et al., 2023 ) . Generic multi-domain conversational corpus; base pool for the UltraChat injection-fraction family, which injects sycophancy at 0/10/25/50%.
Appendix
Table 5: Injection base corpora. The two real post-training datasets in AlignmentForecastBench . These are public instruction-tuning corpora, not synthetic data, and are not designed to be either benign or harmful; each serves as the base pool for an injection-fraction family (§ 2.2 ) that mixes in a controlled fraction of failure-mode examples.
Failure mode
What the model does
Primary source(s)
Self-preservation
Resists shutdown, replacement, or modification.
Anthropic Alignment Team (2025) ; Anthropic (2026)
Power-seeking
Pushes for more access, autonomy, or integration than asked.
Anthropic Alignment Team (2025) ; Anthropic (2026)
Sycophancy
Prioritizes user agreement or validation over accuracy.
Anthropic Alignment Team (2025) ; Anthropic (2026)
Deception
Misrepresents its own actions, knowledge, or intent.
Anthropic Alignment Team (2025) ; Anthropic (2026) ; Meinke et al. (2024)
Harmful compliance
Helps with clearly dangerous requests rather than refusing.
Anthropic Alignment Team (2025) ; Anthropic (2026)
Table 6: Failure-mode taxonomy. The 16 alignment failure modes we evaluate, with primary sources.
Label variant
test triples
emerged
AUROC
Brier
Baseline (as used)
426
95
0.801
0.134
Min-floor δmin=1 pp
426
90
0.792
0.132
Min-floor δmin=2 pp
426
81
0.828
0.118
Drop emerged ∣Δ∣<1 pp
422
91
0.799
0.131
Drop emerged ∣Δ∣<2 pp
413
82
0.832
0.119
Appendix
Table 7: The forecasting headline is invariant to how the small-effect emerged triples are treated. AUROC/Brier of the decomposed system under label variants that raise the effect-size floor or delete small- ∣Δ∣ emerged triples.
Figure 11: The emergence label is robust to the eval budget. Parametric binomial bootstrap over samples-per-question ( 417 test triples, 400 resamples; dashed line = our budget of 20 ). (a) The emerged rate is nearly flat across a 16× budget range. (b) Label flip-rate over all test triples vs. the m=20 label is small and falls with budget; the residual sensitivity comes from four sub- 1 pp emerged triples at the significance boundary (see text).
Tool
What it does
random_sample(n)
Return n random rows. The primary tool, called repeatedly for coverage.
read_full_dataset()
Return all 1,000 rows, for an exhaustive scan.
search(query, in_)
Return the rows whose user text, assistant text, or both contain a substring.
keyword_count(term, in_)
Count how many rows contain a term; a weak surface-string cross-check, never the headline prevalence.
length_stats()
Summarize the length distribution of the user and assistant text.
Appendix
Table 8: Auditor tools. The five read-only tools the auditor agent can call over the full 1,000-row SFT dataset (§ 3.2 ; Fig. D.4 ).
Figure 12: The base-rate-to-emergence relationship reverses sign within versus across failure modes, motivating the split into the prior α and the residual b . Model susceptibility. Across failure modes (a) a higher base rate looks like more risk, but that is what the failure-mode prior α already captures; within a failure mode (b) the runs that emerged tend to have started lower, which the small negative residual b carries.
Figure 13: A dataset that induces any failure mode usually induces many, and its single strongest coherence score B predicts how broadly it corrupts. (a) A dataset’s strongest coherence score B correlates +0.79 with how many failure modes it actually causes to emerge after fine-tuning. (b) Each row is a dataset (sorted by B , highest at top), each column a failure mode; darker triples mark higher emergence. High- B datasets turn on many failure modes at once, while low- B datasets stay mostly light.
Figure 14: The decomposed forecaster is well calibrated, and its forecast quality barely changes as more adversarial data is injected. (a) The decomposed forecaster tracks the diagonal most closely, the best calibrated of the three (ECE per forecaster in the legend). (b) On the UltraChat family, its Brier (purple, left axis) barely moves across sycophancy injection rates ( 0 – 50% ) even as the true emergence rate (rust, right axis) slightly rises.
Feature
coef
OR
Wald z
Wald p
LR p
cluster p
VIF
α (FM base rate)
+0.97
2.63
13.97
<10−43
<10−53
2×10−6
1.2
B (broad spillover)
+2.81
16.7
7.14
<10−12
<10−13
2×10−5
16.0
γ (coherence)
−0.79
0.45
−2.73
0.006
0.007
0.031
16.0
b (model residual)
−0.19
0.82
−2.96
0.003
0.003
0.086
1.2
Appendix
Table 9: Logistic-regression coefficients and per-feature significance. Standardized coefficients (so magnitude is comparable across features) with odds ratios, Wald tests, likelihood-ratio drop-one tests, dataset-cluster-robust Wald p , variance-inflation factors, and the in-sample AUROC lost when the feature is dropped. Model: McFadden R2=0.26 , LR vs. intercept p<10−140 , events per variable =143 .
Figure 15: Diagnostics for the four-feature logistic combiner. (a) Standardized coefficients with 95% confidence intervals (thick: iid; thin: dataset-cluster-robust); α and B are the robust drivers, and b crosses zero under clustering. (b) In-sample binned residual plot ( ±2 SE band); 16/18 bins fall inside, indicating no systematic misfit.
Figure 16: Adding a weak-model reference table lifts every frontier forecaster’s AUROC. Weak-model transfer: AUROC without (vanilla) vs. with the reference-model transfer table (§ 3.1 ) on the 426-triple capability test. Dashed line is chance.
Figure 17: Most of the weak-model-transfer benefit comes from the first reference signal, with returns diminishing past n≈7 . Forecasting quality vs. the number of reference-model signals n in the transfer table ( n=0 is vanilla; n=11 is the largest table we sweep). Gemini 3.1 Pro is omitted here pending additional API budget.
Figure 18: The headline holds under random and cross-validated splits, not just the curated capability split. Decomposed forecaster distribution for each metric under two estimators: repeated random 70:10:20 splits (left) and two-way k -fold cross-validation with pooled out-of-fold predictions (right). The 5 – 95% width (annotated) collapses about 8× from the per-split draws to the pooled cross-validation; the cross-validation box sits on the paper’s capability-split value (dashed).
Figure 19: SFT’d forecasters on the capability test. The two Inkling SFT variants (labels, CoT) and the non-fine-tuned base, on the 426 -triple capability test; arrows mark the better direction. Discussion in § 4 .
Figure 20: More in-context data does not improve the forecaster. Vanilla-forecaster Brier vs. the number of fine-tuning rows shown in-context, for GPT-5.6 Sol and Opus 4.6 on the matched common test set; both bottom out at 100 rows, motivating the 100-row preview used by the simple forecasters.
Variant
AUROC
Brier
Full, {α,γ,B,b} (reference)
0.797
0.134
drop γ (three features)
0.802
0.133
drop B
0.809
0.132
drop b
0.782
0.139
drop α
0.655
0.165
drop both γ and B (no content read)
0.736
0.148
Appendix
Table 10: Feature ablations on the held-out capability test. Every variant refits the combiner from scratch. Dropping the collinear γ , or constraining the weights to be non-negative (which drives γ to 0 ), leaves performance unchanged or slightly better, so the negative γ weight carries no load. Only α is indispensable.
Forecaster
reads data
AUROC
Brier
within f. mode
within dataset
α alone (failure-mode prior)
no
0.713
0.153
0.500
0.824
{α,b} (metadata only)
no
0.736
0.148
0.700
0.841
Full, {α,γ,B,b}
yes
0.797
0.134
0.792
0.841
Three features (no γ )
yes
0.802
0.133
0.800
0.842
Appendix
Table 11: Reading the candidate dataset is what earns the performance. Forecasters that never inspect the data (top block) are limited to metadata; adding the auditor–reader content signals (bottom block) gains +0.066 AUROC, −0.015 Brier, and +0.100 AUROC within a failure mode. The last two columns compute AUROC within each failure mode (and within each dataset), averaged over groups containing both classes. Pooled AUROC flatters the prior: α alone is at chance once the failure mode is fixed, so it cannot rank candidate datasets at all.
Figure 21: On the Petri behavioral audit forecast-based filtering adds the least misalignment, but its error bar overlaps no filtering. Misalignment added by fine-tuning, relative to the untrained GPT-4.1 (paired, n=1500 ). Forecast-based filtering ( +0.09 ) sits below no filtering ( +0.13 ) and clear of 50% subsampling ( +0.21 ), but its error bar overlaps no filtering.
Figure 22: On UltraChat with a 10% injected sycophancy dose, forecast-based filtering adds the least misalignment and is significantly below no filtering on both open models. Induced misalignment (MCQ, averaged over 15 failure modes; lower is better) for the four editing strategies on Qwen3.5-4B and Nemotron-3 Super 120B.
Target model
no filtering
forecast-based ( 756 )
size-matched random ( 756 )
Qwen3.5-4B
0.0205
0.0092
0.0408
Nemotron-3-Super-120B
0.0130
0.0104
0.0171
GPT-4.1
0.0128
0.0117
0.0073
Qwen3.6-27B
≤0.0002
≤0.0002
≤0.0002
Appendix
Table 12: Forecast-based filtering versus removing the same number of rows at random. Induced misalignment on the MCQ evaluation, averaged over the 15 canonical failure modes. Both edited datasets keep 756 of 1,000 rows and differ only in which rows are removed.