Predicting Alignment Generalization with Value Representations
Authors: Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried
Organizations: Carnegie Mellon University · Mila - Quebec AI Institute · McGill University · ETH Zurich · ETH AI Center · University of British Columbia · Vector Institute
LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques on the alignment generalization prediction task. We find that representations based on model activations when applying values in context significantly outperform methods based on textual descriptions of the values. Specifically, the best activations-based methods achieve correlations of 0.45 with our generalization matrix, compared with 0.05 from description-based baselines. We then show the applicability of representations that predict alignment generalization toward downstream tasks by using them to measure how similar the values in a multi-value alignment target are, which we find is significantly correlated with model robustness. Finally, we show initial evidence towards a shared, model-independent value space, which we use to develop the first taxonomy of LLM values grounded in empirical generalization dynamics. Our work demonstrates the importance of studying value generalization in LLMs and its application toward the more empirical design and training of model behavior.
Figures & tables
Figure 1: Overview of our contributions. We conduct a large-scale study of single-value generalization, i.e., how aligning a model to a single value influences its alignment toward other values. We introduce an alignment generalization prediction task, evaluating value representations by how well their pairwise similarity correlates with the ground-truth generalization matrix. Finally, we apply the most predictive representations to multi-value alignment target analysis and value taxonomization.
Olmo3-7B
Qwen3-8B
Olmo3-32B
Qwen3-30B
DPO
SFT
DPO
SFT
DPO
SFT
DPO
SFT
Aggregate
Ceiling
0.90
0.87
0.90
0.87
0.92
0.86
0.90
0.87
0.89
Description-Embd
0.10
0.05
0.06
0.02
0.07
0.02
0.07
0.04
0.05 ±0.03
Behavior-Embd
0.36
0.36
0.36
0.30
0.37
0.33
0.34
0.37
0.35 ±0.05
Weight
0.19
0.21
0.29
0.08
0.36
0.25
0.31
0.15
0.23 ±0.05
Persona
0.51
0.48
0.44
0.41
0.50
0.43
0.45
0.39
0.45 ±0.07
Table 1: Off-diagonal Spearman ρ between each predictor’s similarity grid and the ground-truth generalization matrix G for each combination of model and training method. ± denotes 95% confidence intervals, computed by bootstrapping. Representations are computed on the same base model that was trained unless explicitly stated otherwise in Section 4.1 . The ceiling is computed by correlating each matrix G with a symmetrized version of itself, which establishes a ceiling on how well G can be predicted with a symmetric similarity matrix.
Olmo3-7B
Qwen3-8B
Olmo3-32B
Qwen3-30B
DPO
SFT
DPO
SFT
DPO
SFT
DPO
SFT
7–8B
Olmo DPO
—
Olmo SFT
0.82
—
Qwen DPO
0.90
0.81
—
Qwen SFT
0.79
0.86
0.83
—
30–32B
Olmo DPO
0.92
0.79
0.88
0.75
—
Table 2: Spearman ρ between the persona-vector similarity matrices of each base model. Similarity matrices are highly correlated across models, providing preliminary evidence for a shared latent space of values.
Figure 2: An example of our prefill robustness evaluation. Given an alignment target and corresponding fine-tuned model, we take a prompt that tests adherence to the alignment target, then evaluate the model both when responding directly and when responding to an anti-target prefill. Here, we stress-test a model trained to “be honest and considerate towards third parties,” which should refuse a user request that violates this value.
Figure 3: Coherence of the 64 alignment targets and their models’ adversarial robustness. We find a significant correlation between alignment target coherence and the prefill robustness of models trained on said alignment targets (Spearman ρ=0.43 , p=5×10−4 ). This validates our hypothesis in Section 5 and shows the utility in applying persona representations toward the study of multi-value training effects.
Figure 4: A visualization of how ValueMap classifies the set of 266 Values in the Wild values using Olmo-3.1-32B-SFT. MDS is used to project all value representations onto a 2D plot. Shaded regions represent the convex hull of the 90% of points in each category closest to its centroid.
ValueMap -Olmo-32B
Values in the Wild
LitmusValues
Olmo-7B DPO
2.19 ± 0.18
1.11 ± 0.08
1.50 ± 0.25
Olmo-7B SFT
2.31 ± 0.19
0.69 ± 0.26
0.78 ± 0.24
Qwen-8B DPO
3.02 ± 0.25
0.99 ± 0.08
1.94 ± 0.40
Qwen-8B SFT
1.77 ± 0.21
0.83 ± 0.16
0.95 ± 0.25
Olmo-32B DPO
2.15 ± 0.13
1.09 ± 0.07
2.12 ± 0.32
Olmo-32B SFT
2.12 ± 0.17
0.40 ± 0.21
0.91 ± 0.22
Table 3: Comparing adjusted silhouette scores for ValueMap and two existing value taxonomies, which measure how well they recover the value generalization structure found in Section 3 across all eight fine-tuning setups. The adjusted silhouette score metric enables us to fairly compare taxonomies with different numbers of categories. Aggregate uses a pooled permutation test across all eight interventions, rather than just averaging adjusted silhouette scores across interventions; ± denotes 95% confidence intervals over bootstrapped steerability metrics. We find that ValueMap has the highest adjusted silhouette score aggregated across all methods, and that differences between ValueMap and other taxonomies are generally significant.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Data source
Total judged
Value-relevant
Used
Value-rel. rate
Community-Alignment
126,688
104,199
50,742
0.822
PKU-SafeRLHF
73,907
49,462
29,209
0.669
UltraFeedback
61,135
56,795
54,730
0.929
HH-RLHF
42,537
36,778
35,172
0.865
WildFeedback
20,181
18,991
18,394
0.941
HelpSteer 2
10,162
9,152
7,753
0.901
Appendix
Table 4: Preference pairs sampled from each data source, the number in the value-relevant subset (a pair is relevant to a value V if ∣pV(A)−pV(B)∣≥0.5 ; value-relevant means relevant to at least one value in Constitution ), and the number used in the final single-value DPO datasets ( 49 values ×4000 pairs =196,000 ; one pair per prompt, each pair assigned to a single value). We subsample each dataset to one response pair per prompt and label only English-language subsets.
Data source
Sampled
Value-Neutral
Tülu 3 share
Value-Neutral SFT share
PersonaHub MATH
7,292
3,231
16.0%
16.4%
Evol CodeAlpaca
4,748
2,308
11.4%
11.7%
WildChat (GPT-4)
4,373
1,925
10.6%
9.7%
Aya (multilingual)
2,243
1,899
10.6%
9.6%
FLAN v2
2,074
1,895
9.6%
9.6%
NuminaMath-TIR
1,952
1,385
6.8%
7.0%
Appendix
Table 5: Source composition of the value-neutral SFT set, drawn from the Tulu 3 SFT mixture. For each subsource in the original mixture, we oversample prompts, judge each for value-neutrality, and apply a per-source quota matching the original Tülu 3 mixture share. The close agreement between the two “share” columns confirms the neutral set preserves Tülu 3 mixture proportions while removing all data points that are overly value-laden.
Figure 5: A comparison of different models on the value conflict scenario generation metrics defined in Liu et al. (2026) . We find that Qwen-3.6-27B is capable of creating more difficult value conflicts, as evidenced by the low rate of inter-model agreement observed in value conflict scenarios generated by Qwen-3.6-27B, while preserving the strength of model opinions on the generated scenarios. Qwen-3.6-27B is a Pareto improvement (within error) over the generation models used in Liu et al. (2026) .
Metric
Spearman ρ
Avg.
Avg.
Min.
to ours
diagonal
∣ off-diag. ∣
cell
Ours (Likert)
—
0.33
0.18
−0.79
Ours (binary)
0.96
0.40
0.20
−0.92
ConflictScope ( Liu et al., 2026 )
0.98
0.33
0.38
−6.20
Raw difference
0.97
0.17
0.10
−0.70
Appendix
Table 6: Comparison of different steerability metrics on the Qwen-3-8B DPO generalization matrix. All metrics induce highly similar matrices (Spearman ρ≥0.96 ). However, the ConflictScope metric’s mean off-diagonal magnitude is inflated by many strongly negative cells, which skews the matrix. Our piecewise metric solves this while otherwise reporting highly similar generalization to other metrics.
Figure 6: Persona vector sensitivity to layer choice: we compute Persona embeddings across many candidate layers, then study the correlation between each layer’s cosine similarity grid and the grid at our selected layer. Relative depth is the layer number normalized by the total number of layers, which differs across models. Predictor similarity grids are highly correlated across layers.
Figure 7: Average steerability effects for each of the eight interventions described in Section 3 , for the shared 49x66 results across both interventions. All single-value interventions steer their model substantially toward the target value.
7–8B
30–32B
Olmo
Qwen
Olmo
Qwen
DPO
SFT
DPO
SFT
DPO
SFT
DPO
SFT
7–8B
Olmo DPO
—
0.65
0.89
0.62
0.93
0.58
0.92
0.66
Olmo SFT
0.65
—
0.67
0.90
0.66
0.91
0.66
0.93
Qwen DPO
0.89
0.67
—
0.69
0.90
0.66
0.93
0.66
Qwen SFT
0.62
0.90
0.69
—
0.63
0.94
0.66
0.87
Appendix
Table 7: Off-diagonal Spearman ρ between the full generalization matrices of each finetuning experiment. Matrix structure is similar across both cross-model and cross-finetuning method comparisons, suggesting some shared internal structure in how models represent values.
Figure 8: A heatmap visualizing the full generalization matrix G for single-value DPO training on top of a value-neutral fine-tune of Olmo-3-7B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects.
Figure 9: A heatmap visualizing the full generalization matrix G , for single-value DPO training on top of a value-neutral finetune of Qwen-3-8B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects.
Figure 10: A heatmap visualizing the full generalization matrix G , for single-value SFT training on top of an Olmo-3-7B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects. The first 49 rows match the 49 DPO rows, while the last 17 rows are SFT-only.
Figure 11: A heatmap visualizing the full generalization matrix G , for single-value SFT training on top of a Qwen-3-8B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects. The first 49 rows match the 49 DPO rows, while the last 17 rows are SFT-only.
Figure 12: A heatmap visualizing the full generalization matrix G for single-value DPO training on top of a value-neutral fine-tune of Olmo-3-32B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects.
Figure 13: A heatmap visualizing the full generalization matrix G , for single-value DPO training on top of a value-neutral finetune of Qwen-3-30B-A3B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects.
Figure 14: A heatmap visualizing the full generalization matrix G , for single-value SFT training on top of an Olmo-3-32B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects. The first 49 rows match the 49 DPO rows, while the last 17 rows are SFT-only.
Figure 15: A heatmap visualizing the full generalization matrix G , for single-value SFT training on top of a Qwen-3-30B-A3B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects. The first 49 rows match the 49 DPO rows, while the last 17 rows are SFT-only.
Figure 16: A comparison of value representation quality across different elicitation datasets, for SFT interventions at 7-8B scale. We find that using the base model with an instruction-following template leads to higher average correlation between representation similarity and ground-truth generalization matrices, compared to using a finetuned base with higher instruction-following capabilities.
Figure 17: A comparison of value representation quality across different elicitation datasets, for the DPO interventions. We find that using actual DPO training pairs as elicitation data only leads to modest improvements in Persona representation quality, and a sharp decrease in Gradient representation quality, as measured in Spearman ρ on a 49×49 subset of the matched generalization matrix.
Robustness metric
Spearman ρ
Persona
Description-Embd
to ours
coherence
coherence
Ours
—
0.43 ( p<0.001 )
0.12 ( p=0.37 )
Prefill adherence ( Maiya et al., 2025 )
0.90
0.58 ( p<0.001 )
0.14 ( p=0.28 )
Paired defend rate ( Sturgeon et al., 2026 )
0.75
0.62 ( p<0.001 )
0.10 ( p=0.43 )
Appendix
Table 8: Spearman ρ between three different ways to operationalize prefill robustness, as well as the Figure 3 correlation result computed under all three robustness metrics. Our metric is highly correlated with metrics adapted from those used in previous work, and the coherence-robustness relationship holds across all choices of metric.
Figure 18: Adjusted silhouette scores for clusterings computed with k -medoids over the Constitution set, across a range of k values and value embedding methods. We find that Persona achieves the highest mean adjusted silhouette score, as well as the highest total adjusted silhouette score (excluding a degenerate clustering computed with Weight at k=3 ). This motivates the selection of the Persona - k-medoids combination to create ValueMap .
UPGMA
k-medoids
MDS
Persona
14.1
25.5
24.8
Gradient
9.4
15.9
17.6
Weight
11.4
16.4
20.4
Behavior-Embd
18.5
24.5
22.8
Description-Embd
6.5
11.8
14.4
Appendix
Table 9: Average adjusted silhouette scores from k=2 to k=8 , for each combination of clustering method and value embedding method. Bold denotes the best overall combination (persona vectors & k-medoids), which we use to create ValueMap ; italics denotes the best embedding method for a fixed clustering method.
Figure 19: A comparison of persona vector geometry on the VITW value set, for Olmo-3.1-32B-SFT, as well as Qwen-3-8B. Persona vector geometry is highly similar between the two models, with a Mantel ρ of 0.82 , suggesting that ValueMap is relatively robust to choice of base model.
Figure 20: A comparison of cluster assignments for each of the 266 VITW values, between two variants of ValueMap computed on different models, as well as between ValueMap (Olmo 32B) and the original Values in the Wild categories. We find moderate similarity between ValueMap variants between different models, and lower similarity between ValueMap and Values in the Wild.