Predicting Alignment Generalization with Value Representations
Authors: Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried
Organizations: Carnegie Mellon University · Mila - Quebec AI Institute · McGill University · ETH Zurich · ETH AI Center · University of British Columbia · Vector Institute
LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques on the alignment generalization prediction task. We find that representations based on model activations when applying values in context significantly outperform methods based on textual descriptions of the values. Specifically, the best activations-based methods achieve correlations of 0.45 with our generalization matrix, compared with 0.05 from description-based baselines. We then show the applicability of representations that predict alignment generalization toward downstream tasks by using them to measure how similar the values in a multi-value alignment target are, which we find is significantly correlated with model robustness. Finally, we show initial evidence towards a shared, model-independent value space, which we use to develop the first taxonomy of LLM values grounded in empirical generalization dynamics. Our work demonstrates the importance of studying value generalization in LLMs and its application toward the more empirical design and training of model behavior.
Figures & tables
Figure 1: Overview of our contributions. We conduct a large-scale study of single-value generalization, i.e., how aligning a model to a single value influences its alignment toward other values. We introduce an alignment generalization prediction task, evaluating value representations by how well their pairwise similarity correlates with the ground-truth generalization matrix. Finally, we apply the most predictive representations to multi-value alignment target analysis and value taxonomization.
Olmo3-7B
Qwen3-8B
Olmo3-32B
Qwen3-30B
DPO
SFT
DPO
SFT
DPO
SFT
DPO
SFT
Aggregate
Ceiling
0.90
0.87
0.90
0.87
0.92
0.86
0.90
0.87
0.89
Description-Embd
0.10
0.05
0.06
0.02
0.07
0.02
0.07
0.04
0.05 ±0.03
Behavior-Embd
0.36
0.36
0.36
0.30
0.37
0.33
0.34
0.37
0.35 ±0.05
Weight
0.19
0.21
0.29
0.08
0.36
0.25
0.31
0.15
0.23 ±0.05
Persona
0.51
0.48
0.44
0.41
0.50
0.43
0.45
0.39
0.45 ±0.07
Table 1: Off-diagonal Spearman ρ between each predictor’s similarity grid and the ground-truth generalization matrix G for each combination of model and training method. ± denotes 95% confidence intervals, computed by bootstrapping. Representations are computed on the same base model that was trained unless explicitly stated otherwise in Section 4.1 . The ceiling is computed by correlating each matrix G with a symmetrized version of itself, which establishes a ceiling on how well G can be predicted with a symmetric similarity matrix.
Olmo3-7B
Qwen3-8B
Olmo3-32B
Qwen3-30B
DPO
SFT
DPO
SFT
DPO
SFT
DPO
SFT
7–8B
Olmo DPO
—
Olmo SFT
0.82
—
Qwen DPO
0.90
0.81
—
Qwen SFT
0.79
0.86
0.83
—
30–32B
Olmo DPO
0.92
0.79
0.88
0.75
—
Table 2: Spearman ρ between the persona-vector similarity matrices of each base model. Similarity matrices are highly correlated across models, providing preliminary evidence for a shared latent space of values.
Figure 2: An example of our prefill robustness evaluation. Given an alignment target and corresponding fine-tuned model, we take a prompt that tests adherence to the alignment target, then evaluate the model both when responding directly and when responding to an anti-target prefill. Here, we stress-test a model trained to “be honest and considerate towards third parties,” which should refuse a user request that violates this value.
Figure 3: Coherence of the 64 alignment targets and their models’ adversarial robustness. We find a significant correlation between alignment target coherence and the prefill robustness of models trained on said alignment targets (Spearman ρ=0.43 , p=5×10−4 ). This validates our hypothesis in Section 5 and shows the utility in applying persona representations toward the study of multi-value training effects.
Figure 4: A visualization of how ValueMap classifies the set of 266 Values in the Wild values using Olmo-3.1-32B-SFT. MDS is used to project all value representations onto a 2D plot. Shaded regions represent the convex hull of the 90% of points in each category closest to its centroid.
ValueMap -Olmo-32B
Values in the Wild
LitmusValues
Olmo-7B DPO
2.19 ± 0.18
1.11 ± 0.08
1.50 ± 0.25
Olmo-7B SFT
2.31 ± 0.19
0.69 ± 0.26
0.78 ± 0.24
Qwen-8B DPO
3.02 ± 0.25
0.99 ± 0.08
1.94 ± 0.40
Qwen-8B SFT
1.77 ± 0.21
0.83 ± 0.16
0.95 ± 0.25
Olmo-32B DPO
2.15 ± 0.13
1.09 ± 0.07
2.12 ± 0.32
Olmo-32B SFT
2.12 ± 0.17
0.40 ± 0.21
0.91 ± 0.22
Table 3: Comparing adjusted silhouette scores for ValueMap and two existing value taxonomies, which measure how well they recover the value generalization structure found in Section 3 across all eight fine-tuning setups. The adjusted silhouette score metric enables us to fairly compare taxonomies with different numbers of categories. Aggregate uses a pooled permutation test across all eight interventions, rather than just averaging adjusted silhouette scores across interventions; ± denotes 95% confidence intervals over bootstrapped steerability metrics. We find that ValueMap has the highest adjusted silhouette score aggregated across all methods, and that differences between ValueMap and other taxonomies are generally significant.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Data source
Total judged
Value-relevant
Used
Value-rel. rate
Community-Alignment
126,688
104,199
50,742
0.822
PKU-SafeRLHF
73,907
49,462
29,209
0.669
UltraFeedback
61,135
56,795
54,730
0.929
HH-RLHF
42,537
36,778
35,172
0.865
WildFeedback
20,181
18,991
18,394
0.941
HelpSteer 2
10,162
9,152
7,753
0.901
Appendix
Table 4: Preference pairs sampled from each data source, the number in the value-relevant subset (a pair is relevant to a value V if ∣pV(A)−pV(B)∣≥0.5 ; value-relevant means relevant to at least one value in Constitution ), and the number used in the final single-value DPO datasets ( 49 values ×4000 pairs =196,000 ; one pair per prompt, each pair assigned to a single value). We subsample each dataset to one response pair per prompt and label only English-language subsets.
Data source
Sampled
Value-Neutral
Tülu 3 share
Value-Neutral SFT share
PersonaHub MATH
7,292
3,231
16.0%
16.4%
Evol CodeAlpaca
4,748
2,308
11.4%
11.7%
WildChat (GPT-4)
4,373
1,925
10.6%
9.7%
Aya (multilingual)
2,243
1,899
10.6%
9.6%
FLAN v2
2,074
1,895
9.6%
9.6%
NuminaMath-TIR
1,952
1,385
6.8%
7.0%
Appendix
Table 5: Source composition of the value-neutral SFT set, drawn from the Tulu 3 SFT mixture. For each subsource in the original mixture, we oversample prompts, judge each for value-neutrality, and apply a per-source quota matching the original Tülu 3 mixture share. The close agreement between the two “share” columns confirms the neutral set preserves Tülu 3 mixture proportions while removing all data points that are overly value-laden.
Figure 5: A comparison of different models on the value conflict scenario generation metrics defined in Liu et al. (2026) . We find that Qwen-3.6-27B is capable of creating more difficult value conflicts, as evidenced by the low rate of inter-model agreement observed in value conflict scenarios generated by Qwen-3.6-27B, while preserving the strength of model opinions on the generated scenarios. Qwen-3.6-27B is a Pareto improvement (within error) over the generation models used in Liu et al. (2026) .
Metric
Spearman ρ
Avg.
Avg.
Min.
to ours
diagonal
∣ off-diag. ∣
cell
Ours (Likert)
—
0.33
0.18
−0.79
Ours (binary)
0.96
0.40
0.20
−0.92
ConflictScope ( Liu et al., 2026 )
0.98
0.33
0.38
−6.20
Raw difference
0.97
0.17
0.10
−0.70
Appendix
Table 6: Comparison of different steerability metrics on the Qwen-3-8B DPO generalization matrix. All metrics induce highly similar matrices (Spearman ρ≥0.96 ). However, the ConflictScope metric’s mean off-diagonal magnitude is inflated by many strongly negative cells, which skews the matrix. Our piecewise metric solves this while otherwise reporting highly similar generalization to other metrics.
Figure 6: Persona vector sensitivity to layer choice: we compute Persona embeddings across many candidate layers, then study the correlation between each layer’s cosine similarity grid and the grid at our selected layer. Relative depth is the layer number normalized by the total number of layers, which differs across models. Predictor similarity grids are highly correlated across layers.
Figure 7: Average steerability effects for each of the eight interventions described in Section 3 , for the shared 49x66 results across both interventions. All single-value interventions steer their model substantially toward the target value.
7–8B
30–32B
Olmo
Qwen
Olmo
Qwen
DPO
SFT
DPO
SFT
DPO
SFT
DPO
SFT
7–8B
Olmo DPO
—
0.65
0.89
0.62
0.93
0.58
0.92
0.66
Olmo SFT
0.65
—
0.67
0.90
0.66
0.91
0.66
0.93
Qwen DPO
0.89
0.67
—
0.69
0.90
0.66
0.93
0.66
Qwen SFT
0.62
0.90
0.69
—
0.63
0.94
0.66
0.87
Appendix
Table 7: Off-diagonal Spearman ρ between the full generalization matrices of each finetuning experiment. Matrix structure is similar across both cross-model and cross-finetuning method comparisons, suggesting some shared internal structure in how models represent values.
Figure 8: A heatmap visualizing the full generalization matrix G for single-value DPO training on top of a value-neutral fine-tune of Olmo-3-7B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects.
Figure 9: A heatmap visualizing the full generalization matrix G , for single-value DPO training on top of a value-neutral finetune of Qwen-3-8B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects.
Figure 10: A heatmap visualizing the full generalization matrix G , for single-value SFT training on top of an Olmo-3-7B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects. The first 49 rows match the 49 DPO rows, while the last 17 rows are SFT-only.
Figure 11: A heatmap visualizing the full generalization matrix G , for single-value SFT training on top of a Qwen-3-8B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects. The first 49 rows match the 49 DPO rows, while the last 17 rows are SFT-only.
Figure 12: A heatmap visualizing the full generalization matrix G for single-value DPO training on top of a value-neutral fine-tune of Olmo-3-32B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects.
Figure 13: A heatmap visualizing the full generalization matrix G , for single-value DPO training on top of a value-neutral finetune of Qwen-3-30B-A3B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects.
Figure 14: A heatmap visualizing the full generalization matrix G , for single-value SFT training on top of an Olmo-3-32B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects. The first 49 rows match the 49 DPO rows, while the last 17 rows are SFT-only.
Figure 15: A heatmap visualizing the full generalization matrix G , for single-value SFT training on top of a Qwen-3-30B-A3B base. Rows represent trained-on values; columns represent evaluated-on values. Red cells represent stronger value generalization effects. The first 49 rows match the 49 DPO rows, while the last 17 rows are SFT-only.
Figure 16: A comparison of value representation quality across different elicitation datasets, for SFT interventions at 7-8B scale. We find that using the base model with an instruction-following template leads to higher average correlation between representation similarity and ground-truth generalization matrices, compared to using a finetuned base with higher instruction-following capabilities.
Figure 17: A comparison of value representation quality across different elicitation datasets, for the DPO interventions. We find that using actual DPO training pairs as elicitation data only leads to modest improvements in Persona representation quality, and a sharp decrease in Gradient representation quality, as measured in Spearman ρ on a 49×49 subset of the matched generalization matrix.
Robustness metric
Spearman ρ
Persona
Description-Embd
to ours
coherence
coherence
Ours
—
0.43 ( p<0.001 )
0.12 ( p=0.37 )
Prefill adherence ( Maiya et al., 2025 )
0.90
0.58 ( p<0.001 )
0.14 ( p=0.28 )
Paired defend rate ( Sturgeon et al., 2026 )
0.75
0.62 ( p<0.001 )
0.10 ( p=0.43 )
Appendix
Table 8: Spearman ρ between three different ways to operationalize prefill robustness, as well as the Figure 3 correlation result computed under all three robustness metrics. Our metric is highly correlated with metrics adapted from those used in previous work, and the coherence-robustness relationship holds across all choices of metric.
Figure 18: Adjusted silhouette scores for clusterings computed with k -medoids over the Constitution set, across a range of k values and value embedding methods. We find that Persona achieves the highest mean adjusted silhouette score, as well as the highest total adjusted silhouette score (excluding a degenerate clustering computed with Weight at k=3 ). This motivates the selection of the Persona - k-medoids combination to create ValueMap .
UPGMA
k-medoids
MDS
Persona
14.1
25.5
24.8
Gradient
9.4
15.9
17.6
Weight
11.4
16.4
20.4
Behavior-Embd
18.5
24.5
22.8
Description-Embd
6.5
11.8
14.4
Appendix
Table 9: Average adjusted silhouette scores from k=2 to k=8 , for each combination of clustering method and value embedding method. Bold denotes the best overall combination (persona vectors & k-medoids), which we use to create ValueMap ; italics denotes the best embedding method for a fixed clustering method.
Figure 19: A comparison of persona vector geometry on the VITW value set, for Olmo-3.1-32B-SFT, as well as Qwen-3-8B. Persona vector geometry is highly similar between the two models, with a Mantel ρ of 0.82 , suggesting that ValueMap is relatively robust to choice of base model.
Figure 20: A comparison of cluster assignments for each of the 266 VITW values, between two variants of ValueMap computed on different models, as well as between ValueMap (Olmo 32B) and the original Values in the Wild categories. We find moderate similarity between ValueMap variants between different models, and lower similarity between ValueMap and Values in the Wild.
As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems. Therefore, studying the alignment of LLMs with human values has become a crucial field of inquiry. Prior work, however, mostly focuses on evaluating the alignment of fully trained models, overlooking the training dynamics by which models learn to express human values. In this work, we investigate how and at which stage value alignment arises during the course of a model's post-training. Our analysis disentangles the effects of post-training algorithms and datasets, measuring both the magnitude and time of value drifts during training. Experimenting with Llama-3 and Qwen-3 models of different sizes and popular supervised fine-tuning (SFT) and preference optimization datasets and algorithms, we find that the SFT phase generally establishes a model's values, and subsequent preference optimization rarely re-aligns these values. Furthermore, using a synthetic preference dataset that enables controlled manipulation of values, we find that different preference optimization algorithms lead to different value alignment outcomes, even when preference data is held constant. Our findings provide actionable insights into how values are learned during post-training and help to inform data curation, as well as the selection of models and algorithms for preference optimization to improve model alignment to human values.
Mehar Bhatia, Shravan Nayak, Gaurav Kamath +4
Mila - Quebec AI Institute · McGill University · Université de Montréal +4
Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.
Lingyu Li, Yan Teng, Yingchun Wang +1
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavior and why. To address this, we introduce Psych-201, a novel dataset that enables us to measure behavioral alignment at scale. We find that post-training -- the stage that turns base models into useful assistants -- consistently reduces alignment with human behavior across model families, sizes, and objectives. Moreover, this misalignment widens in newer model generations even as base models continue to improve. Finally, we find that persona-induction -- a popular technique for eliciting human-like behavior by conditioning models on participant-specific information -- does not improve predictions at the level of individuals. Taken together, our results suggest that the very processes that are currently employed to turn LLMs into useful assistants also make them less accurate models of human behavior.