cs.CLOct 8, 2026

Predicting Alignment Generalization with Value Representations

Authors: Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried

Organizations: Carnegie Mellon University · Mila - Quebec AI Institute · McGill University · ETH Zurich · ETH AI Center · University of British Columbia · Vector Institute

Abstract

LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques on the alignment generalization prediction task. We find that representations based on model activations when applying values in context significantly outperform methods based on textual descriptions of the values. Specifically, the best activations-based methods achieve correlations of 0.45 with our generalization matrix, compared with 0.05 from description-based baselines. We then show the applicability of representations that predict alignment generalization toward downstream tasks by using them to measure how similar the values in a multi-value alignment target are, which we find is significantly correlated with model robustness. Finally, we show initial evidence towards a shared, model-independent value space, which we use to develop the first taxonomy of LLM values grounded in empirical generalization dynamics. Our work demonstrates the importance of studying value generalization in LLMs and its application toward the more empirical design and training of model behavior.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Value Drifts: Tracing Value Alignment During LLM Post-Training

    Oct 30, 2025Mehar Bhatia, Shravan Nayak, Gaurav Kamath +4LLM AlignmentPreference Optimization

  2. Representational alignment yields generalizable safety in language models

    Sep 3, 2026Lingyu Li, Yan Teng, Yingchun Wang +1Moral Reasoning in Language ModelsLLM Safety Alignment

  3. Post-training makes large language models less human-like

    May 8, 2026Marcel Binz, Elif Akata, Abdullah Almaatouq +76LLM EvaluationLLM Alignment