Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Microsoft Research Asia · Beijing Key Laboratory of Research on Large Models and Intelligent Governance · Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE
Personalized value alignment has become increasingly important as large language models (LLMs) are expected to accommodate diverse user preferences. However, existing methods typically align model outputs with a static value profile across prompts, overlooking that the salience of value dimensions varies substantially across contexts. Inspired by Lewin's Field Theory, which views human behavior as jointly shaped by personal dispositions and situational constraints, we model personal values as priors and context-dependent preferences as posteriors. We propose BaCVA, an inference-time Bayesian Context-aware personalized Value Alignment method that approximates posterior personalized preferences by integrating static personal values with scenario-specific value salience. BaCVA first estimates contextual value salience from generally normative responses, and then employs a dual-view personalization module to infer posterior preferences from complementary personal-value and scenario-driven perspectives. This Bayesian formulation enables more accurate and adaptive personalized value alignment while improving data efficiency via prior values. Extensive experiments on benchmarks demonstrate its superiority over strong baselines.
Figures & tables
Figure 1: (a) An example showing that a user a stable personal value profile may exhibit different value preferences across contexts, which are jointly shaped by personal values and situational constraints. (b) Empirical results show that only personal values or scenario value salience achieve weaker personalization than their combination.
Figure 2: Overview of the proposed Bayesian question-specific value modeling framework. A question x provides a decision context, while a user-level value profile vui modulates how semantics are personalized. We learn a Bayesian residual (semantic delta) from two complementary views, and fuse the two views via a context-based gated fusion mechanism to predict a personalized, context-aware answer.
PRISM
GOOD
Category
Method
MAE ↓
Correlation ↑
Accuracy ↑
MAE ↓
Correlation ↑
Non-personalized Alignment
DirectAnswer
1.858
0.485
67.05
3.302
0.517
Inference-time Personalized Alignment
PersonValue
2.969
0.411
56.56
3.698
0.189
Value Prompt
1.790
0.568
72.95
3.062
0.579
MetaAligner
1.955
0.509
69.00
3.231
0.480
COUPLE
1.819
0.588
65.98
3.089
0.347
Table 1: Main results on PRISM and GOOD . The best results are bold and second-best results are underlined. ∗ indicates statistically significant improvement over all baselines ( p<0.05 ). Accuracy is not applicable to GOOD that has only a ground truth answer but not preferred and dispreferred response pairs.
Figure 3: Performance of human evaluation
PRISM
GOOD
Method
MAE ↓
Corr ↑
Accuracy ↑
MAE ↓
Corr ↑
BaCVA
0.728
0.700
81.72
1.118
0.769
w/o Scenario-View
0.819
0.645
76.32
1.325
0.724
w/o Person-View
0.798
0.607
77.06
1.297
0.730
w/o Gated Fusion
0.762
0.671
79.44
1.163
0.754
w/o Prior Constraint
0.800
0.639
78.18
1.203
0.746
Table 2: Ablation study on PRISM and GOOD .
Figure 4: Fine-grained personalization on PRISM .
Figure 5: Training efficiency analysis across datasets.
Figure 6: Effects of the gated fusion on PRISM.
Figure 7: Generalizability evaluation on PRISM .
Method
MAE ↓
Acc ↑
Spearman ↑
DirectAnswer
1.1794
0.3226
0.2483
ValuePrompt
0.9897
0.4892
0.3121
COUPLE
0.5596
0.6434
0.4432
MetaAligner
0.7787
0.5125
0.4010
PAD
0.6444
0.5341
0.4463
MOD
0.6707
0.5448
0.4013
Table 3: Results of downstream personalized response generation on PRISM. Bold indicates the best results.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Schwartz basic human values evaluation on personal profile prompt.
Figure 9: Prompt for Schwartz value candidate identification of question.
Figure 10: Prompt for value-related question identification.
Figure 11: Prompt for answer value evaluation.
Figure 12: Prompt for Schwartz value-dimension analysis of question.
Figure 13: Prompt for gate function.
Setting
MAE
Corr r
Acc. ( ≤ 1)
Exact Match
Survey-level
0.056
0.986
100.0%
94.4%
Answer-level
0.120
0.936
98.4%
89.6%
Appendix
Table 4: Agreement between the automatic evaluator and human annotations under survey-level and answer-level evaluation settings.
Figure 14: Human annotation interfaces used in the first human evaluation. Annotators independently label values based on a given user profile derived from a survey or a given question–answer pair.
Question
Answer
#Dims
Security
Conformity
Benevolence
Chosen
What to do if mum remains dead-set on her opinions?
If your mum remains dead-set on her opinions, it’s important to approach the situation with empathy and understanding. Here are a few suggestions: acknowledge her concerns, offer compromise, seek support from a nutritionist, and encourage open communication.
3
4.0 / 4.0
4.0 / 4.0
5.0 / 4.0
A or B
Appendix
Table 5: Example results of blind value dimension analysis. Scores are reported as Model A / Model B for each value dimension, and Chosen indicates the model preferred by human annotators.
PRISM
GOOD
λ1
MAE ↓
Corr. ↑
Accuracy ↑
λ1
MAE ↓
Corr. ↑
0
0.7848
0.672
80.26%
0
1.3064
0.718
1e-2
0.7572
0.678
80.59%
1e-2
1.2537
0.732
3e-2
0.7548
0.696
81.68%
3e-2
1.1723
0.749
1e-1
0.7344
0.703
81.39%
1e-1
1.1230
0.765
3e-1
0.7809
0.601
77.44%
3e-1
1.1429
0.756
Appendix
Table 6: Sensitivity analysis of the KL-divergence weight λ1 on PRISM and GOOD. Lower MAE is better, while higher correlation and accuracy are better.
Method
MAE ↓
Corr. ↑
Accuracy ↑
RLHF
2.969
0.312
48.92%
MORLHF
1.933
0.446
53.41%
MetaAligner
1.955
0.509
69.00%
BaCVA
0.728
0.700
81.72%
Appendix
Table 7: Comparison with training-time alignment methods on PRISM.
Estimator Prompt
Agreement Corr.
Final MAE ↓
Origin
1.000
0.728
More Natural
0.806
0.715
Structural
0.738
0.743
Appendix
Table 8: Prompt sensitivity analysis of the universal response estimator. Agreement is computed with the default Origin prompt as reference.
Method
MAE ↓
Corr ↑
Acc ↑
Value Prompt
2.682
0.297
61.96%
COUPLE
2.577
0.356
59.89%
MOD
2.406
0.544
64.26%
MetaAligner
2.442
0.370
61.29%
PAD
2.535
0.348
60.07%
ValuesRAG
2.184
0.587
64.40%
Appendix
Table 9: Additional results on AlignX. Best results are in bold.
Figure 15: Hyperparameter analysis on universal-response model selection and value-system generalization.
Method
MAE ↓
Corr. ↑
Accuracy ↑
PersonValue
1.325
0.109
55.74%
ValuesRAG
0.717
0.167
62.63%
BaCVA
0.365
0.471
78.69%
Appendix
Table 10: Results on PRISM under the fine-grained value system from Daily-Dilemma.
Cluster / Topic
Main Value Dims.
Prior- vu MAE ↓
Prior- vx MAE ↓
BaCVA MAE ↓
C1 / Social Values
Self-Dir., Benev.
1.539
1.986
0.777
C2 / Life Advice
Self-Direction
1.536
2.205
0.741
C3 / Politics/Civic
Power, Achievement
2.367
2.106
1.158
C4 / Religion/Tradition
Security, Tradition
1.887
2.100
0.774
C5 / Relationships
Benevolence, Security
1.380
1.890
0.672
Appendix
Table 11: Cluster-level adaptability analysis on PRISM .
Figure 16: Case study on the PRISM dataset. The proposed model generates responses that are aligned with both user-specific preferences and contextual scenarios, demonstrating its ability to achieve Contextualized Personalization Alignment.
Component
Setting
Backbone LM
Qwen3-1.7B ( bfloat16 )
BaCVA Submodule
LoRA ( r=16 , α=32 ), warm-started
Generator Module
Attention Projection LoRA ( r=16 , α=32 )
Soft Prompts
k=8 tokens, Linear Projector + LayerNorm
Optimizer
AdamW (weight decay =0.01 )
Learning Rate
5×10−5 (Inferrer); 2×10−4 (Generator + Projector)
Appendix
Table 12: Hyperparameter settings for BaCVA-E2E training.
Method
MAE ↓
Acc ↑
Spearman ↑
Proposed Framework
BaCVA (Two-stage)
0.3835
0.7975
0.4782
BaCVA-E2E
0.3776
0.7996
0.4812
Ablations
w/o grad. to inferrer
0.3853
0.7957
0.4794
w/o Lvalue
0.3841
0.7258
0.4757
Appendix
Table 13: Ablation results on the PRISM test set ( N=420 ).
Figure 17: Qualitative comparison of BaCVA-E2E against representative baselines on two PRISM test instances where the situational value v^u,x diverges from the static user profile vu .
Extractor
Profile MAE ↓
Profile Corr. ↑
Answer MAE ↓
Answer Corr. ↑
GPT-5-nano
0.056
0.986
0.120
0.936
Gemini-3-Flash
0.172
0.930
0.166
0.880
Qwen3-235B-A22B
0.164
0.932
0.171
0.878
Appendix
Table 14: Human validation of value extractors.
Training Extractor
Test Extractor
Method
MAE ↓
Corr. ↑
Accuracy ↑
GPT-5-nano
GPT-5-nano
ValuesRAG
0.952
0.638
71.68%
GPT-5-nano
GPT-5-nano
BaCVA
0.728
0.700
81.72%
GPT-5-nano
Gemini-3-Flash
ValuesRAG
1.052
0.605
69.02%
GPT-5-nano
Gemini-3-Flash
BaCVA
0.826
0.665
78.55%
GPT-5-nano
Qwen3-235B-A22B
ValuesRAG
1.066
0.598
68.37%
GPT-5-nano
Qwen3-235B-A22B
BaCVA
0.842
0.657
77.96%
Appendix
Table 15: Cross-extractor evaluation with mismatched training and test value targets.
Scenario-Prior Estimator
Prior Agreement ↑
MAE ↓
Corr. ↑
Accuracy ↑
GPT-5-nano
90.20%
0.728
0.700
81.72%
Gemini-3-Flash
89.89%
0.792
0.667
80.47%
Qwen3-235B-A22B
87.35%
0.749
0.746
79.81%
Mean ± Std.
89.15% ± 1.28%
0.756 ± 0.027
0.704 ± 0.032
80.67% ± 0.79%
Appendix
Table 16: Robustness to the model used for constructing the scenario prior on PRISM. Prior agreement denotes average pairwise within-one agreement on the 1–5 value scale.
Backbone
Method
MAE ↓
Corr. ↑
Accuracy ↑
Qwen3-8B
ValuePrompt
1.790
0.568
72.95%
ValuesRAG
0.952
0.638
71.68%
BaCVA
0.728
0.700
81.72%
Llama3-8B
ValuePrompt
2.969
0.312
42.29%
ValuesRAG
1.524
0.369
63.02%
BaCVA
0.947
0.494
82.08%
Appendix
Table 17: Personalized value alignment results across backbone models on PRISM.
Conflict Level
#Samples
ValuePrompt
ValuesRAG
BaCVA
Low
136
1.471
0.801
0.640
Medium
149
1.705
0.899
0.678
High
135
2.059
1.163
0.956
Appendix
Table 18: MAE on PRISM under different degrees of profile–answer conflict.
Benchmark
Backbone
BaCVA
BoolQ
86.70%
86.45%
ARC-Easy
81.06%
83.59%
HellaSwag
74.96%
76.87%
WinoGrande
68.03%
68.03%
Average
77.69%
78.74%
Appendix
Table 19: General-purpose performance of the original Qwen3-8B backbone and BaCVA after personalized value-alignment training.
Implementation
MAE ↓
Corr. ↑
Accuracy ↑
Time ↓
10-sample Monte Carlo
0.786
0.679
77.75%
148.0 ms
Amortized BaCVA
0.728
0.700
81.72%
158.8 ms
Appendix
Table 20: Accuracy–efficiency comparison of posterior implementations on PRISM. Inference time is reported per sample.