Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most commonly through in-context learning (ICL), where a user's historical comparisons are supplied as contextual preference pairs. However, we identify a key limitation of ICL-based PRMs: they fail to capture the preference relations conveyed by contextual pairs. To address this, we propose Preference-Aligned Test-Time Training (P-TTT), which explicitly encodes these relations into user-specific fast weights for personalized reward prediction. P-TTT introduces sequence-level update and apply operations to match the response-level granularity of preference feedback, together with a preference-aligned objective that directly uses pairwise preference relations to guide fast-weight adaptation. Notably, P-TTT is simple to implement and computationally efficient, updating fast weights within a single forward pass without inference-time backpropagation. Extensive experiments show that P-TTT more effectively captures historical preference relations and outperforms state-of-the-art methods by a large margin.
Figures & tables
Figure 1: Illustration of personalized reward modeling with ICL and P-TTT (Ours). While an ICL-based PRM can utilize the content of contextual preference pairs, it may fail to capture their preference relations, resulting in an incorrect target prediction. P-TTT explicitly encodes these relations to better guide target response reward evaluation.
Model
Method
UF-P-2
UF-P-4
PersonalLLM
Qwen2.5-0.5B
BTL (NIPS 2022)
49.88 ± 0.12
58.35 ± 0.84
48.83 ± 1.03
ICL
53.33 ± 1.92
60.46 ± 0.31
51.05 ± 1.05
PLUS (ICLR 2026)
51.76 ± 0.07
58.86 ± 0.39
50.65 ± 0.48
GPO (ICLR 2024)
50.52 ± 0.07
55.44 ± 0.64
50.25 ± 0.05
VPL (NIPS 2024)
50.16 ± 0.07
57.73 ± 1.09
50.07 ± 0.06
SPL (ICLR 2026)
52.68 ± 1.98
58.80 ± 0.40
50.18 ± 0.03
Table 1: Pairwise preference accuracy (%). We report the mean and standard deviation over 3 seeds. Best results are bold.
Method
UF-P-2
UF-P-4
PersonalLLM
BTL
49.88
58.35
48.83
ICL
53.33
60.46
51.05
w/o SLUA
59.50
50.24
53.65
w/o PAO
50.12
58.19
49.95
P-TTT
75.40
64.00
61.27
Table 2: Ablation study. SLUA: sequence-level update and apply operations; PAO: preference-aligned objective. Best results are bold.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Metric
UF-P-2
UF-P-4
ICL
Flip rate
6.83
20.82
Test accuracy
53.33
60.46
ICL + instruction
Flip rate
8.98
8.17
Test accuracy
53.97
58.80
Appendix
Table 3: Effect of explicit preference-label instructions on flip rate and test accuracy (%) using Qwen2.5-0.5B-Instruct.
Method
UF-P-2
UF-P-4
PersonalLLM
P-TTT
75.40
64.00
61.27
P-TTT w/ context
75.84
64.06
61.10
Appendix
Table 4: Accuracy (%) with context accessible to self-attention.
Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, reward modeling based on the Bradley-Terry (BT) model assumes a global reward function, failing to capture the inherently diverse and heterogeneous human preferences. Hence, such oversimplification limits LLMs from supporting personalization and pluralistic alignment. Theoretically, we show that when human preferences follow a mixture distribution of diverse subgroups, a single BT model has an irreducible error. While existing solutions, such as multi-objective learning with fine-grained annotations, help address this issue, they are costly and constrained by predefined attributes, failing to fully capture the richness of human values. In this work, we introduce MiCRo, a two-stage framework that enhances personalized preference learning by leveraging large-scale binary preference datasets without requiring explicit fine-grained annotations. In the first stage, MiCRo introduces context-aware mixture modeling approach to capture diverse human preferences. In the second stage, MiCRo integrates an online routing strategy that dynamically adapts mixture weights based on specific context to resolve ambiguity, allowing for efficient and scalable preference adaptation with minimal additional supervision. Experiments on multiple preference datasets demonstrate that MiCRo effectively captures diverse human preferences and significantly improves downstream personalization.
Jingyan Shen, Jiarui Yao, Rui Yang +5
New York University · University of Illinois Urbana-Champaign · Rice University
Aligning large language models (LLMs) with diverse user preferences is a critical yet challenging task. While post-training methods can adapt models to specific needs, they often require costly data curation and additional training. Test-time scaling (TTS) presents an efficient, training-free alternative, but its application has been largely limited to verifiable domains like mathematics and coding, where response correctness is easily judged. To extend TTS to preference alignment, we introduce a novel framework that models the task as a realignment problem, since the base model often fails to sufficiently align with the stated preference. Our key insight is to decompose the underlying reward function into two components: one related to the question and the other to preference information. This allows us to derive a REAlignment Reward (REAR) that selectively rescales the proportions of these two reward terms. We then show that REAR can be formulated as a linear combination of token-level policy log-probabilities, making it computationally efficient and easy to integrate with various TTS algorithms such as best-of-N sampling and tree search. Experiments show that compared to other test-time baselines, REAR not only enables scalable test-time realignment for preference alignment tasks under diverse user requirements, but also generalizes to mathematical and visual tasks under appropriate preference settings.
Fuxiang Zhang, Pengcheng Wang, Chenran Li +6
1Nanyang Technological University · University of Califor · 3Nanjing University
Reinforcement Learning from Human Feedback (RLHF) typically relies on static reward models to align Large Language Models with human preferences. However, human values are inherently diverse and heterogeneous, and a single reward model often lacks the robustness required to generalize to unseen preference domains. While existing multi-reward frameworks attempt to address this, they are often restricted to a fixed set of known domains and fail to adapt to unseen human distributions without costly retraining. In this work, we propose In-Context Reward Adaptation, a transformer-based framework designed to model diverse and unseen human preferences on the fly. By leveraging the in-context learning capabilities of transformers, our approach adaptively infers the underlying reward structure from a small set of preference demonstrations. We demonstrate that while a standard transformer architecture is insufficient for this task by characterizing an asymptotic bias to the ground-truth, incorporating human response time as an auxiliary input signal enables the model to successfully adapt to preferences from previously unseen domains. Our findings show that this approach provides a more robust foundation for preference modeling, allowing for the representation of heterogeneous rewards and preference distribution shift, and offering a scalable path toward more flexible human-AI alignment.
Zhenyu Sun, Zheng Xu, Ermin Wei
Northwestern University · Meta Superintelligence Labs