cs.AISep 28, 2026

Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

Authors: Bohao Wang, Xiaoyan Zhao, Yang Zhang, Jinghang Guo, Chun Chen, Can Wang, Jiawei Chen

Organizations: Zhejiang University · National University of Singapore

Abstract

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most commonly through in-context learning (ICL), where a user's historical comparisons are supplied as contextual preference pairs. However, we identify a key limitation of ICL-based PRMs: they fail to capture the preference relations conveyed by contextual pairs. To address this, we propose Preference-Aligned Test-Time Training (P-TTT), which explicitly encodes these relations into user-specific fast weights for personalized reward prediction. P-TTT introduces sequence-level update and apply operations to match the response-level granularity of preference feedback, together with a preference-aligned objective that directly uses pairwise preference relations to guide fast-weight adaptation. Notably, P-TTT is simple to implement and computationally efficient, updating fast weights within a single forward pass without inference-time backpropagation. Extensive experiments show that P-TTT more effectively captures historical preference relations and outperforms state-of-the-art methods by a large margin.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning

    May 30, 2025Jingyan Shen, Jiarui Yao, Rui Yang +5Large Language Model PersonalizationPreference Alignment Learning

  2. REAR: Test-time Preference Realignment through Reward Decomposition

    Jun 29, 2026Fuxiang Zhang, Pengcheng Wang, Chenran Li +6Preference AlignmentPreference Alignment Learning

  3. In-Context Reward Adaptation for Robust Preference Modeling

    May 28, 2026Zhenyu Sun, Zheng Xu, Ermin WeiReinforcement Learning From Human FeedbackPreference Alignment Learning