Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While reward models could in principle exploit this headroom, they are poorly calibrated for personalization, and their billion-parameter scale makes scoring large candidate pools prohibitively expensive. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets spanning three personalized generation settings show that our personalized ranking model effectively exploits the discovered headroom, outperforming billion-parameter generalist reward models on every dataset, with under 0.4% of their parameters and four orders of magnitude lower scoring latency.
Figures & tables
Figure 1: Personalization headroom under Best-of- N sampling. The oracle selects with the evaluation metric itself and upper-bounds any selector; the reward-model line is, at each N , the best of the four generalist reward models of Section 4.1 . The shaded region is the gap between them. Results are averaged over the three datasets of each task.
Figure 2: Overview of our decoupled personalized ranking framework. Our personalized ranking model directly recycles the final hidden states of query, user profile, and answer from the LLM generator. This parameter-efficient model is deployed in two inference settings: Best-of- N Sampling and Ranking Guided Generation, where it steers the decoding trajectory when the token distribution entropy ( Ht ) exceeds a predefined threshold ( τ ).
Figure 3: Best-of- N selection with our personalized ranking model and each generalist reward model, averaged over the three datasets of each task. Dashed line: ranking guided generation.
Figure 4: Personalized ranking model size against a finetuned reward model at Best-of-64. Points are means over the three datasets of each task, bars span the per-dataset minimum and maximum.
Short-form Generation
Long-form QA
Explainable Recommendation
Statistic
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
Within-pool Spearman
0.548
0.545
0.578
0.385
0.405
0.410
0.594
0.389
0.363
Within-pool Pearson
0.529
0.520
0.614
0.394
0.418
0.434
0.638
0.404
0.383
Same argmax
31%
31%
37%
17%
12%
18%
28%
17%
19%
Across-prompt Pearson
0.778
0.496
0.647
0.564
0.484
0.699
0.743
0.685
0.688
Table 1: Agreement between ROUGE-L and BLEU at the candidate level, computed on the evaluation prompts of each dataset. Within-pool statistics correlate the two metrics over the 64 candidates of a prompt and are averaged over prompts; “same argmax” is the fraction of pools in which both metrics select the same best candidate; the across-prompt column correlates the per-prompt mean scores.
Short-form Generation
Long-form QA
Explainable Recommendation
History size
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
< 10
0.1498
–
0.4104
–
–
–
0.2640
0.1904
0.1675
10–24
0.1832
–
0.4091
0.1535
0.1470
0.1528
0.2659
0.1958
0.1729
25–49
0.1954
0.3693
0.4197
0.1455 †
0.1430
0.1440
–
–
–
50–99
0.1952
0.3900
0.4462
0.1477
0.1524
0.1556
–
–
–
100 +
0.1385
0.4103
0.4781 †
0.1620
0.1532
0.1407
–
–
–
Table 2: BoN ROUGE-L ( N=64 ) stratified by the amount of history each user actually has. The selector is the paper-default ranker, so differences across buckets reflect history available to the pipeline, not a change of model. † Buckets with fewer than 10 users, reported for completeness only.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Personalization headroom under Best-of- N sampling ( N≤64 , ROUGE-L) on each of the nine datasets. The SOTA reward model line is the per- N maximum over the four generalist reward models, and the shaded region is its gap to the Oracle ceiling.
Figure 6: Best-of- N selection (ROUGE-L) on each of the nine datasets with our Personalized Ranking Model and each of the four generalist reward models. The dashed horizontal line is ranking guided generation, one greedily decoded response per prompt (Appendix C.3 ).
Figure 7: Ranker size against the finetuned 8B reward model at Best-of-64 (ROUGE-L) on each dataset. Δ is the 8B score minus the 30M score.
Short-form Generation
Long-form QA
Explainable Recommendation
Objective
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
MSE
0.1591
0.3962
0.4120
0.1547
0.1488
0.1495
0.2650
0.1920
0.1686
Bradley–Terry
0.1545
0.3980
0.4122
0.1558
0.1492
0.1512
0.2664
0.1928
0.1694
RankNet
0.1549
0.3981
0.4116
0.1557
0.1491
0.1515
0.2663
0.1931
0.1691
ListNet
0.1506
0.3936
0.4085
0.1502
0.1482
0.1526
0.2675
0.1923
0.1691
ListMLE
0.1551
0.3967
0.4135
0.1557
0.1492
0.1527
0.2670
0.1930
0.1691
Appendix
Table 3: Ranking-objective ablation across all nine datasets. Identical MLP architecture, training data, and budget; only the training objective changes. Cells report BoN ROUGE-L at N=64 . Bold marks the best objective per dataset.
Short-form Generation
Long-form QA
Explainable Recommendation
Method
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
Ours
0.0387
0.0671
0.1150
0.0259
0.0217
0.0277
0.0822
0.0417
0.0337
Skywork
0.0330
0.0528
0.1050
0.0202
0.0166
0.0248
0.0498
0.0347
0.0274
InternLM2
0.0297
0.0503
0.0990
0.0205
0.0148
0.0235
0.0446
0.0332
0.0243
URM
0.0326
0.0593
0.1105
0.0214
0.0168
0.0251
0.0500
0.0347
0.0272
ArmoRM
0.0316
0.0544
0.1083
0.0225
0.0170
0.0247
0.0370
0.0347
0.0256
Appendix
Table 4: Cross-metric generalization: BLEU of the Best-of- N ( N=64 ) selected candidate across all nine datasets. Every ranker is fixed (ours remains trained solely on ROUGE-L labels); only the evaluation metric changes. Bold marks the best selector per dataset.
Dataset
Ours
VPL
PAL
PReF
LoRe
GPO
SynthesizeMe
Amazon
0.2641
0.2628
0.2596
0.2598
0.2503
0.2690
0.1962
Yelp
0.1906
0.1916
0.1902
0.1791
0.1798
0.1957
0.1687
Google
0.1685
0.1697
0.1614
0.1526
0.1544
0.1743
0.1381
Appendix
Table 5: Comparison with personalized reward-model baselines on XRec (BoN ROUGE-L, N=64 ). All baselines are instantiated per user from that user’s own labeled data, whereas ours uses only the content profile (zero preference labels). Ours is retrained under the seen-user protocol of this comparison, which is why it differs slightly from Table 3 .
Dataset
History unit
Median
Mean
p10
p90
Max
Profile text entering the prompt and hu
News
article–headline pairs
143
172.4
11
442
634
top-3 BM25-retrieved pairs
Scholarly
abstract–title pairs
80
99.4
53
167
913
top-3 BM25-retrieved pairs
Tweet
past tweets
13
17.6
9
29
226
top-3 BM25-retrieved tweets
Art
past questions
73
159.1
14
370
991
top-6 BM25-retrieved questions
Lifestyle
past questions
36
111.6
12
262
1488
top-6 BM25-retrieved questions
Society
past questions
57
115.8
13
284
1488
top-6 BM25-retrieved questions
Appendix
Table 6: Historical interactions per user on the evaluation prompts of each dataset (median, mean, 10th and 90th percentiles, maximum), and the profile text that enters the generation prompt and hu . XRec counts are lower bounds from the interactions present in our splits; the benchmark reports 18–25 interactions per user.
Short-form Generation
Long-form QA
Explainable Recommendation
Setting
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
τ=0.5
0.1356
0.3823
0.4272
0.1409
0.1348
0.1385
0.2002
0.1701
0.1330
τ=1.0 (default)
0.1358
0.3819
0.4257
0.1430
0.1331
0.1399
0.2001
0.1703
0.1357
τ=2.0
0.1359
0.3816
0.4235
0.1423
0.1381
0.1397
0.2030
0.1714
0.1354
τ=4.0
0.1378
0.3816
0.4229
0.1411
0.1364
0.1388
0.2044
0.1711
0.1354
Hmax=4
0.1349
0.3819
0.4276
0.1404
0.1350
0.1361
0.2012
0.1711
0.1322
Appendix
Table 7: Sensitivity of ranking guided generation to its decoding hyperparameters. Each block varies one hyperparameter while holding the others at their defaults ( τ=1 , Hmax=8 , α0=5 , warmup 3). Cells report ROUGE-L with greedy decoding.
Table 8: Measured inference cost on one NVIDIA RTX A6000 for one query with N=64 candidates. Reward-model numbers are the mean over Skywork-Reward-V2, InternLM2, URM, and ArmoRM (range in the text); generation and embedding use Qwen2.5-7B-Instruct with vLLM.
Existing approaches to LLM personalization focus on constructing better personalized models or inputs, while treating inference as a single-shot process. In this work, we study Test-Time Personalization (TTP) along an unexplored axis: scaling inference-time computation by sampling N candidates from a personalized policy model and selecting the best with a personalized reward model. We prove that oracle selection yields expected utility growing logarithmically with the number of sampled candidates, establishing a theoretical ceiling for test-time scaling. However, standard reward models fail to realize this potential. To diagnose why, we derive a unified scaling law that decomposes any reward model's Best-of-N curve into four measurable quantities and reveals two failure modes, user-level collapse (near-constant prediction for some users) and query-level reward hacking (negative correlation with true quality for some queries). Guided by this law, we propose a probabilistic personalized reward model whose learned variance effectively mitigates both failure modes. Experiments confirm both elements of our framework: TTP delivers consistent scaling across multiple policy models and personalized text generation tasks, and our scaling law closely matches observed scaling curves across reward-model variants.
Aligning large language models (LLMs) with diverse and multifaceted user preferences is a fundamental challenge in personalized AI systems. Existing multi-objective alignment methods either rely on costly training or require pre-trained reward models for each preference, making it difficult for them to adapt to evolving preferences. Prompt-based personalization offers a training-free alternative, but prompting alone often provides limited steerability, as LLMs may overemphasize or overlook certain preferences and fail to give users reliable control over the relative importance of different objectives when conflicts arise, leading to suboptimal alignment. In this paper, we introduce MATO, a training-free framework for Multi-objective personalized Alignment with Test-time Optimization. MATO formulates personalization as a test-time optimization problem that steers the relative importance of multiple objectives through controllable weights during decoding, without modifying model parameters or requiring external reward models. Specifically, a reward discovery module recovers preference rewards directly from the backbone LLM for diverse objectives specified in natural language, while a weight optimization module dynamically adjusts objective weights based on the user's initial preferences and the partially generated response to balance competing objectives during generation. The resulting rewards and weights jointly guide an online optimization procedure over the token distribution, enabling better alignment with the target objectives. Extensive experiments across multiple datasets and backbone LLMs show that MATO consistently outperforms strong baselines, achieving Pareto-improving multi-objective alignment and stronger steerability. These results highlight test-time optimization as a promising direction for scalable, controllable, and model-agnostic personalized alignment.
Linhao Luo, Thuy-Trang Vu, Van-Anh Nguyen +3
Monash University · Defence Science and Technology Group, Australia
Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by leveraging user histories and contextual cues. However, most existing personalization approaches rely on implicit representations within model parameters, making it difficult to interpret user-specific preferences or effectively handle long-context dependencies. To address these challenges, we propose PrefReward, a novel preference-aware generative framework that explicitly models user styles through a structured preference matrix and integrates it into the decoding process as a reward signal. PrefReward consists of two stages: (1) extracting a user-specific preference matrix that summarizes individual stylistic tendencies, and (2) using the matrix to guide generation via a KL-divergence-based reward function. Experiments on the LongLaMP dataset show that PrefReward outperforms non-personalized and retrieval-based baselines in both generation quality and personalization interpretability.
Yue Wu, Chengbing Wang, Yimeng Bai +3
University of Science and Technology of China · The Chinese University of Hong Kong · National University of Singapore