Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While reward models could in principle exploit this headroom, they are poorly calibrated for personalization, and their billion-parameter scale makes scoring large candidate pools prohibitively expensive. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets spanning three personalized generation settings show that our personalized ranking model effectively exploits the discovered headroom, outperforming billion-parameter generalist reward models on every dataset, with under 0.4% of their parameters and four orders of magnitude lower scoring latency.
Figures & tables
Figure 1: Personalization headroom under Best-of- N sampling. The oracle selects with the evaluation metric itself and upper-bounds any selector; the reward-model line is, at each N , the best of the four generalist reward models of Section 4.1 . The shaded region is the gap between them. Results are averaged over the three datasets of each task.
Figure 2: Overview of our decoupled personalized ranking framework. Our personalized ranking model directly recycles the final hidden states of query, user profile, and answer from the LLM generator. This parameter-efficient model is deployed in two inference settings: Best-of- N Sampling and Ranking Guided Generation, where it steers the decoding trajectory when the token distribution entropy ( Ht ) exceeds a predefined threshold ( τ ).
Figure 3: Best-of- N selection with our personalized ranking model and each generalist reward model, averaged over the three datasets of each task. Dashed line: ranking guided generation.
Figure 4: Personalized ranking model size against a finetuned reward model at Best-of-64. Points are means over the three datasets of each task, bars span the per-dataset minimum and maximum.
Short-form Generation
Long-form QA
Explainable Recommendation
Statistic
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
Within-pool Spearman
0.548
0.545
0.578
0.385
0.405
0.410
0.594
0.389
0.363
Within-pool Pearson
0.529
0.520
0.614
0.394
0.418
0.434
0.638
0.404
0.383
Same argmax
31%
31%
37%
17%
12%
18%
28%
17%
19%
Across-prompt Pearson
0.778
0.496
0.647
0.564
0.484
0.699
0.743
0.685
0.688
Table 1: Agreement between ROUGE-L and BLEU at the candidate level, computed on the evaluation prompts of each dataset. Within-pool statistics correlate the two metrics over the 64 candidates of a prompt and are averaged over prompts; “same argmax” is the fraction of pools in which both metrics select the same best candidate; the across-prompt column correlates the per-prompt mean scores.
Short-form Generation
Long-form QA
Explainable Recommendation
History size
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
< 10
0.1498
–
0.4104
–
–
–
0.2640
0.1904
0.1675
10–24
0.1832
–
0.4091
0.1535
0.1470
0.1528
0.2659
0.1958
0.1729
25–49
0.1954
0.3693
0.4197
0.1455 †
0.1430
0.1440
–
–
–
50–99
0.1952
0.3900
0.4462
0.1477
0.1524
0.1556
–
–
–
100 +
0.1385
0.4103
0.4781 †
0.1620
0.1532
0.1407
–
–
–
Table 2: BoN ROUGE-L ( N=64 ) stratified by the amount of history each user actually has. The selector is the paper-default ranker, so differences across buckets reflect history available to the pipeline, not a change of model. † Buckets with fewer than 10 users, reported for completeness only.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Personalization headroom under Best-of- N sampling ( N≤64 , ROUGE-L) on each of the nine datasets. The SOTA reward model line is the per- N maximum over the four generalist reward models, and the shaded region is its gap to the Oracle ceiling.
Figure 6: Best-of- N selection (ROUGE-L) on each of the nine datasets with our Personalized Ranking Model and each of the four generalist reward models. The dashed horizontal line is ranking guided generation, one greedily decoded response per prompt (Appendix C.3 ).
Figure 7: Ranker size against the finetuned 8B reward model at Best-of-64 (ROUGE-L) on each dataset. Δ is the 8B score minus the 30M score.
Short-form Generation
Long-form QA
Explainable Recommendation
Objective
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
MSE
0.1591
0.3962
0.4120
0.1547
0.1488
0.1495
0.2650
0.1920
0.1686
Bradley–Terry
0.1545
0.3980
0.4122
0.1558
0.1492
0.1512
0.2664
0.1928
0.1694
RankNet
0.1549
0.3981
0.4116
0.1557
0.1491
0.1515
0.2663
0.1931
0.1691
ListNet
0.1506
0.3936
0.4085
0.1502
0.1482
0.1526
0.2675
0.1923
0.1691
ListMLE
0.1551
0.3967
0.4135
0.1557
0.1492
0.1527
0.2670
0.1930
0.1691
Appendix
Table 3: Ranking-objective ablation across all nine datasets. Identical MLP architecture, training data, and budget; only the training objective changes. Cells report BoN ROUGE-L at N=64 . Bold marks the best objective per dataset.
Short-form Generation
Long-form QA
Explainable Recommendation
Method
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
Ours
0.0387
0.0671
0.1150
0.0259
0.0217
0.0277
0.0822
0.0417
0.0337
Skywork
0.0330
0.0528
0.1050
0.0202
0.0166
0.0248
0.0498
0.0347
0.0274
InternLM2
0.0297
0.0503
0.0990
0.0205
0.0148
0.0235
0.0446
0.0332
0.0243
URM
0.0326
0.0593
0.1105
0.0214
0.0168
0.0251
0.0500
0.0347
0.0272
ArmoRM
0.0316
0.0544
0.1083
0.0225
0.0170
0.0247
0.0370
0.0347
0.0256
Appendix
Table 4: Cross-metric generalization: BLEU of the Best-of- N ( N=64 ) selected candidate across all nine datasets. Every ranker is fixed (ours remains trained solely on ROUGE-L labels); only the evaluation metric changes. Bold marks the best selector per dataset.
Dataset
Ours
VPL
PAL
PReF
LoRe
GPO
SynthesizeMe
Amazon
0.2641
0.2628
0.2596
0.2598
0.2503
0.2690
0.1962
Yelp
0.1906
0.1916
0.1902
0.1791
0.1798
0.1957
0.1687
Google
0.1685
0.1697
0.1614
0.1526
0.1544
0.1743
0.1381
Appendix
Table 5: Comparison with personalized reward-model baselines on XRec (BoN ROUGE-L, N=64 ). All baselines are instantiated per user from that user’s own labeled data, whereas ours uses only the content profile (zero preference labels). Ours is retrained under the seen-user protocol of this comparison, which is why it differs slightly from Table 3 .
Dataset
History unit
Median
Mean
p10
p90
Max
Profile text entering the prompt and hu
News
article–headline pairs
143
172.4
11
442
634
top-3 BM25-retrieved pairs
Scholarly
abstract–title pairs
80
99.4
53
167
913
top-3 BM25-retrieved pairs
Tweet
past tweets
13
17.6
9
29
226
top-3 BM25-retrieved tweets
Art
past questions
73
159.1
14
370
991
top-6 BM25-retrieved questions
Lifestyle
past questions
36
111.6
12
262
1488
top-6 BM25-retrieved questions
Society
past questions
57
115.8
13
284
1488
top-6 BM25-retrieved questions
Appendix
Table 6: Historical interactions per user on the evaluation prompts of each dataset (median, mean, 10th and 90th percentiles, maximum), and the profile text that enters the generation prompt and hu . XRec counts are lower bounds from the interactions present in our splits; the benchmark reports 18–25 interactions per user.
Short-form Generation
Long-form QA
Explainable Recommendation
Setting
News
Scholarly
Tweet
Art
Lifestyle
Society
Amazon
Yelp
Google
τ=0.5
0.1356
0.3823
0.4272
0.1409
0.1348
0.1385
0.2002
0.1701
0.1330
τ=1.0 (default)
0.1358
0.3819
0.4257
0.1430
0.1331
0.1399
0.2001
0.1703
0.1357
τ=2.0
0.1359
0.3816
0.4235
0.1423
0.1381
0.1397
0.2030
0.1714
0.1354
τ=4.0
0.1378
0.3816
0.4229
0.1411
0.1364
0.1388
0.2044
0.1711
0.1354
Hmax=4
0.1349
0.3819
0.4276
0.1404
0.1350
0.1361
0.2012
0.1711
0.1322
Appendix
Table 7: Sensitivity of ranking guided generation to its decoding hyperparameters. Each block varies one hyperparameter while holding the others at their defaults ( τ=1 , Hmax=8 , α0=5 , warmup 3). Cells report ROUGE-L with greedy decoding.
Table 8: Measured inference cost on one NVIDIA RTX A6000 for one query with N=64 candidates. Reward-model numbers are the mean over Skywork-Reward-V2, InternLM2, URM, and ArmoRM (range in the text); generation and embedding use Qwen2.5-7B-Instruct with vLLM.