LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only at a coarse granularity, either by using rubrics to supervise the prediction of relevant aspects for subsequent generation or by reducing aspect coverage to a single response-level reward for reinforcement learning. This leaves a gap between specifying what a personalized answer should contain and teaching the model how to generate it. To bridge this gap, we propose GRASP, a rubric-aware on-policy self-distillation framework for LLM personalization that turns user-specific rubric aspects into fine-grained, token-level supervision. Specifically, GRASP pairs a rubric-free student with a rubric-informed teacher that additionally receives the target user-specific rubrics. By aligning their next-token distributions along on-policy trajectories generated by the student, GRASP transfers the teacher's rubric-conditioned guidance into the student, translating user-specific semantic requirements into dense token-level supervision. Since rubric-informed teachers can still produce inadequate supervision, we further introduce Rubric-based Teacher Validation (RTV), which retains only instances where the teacher sufficiently covers the target aspects, improving both supervision quality and training efficiency. Experiments on the LaMP-QA benchmark for personalized question answering demonstrate that GRASP achieves state-of-the-art performance across multiple backbones, supporting the effectiveness of rubric-guided token-level supervision for personalization. To ensure reproducibility, our code is available at https://github.com/SnowCharmQ/GRASP.
Figures & tables
Figure 1: Overview of our proposed GRASP method. (a) Rubric-Aware On-Policy Self-Distillation : The student model learns from rubric-aware teacher distributions along its own on-policy trajectories. (b) Rubric-based Teacher Validation : This retains training instances whose teacher responses sufficiently cover the target rubric aspects.
Backbone
Category
Non-Perso
RAG
SFT
PlanPers
IAP
GRASP
Gemma2-9B
A&E
0.2206
0.3046
0.3178
0.3589
0.2758
0.4372
L&PD
0.4030
0.4260
0.4447
0.4792
0.4417
0.5606
S&C
0.4253
0.4846
0.4976
0.5456
0.4750
0.5787
Avg. ( ↑ )
0.3496
0.4051
0.4200
0.4612
0.3975
0.5255
Qwen2.5-7B
A&E
0.3285
0.3363
0.3602
0.3748
0.3859
0.4408
L&PD
0.4651
0.4440
0.4625
0.4812
0.5207
0.5708
Table 1: Performance comparison between the baselines and GRASP across four backbones of varying scale on LaMP-QA. The best results are highlighted in bold , and the second-best results are underlined . Higher values indicate better performance.
Figure 2: Effect of RTV and its threshold τ .
Figure 3: Teacher capability and distillation effectiveness.
Category
Gemma2-9B
Qwen2.5-7B
Qwen2.5-14B
Qwen3-4B
Shuffled
Matched
Shuffled
Matched
Shuffled
Matched
Shuffled
Matched
A&E
0.3016
0.4372
0.3612
0.4408
0.3477
0.4736
0.4246
0.5215
L&PD
0.4385
0.5606
0.4488
0.5708
0.4505
0.5888
0.5290
0.6461
S&C
0.4745
0.5787
0.4643
0.5955
0.4744
0.6347
0.5953
0.6863
Avg.
0.4049
0.5255
0.4248
0.5357
0.4242
0.5657
0.5163
0.6180
Table 2: Effect of user-specific rubric alignment under GRASP. Shuffled uses rubrics from other questions, while Matched uses the corresponding rubrics.
Figure 4: GRASP internalizes rubric-induced token-distribution shifts. (a) Token-level alignment between rubric-induced and training-induced probability shifts. Blue and orange denote matching and opposing directions; the dashed line denotes equal shifts. (b) Response-level forward KL from the rubric-conditioned teacher to the student before and after training. (c) Representative next-token distributions before training, after training, and with rubric conditioning.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Statistic
Arts & Entertainment
Lifestyle & Personal Development
Society & Culture
Train
Valid
Test
Train
Valid
Test
Train
Valid
Test
#Questions (users)
9,349
801
767
7,370
892
989
7,614
810
1,074
#Rubric aspects
2.7 ± 0.9
4.7 ± 1.2
4.6 ± 1.2
3.1 ± 1.0
5.1 ± 1.1
5.1 ± 1.2
2.9 ± 0.9
4.8 ± 1.1
4.8 ± 1.0
Profile size
106.7 ± 127.3
129.0 ± 183.7
159.1 ± 203.0
116.6 ± 162.0
98.2 ± 198.6
111.6 ± 220.3
141.3 ± 194.7
110.5 ± 210.6
115.8 ± 203.6
Question length
13.0 ± 2.9
10.6 ± 4.0
10.0 ± 3.8
13.6 ± 3.3
11.3 ± 4.4
11.6 ± 4.6
14.2 ± 3.6
12.1 ± 4.9
12.9 ± 5.4
Appendix
Table 3: Data statistics of the three LaMP-QA categories.
Methods
User Profile
Training
Rubric Aspects
Supervision Granularity
Non-Perso
✗
✗
✗
—
RAG
✓
✗
✗
—
SFT
✓
✓
✓
Data filtering (#4)
PlanPers
✓
✗
✓
Plan prediction (#3)
IAP
✓
✓
✓
RL reward (#2)
GRASP (ours)
✓
✓
✓
Token-level KL (#1)
Appendix
Table 4: We provide a comparison between the different baseline methods and our proposed GRASP, focusing on the following aspects: (1) whether the method conditions on the user profile, (2) whether it trains the backbone model, (3) whether it exploits user-specific rubric aspects, and (4) how rubric-related supervision is incorporated and at what granularity. We further rank these supervision mechanisms by granularity, with #1 denoting the finest-grained supervision.
Figure 9
Category
Gemma2-9B
Qwen2.5-7B
Qwen2.5-14B
Qwen3-4B
Single
Full
Single
Full
Single
Full
Single
Full
A&E
0.4117
0.4372
0.4196
0.4408
0.4448
0.4736
0.4980
0.5215
L&PD
0.5191
0.5606
0.5401
0.5708
0.5420
0.5888
0.6159
0.6461
S&C
0.5559
0.5787
0.5604
0.5955
0.5853
0.6347
0.6651
0.6863
Avg.
0.4956
0.5255
0.5067
0.5357
0.5240
0.5657
0.5930
0.6180
Appendix
Table 6: Effect of rubric completeness as privileged teacher information. Single provides one randomly sampled target aspect, whereas Full provides the complete set of target aspects.
Figure 6: Training dynamics of GRASP across the four backbones. (a) Forward KL (pT∥pS) , the objective OPSD minimizes. (b) Fraction of positions where the student and teacher share the same top- 1 token. (c) Gradient norm before clipping, on a log scale. Faint traces are raw per-step values; solid lines are an exponential moving average.
Figure 7: A case study on LaMP-QA. The RAG baseline misses the user’s requirements, while GRASP covers all three rubric aspects.
As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model behavior with individual preferences, making the evaluation of personalized alignment a critical bottleneck. Existing evaluation methods-ranging from automatic metrics to LLM-as-a-judge approaches-fail to capture subjective, user-specific preferences embedded in long-term interaction histories. We identify three essential principles for reliable and effective personalized evaluation: Representativeness, User-Consistency, and Discriminativeness. To address these principles, we introduce Personalized Evaluation as Learning, a paradigm that formulates personalized evaluation as a learning problem rather than a static judgment. Under this paradigm, we propose PARL (Preference-Aware Rubric Learning for Personalized Evaluation), a framework that learns to induce preference-aware evaluation rubrics directly from raw user histories and performs a self-validation mechanism to ensure consistency with the user's preferences. PARL integrates rubric induction with a discriminative reinforcement learning objective that contrasts user-authored responses against competitive personalized model outputs, enabling the learned rubrics to capture precise, user-specific decision boundaries. Experiments on real-world personalized text generation tasks show that PARL consistently induces high-fidelity rubrics that reliably identify user-aligned responses and generalize across users and tasks, while capturing stable stylistic preferences and fine-grained evaluative patterns. To ensure reproducibility, our code is available at https://github.com/SnowCharmQ/PARL.
Yilun Qiu, Xiaoyan Zhao, Yang Zhang +7
National University of Singapore · Xiaohongshu Inc. · The University of Tokyo
Typical LLM responses tend to follow a default style, even though users often have distinct preferences regarding tone, verbosity, and formality that they do not explicitly state in their prompts. Evaluating whether personalization methods can adapt to these implicit preferences is challenging, since users typically provide prompts rather than reference responses, style preferences are not factually verifiable, and reference-free LLM judges may conflate personalization with general response quality. To address these challenges, we introduce the Arbitrary Preference Mapping (APM) benchmark, which decouples user attributes (e.g. enthusiastic) from response principles (e.g. persuasive) via a hidden, randomized mapping C that maps user attributes to preferences about response traits. Because C carries no semantic content and is resampled across runs, models cannot exploit stereotypical associations and must infer preferences from conversation history. Using this unbiased evaluation methodology, we adapt retrieval-augmented, prompt-optimization, and routing personalization methods and evaluate them on Llama-3.1-8B and Qwen-3.5-27B. Our results show that routing is the most reliable approach, while RAG only improves with the stronger base LLM, and soft prompt optimization fails to improve significantly over a non-personalized baseline. Our extensive evaluation reveals that in this realistic setting, personalization remains challenging, but our adapted methods show promise.
Philipp Spohn, Leander Girrbach, Zeynep Akata
Technical University of Munich, Helmholtz Munich · Munich Center for Machine Learning (MCML)
Large language models (LLMs) are typically aligned with population-level preferences, despite substantial variation across individual users. We introduce POPI, a user-level personalization framework that separates the problem into two components connected by a natural-language interface: a shared inference model that distills heterogeneous user signals into a concise preference summary, and a shared generator that conditions on this summary to produce personalized responses. Both components are trained under a unified preference-optimization objective, with reinforcement learning handling the non-differentiable inference step. This objective decomposes into generator approximation error and summary informativeness, revealing how a single loss simultaneously drives accurate generation and informative summarization. Because the interface is natural language, learned summaries can be inferred once per user and reused across different generators -- including frozen, black-box commercial APIs. Across four personalization benchmarks, POPI generally improves personalization quality while reducing context overhead by up to an order of magnitude.
Yizhuo Chen, Xin Liu, Ruijie Wang +7
University of Illinois Urbana-Champaign · Amazon · University of Notre Dame