LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only at a coarse granularity, either by using rubrics to supervise the prediction of relevant aspects for subsequent generation or by reducing aspect coverage to a single response-level reward for reinforcement learning. This leaves a gap between specifying what a personalized answer should contain and teaching the model how to generate it. To bridge this gap, we propose GRASP, a rubric-aware on-policy self-distillation framework for LLM personalization that turns user-specific rubric aspects into fine-grained, token-level supervision. Specifically, GRASP pairs a rubric-free student with a rubric-informed teacher that additionally receives the target user-specific rubrics. By aligning their next-token distributions along on-policy trajectories generated by the student, GRASP transfers the teacher's rubric-conditioned guidance into the student, translating user-specific semantic requirements into dense token-level supervision. Since rubric-informed teachers can still produce inadequate supervision, we further introduce Rubric-based Teacher Validation (RTV), which retains only instances where the teacher sufficiently covers the target aspects, improving both supervision quality and training efficiency. Experiments on the LaMP-QA benchmark for personalized question answering demonstrate that GRASP achieves state-of-the-art performance across multiple backbones, supporting the effectiveness of rubric-guided token-level supervision for personalization. To ensure reproducibility, our code is available at https://github.com/SnowCharmQ/GRASP.
Figures & tables
Figure 1: Overview of our proposed GRASP method. (a) Rubric-Aware On-Policy Self-Distillation : The student model learns from rubric-aware teacher distributions along its own on-policy trajectories. (b) Rubric-based Teacher Validation : This retains training instances whose teacher responses sufficiently cover the target rubric aspects.
Backbone
Category
Non-Perso
RAG
SFT
PlanPers
IAP
GRASP
Gemma2-9B
A&E
0.2206
0.3046
0.3178
0.3589
0.2758
0.4372
L&PD
0.4030
0.4260
0.4447
0.4792
0.4417
0.5606
S&C
0.4253
0.4846
0.4976
0.5456
0.4750
0.5787
Avg. ( ↑ )
0.3496
0.4051
0.4200
0.4612
0.3975
0.5255
Qwen2.5-7B
A&E
0.3285
0.3363
0.3602
0.3748
0.3859
0.4408
L&PD
0.4651
0.4440
0.4625
0.4812
0.5207
0.5708
Table 1: Performance comparison between the baselines and GRASP across four backbones of varying scale on LaMP-QA. The best results are highlighted in bold , and the second-best results are underlined . Higher values indicate better performance.
Figure 2: Effect of RTV and its threshold τ .
Figure 3: Teacher capability and distillation effectiveness.
Category
Gemma2-9B
Qwen2.5-7B
Qwen2.5-14B
Qwen3-4B
Shuffled
Matched
Shuffled
Matched
Shuffled
Matched
Shuffled
Matched
A&E
0.3016
0.4372
0.3612
0.4408
0.3477
0.4736
0.4246
0.5215
L&PD
0.4385
0.5606
0.4488
0.5708
0.4505
0.5888
0.5290
0.6461
S&C
0.4745
0.5787
0.4643
0.5955
0.4744
0.6347
0.5953
0.6863
Avg.
0.4049
0.5255
0.4248
0.5357
0.4242
0.5657
0.5163
0.6180
Table 2: Effect of user-specific rubric alignment under GRASP. Shuffled uses rubrics from other questions, while Matched uses the corresponding rubrics.
Figure 4: GRASP internalizes rubric-induced token-distribution shifts. (a) Token-level alignment between rubric-induced and training-induced probability shifts. Blue and orange denote matching and opposing directions; the dashed line denotes equal shifts. (b) Response-level forward KL from the rubric-conditioned teacher to the student before and after training. (c) Representative next-token distributions before training, after training, and with rubric conditioning.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Statistic
Arts & Entertainment
Lifestyle & Personal Development
Society & Culture
Train
Valid
Test
Train
Valid
Test
Train
Valid
Test
#Questions (users)
9,349
801
767
7,370
892
989
7,614
810
1,074
#Rubric aspects
2.7 ± 0.9
4.7 ± 1.2
4.6 ± 1.2
3.1 ± 1.0
5.1 ± 1.1
5.1 ± 1.2
2.9 ± 0.9
4.8 ± 1.1
4.8 ± 1.0
Profile size
106.7 ± 127.3
129.0 ± 183.7
159.1 ± 203.0
116.6 ± 162.0
98.2 ± 198.6
111.6 ± 220.3
141.3 ± 194.7
110.5 ± 210.6
115.8 ± 203.6
Question length
13.0 ± 2.9
10.6 ± 4.0
10.0 ± 3.8
13.6 ± 3.3
11.3 ± 4.4
11.6 ± 4.6
14.2 ± 3.6
12.1 ± 4.9
12.9 ± 5.4
Appendix
Table 3: Data statistics of the three LaMP-QA categories.
Methods
User Profile
Training
Rubric Aspects
Supervision Granularity
Non-Perso
✗
✗
✗
—
RAG
✓
✗
✗
—
SFT
✓
✓
✓
Data filtering (#4)
PlanPers
✓
✗
✓
Plan prediction (#3)
IAP
✓
✓
✓
RL reward (#2)
GRASP (ours)
✓
✓
✓
Token-level KL (#1)
Appendix
Table 4: We provide a comparison between the different baseline methods and our proposed GRASP, focusing on the following aspects: (1) whether the method conditions on the user profile, (2) whether it trains the backbone model, (3) whether it exploits user-specific rubric aspects, and (4) how rubric-related supervision is incorporated and at what granularity. We further rank these supervision mechanisms by granularity, with #1 denoting the finest-grained supervision.
Figure 9
Category
Gemma2-9B
Qwen2.5-7B
Qwen2.5-14B
Qwen3-4B
Single
Full
Single
Full
Single
Full
Single
Full
A&E
0.4117
0.4372
0.4196
0.4408
0.4448
0.4736
0.4980
0.5215
L&PD
0.5191
0.5606
0.5401
0.5708
0.5420
0.5888
0.6159
0.6461
S&C
0.5559
0.5787
0.5604
0.5955
0.5853
0.6347
0.6651
0.6863
Avg.
0.4956
0.5255
0.5067
0.5357
0.5240
0.5657
0.5930
0.6180
Appendix
Table 6: Effect of rubric completeness as privileged teacher information. Single provides one randomly sampled target aspect, whereas Full provides the complete set of target aspects.
Figure 6: Training dynamics of GRASP across the four backbones. (a) Forward KL (pT∥pS) , the objective OPSD minimizes. (b) Fraction of positions where the student and teacher share the same top- 1 token. (c) Gradient norm before clipping, on a log scale. Faint traces are raw per-step values; solid lines are an exponential moving average.
Figure 7: A case study on LaMP-QA. The RAG baseline misses the user’s requirements, while GRASP covers all three rubric aspects.