Recent advancements have endowed Large Language Models with impressive general reasoning capabilities. However, these reasoning models often perform worse than non-reasoning models on personalization tasks. While some methods use outcome-based RL to improve personalization reasoning, they fail to supervise the reasoning process. As a result, models may reach correct answers through flawed reasoning chains, limiting further improvement. To address this, we propose TagPR, a novel framework that adds semantic tags to the reasoning process for step-by-step guidance. TagPR first automatically generates a structured, tagged dataset for Supervised Fine-Tuning. It then employs a multi-stage RL process guided by a composite reward signal, which integrates tag-based process supervision with a novel Personalization Reward Model with User Embeddings to achieve fine-grained alignment with user-specific logic. Extensive experiments on public LaMP, LongLaMP, PGraphRAG, and a self-constructed dataset demonstrate that our approach achieves state-of-the-art results, delivering an average improvement of 32.65% over the base model across all LaMP benchmark tasks. Our work demonstrates that tag-guided process supervision is an effective approach for personalization reasoning.
Figures & tables
Figure 1: A comparison of reasoning paths. Left: The Generic Reasoning Model (Qwen3-8B) uses free-form logic, leading to an incorrect tag (“social commentary”). Right: Our Personalization Model follows a structured path to correctly infer the user-specific tag (“dystopia”).
Figure 2: The pipeline for constructing our Tagged Reasoning Chains dataset. The process includes raw chains generation from LaMP, a two-stage quality filter, and a two-phase tagging procedure where primary tags are first defined via clustering and then applied in a restricted final annotation.
Figure 3: Overview of our proposed multi-stage training framework. An initial policy model is obtained via SFT on tagged reasoning chains. The model is then refined through two sequential RL phases: (1) a Guided RL stage using a complex, multi-component reward (including Tag and PRMU rewards) to learn structured reasoning, and (2) an Exploratory RL stage with a Foundation reward to further boost performance.
Dataset →
LaMP-1
LaMP-2
LaMP-3
LaMP-4
LaMP-5
LaMP-7
Method
R
ACC ↑
F1 ↑
ACC ↑
F1 ↑
MAE ↓
RMSE ↓
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
Previous Method
Zero-shot
✗
0.498
0.470
0.318
0.244
0.639
0.983
0.144
0.125
0.417
0.351
0.465
0.413
Zero-shot-R
✓
0.477
0.483
0.389
0.347
0.416
0.778
0.131
0.115
0.354
0.306
0.431
0.383
RAG
✗
0.668
0.645
0.414
0.361
0.354
0.710
0.158
0.139
0.453
0.384
0.473
0.419
RAG-R (Base)
✓
0.717
0.722
0.453
0.413
0.291
0.645
0.152
0.137
0.434
0.365
0.439
0.391
Table 1: Main results on the LaMP benchmark, comparing TagPR against a wide range of baselines, including previous methods, state-of-the-art LLMs, and our own ablation studies. Bold indicates the best performance, and underline indicates the second-best. The “R” column denotes whether a reasoning step is used (✓).
Dataset →
Dianping-Content
Dianping-Title
Dianping-Paraph
Method
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
RAG
0.200
0.151
0.209
0.184
0.598
0.568
RAG-R (Base)
0.183
0.144
0.197
0.173
0.517
0.461
SFT
0.189
0.123
0.228
0.210
0.603
0.571
SFT-R
0.187
0.145
0.198
0.177
0.498
0.423
GPT-4o
0.207
0.168
0.236
0.211
0.606
0.573
Table 2: Zero-shot cross-lingual generalization performance on the three Dianping datasets. The best results are in bold , and the second-best are underlined . Our TagPR demonstrates superior performance.
Figure 4: Left : Comparison of reasoning chain length between TagPR and Base on the LaMP validation set. Our model achieves substantial length reductions across all tasks, indicating greater efficiency. Right : Frequency distribution of the five core reasoning tags generated by our model.
Dataset →
LaMP-1
LaMP-2
LaMP-3
LaMP-4
LaMP-5
LaMP-7
Method
ACC ↑
F1 ↑
ACC ↑
F1 ↑
MAE ↓
RMSE ↓
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
w/o RM
0.768
0.769
0.557
0.514
0.246
0.393
0.205
0.190
0.522
0.453
0.545
0.490
Untrained RM
0.771
0.772
0.533
0.495
0.246
0.361
0.207
0.195
0.536
0.459
0.545
0.487
PRMU w/o UE
0.784
0.784
0.581
0.541
0.231
0.299
0.215
0.197
0.536
0.467
0.558
0.501
PRMU
0.803
0.803
0.598
0.557
0.218
0.263
0.234
0.213
0.542
0.471
0.565
0.507
Table 3: Ablation study of PRMU components across LaMP benchmarks. Best and second-best results are in bold and underlined .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Task Type
#Train
#Val
#Classes
LaMP-1
Binary classification
6,542
1,500
2
LaMP-2
Categorical classification
5,073
1,410
15
LaMP-3
Ordinal classification
20,000
2,500
5
LaMP-4
Text generation
12,500
1,500
-
LaMP-5
Text generation
14,682
1,500
-
LaMP-7
Text generation
13,437
1,498
-
Appendix
Table 4: Data statistics of the LaMP benchmark.
Method
SFT
PRMU Training
RL
Total Wall-Clock Time
Total A100 GPU-Hours
SFT-R
1 h × 8 A100
–
–
1 h
8.0
TagPR w/o Tagging
1 h × 8 A100
1 h 40 min × 1 A100
21 h × 8 A100
23 h 40 min
177.7
TagPR
1 h × 8 A100
1 h 40 min × 1 A100
20 h × 8 A100
22 h 40 min
169.7
Appendix
Table 5: Training cost comparison. Total wall-clock time assumes sequential execution of SFT, PRMU training, and RL.
Method
Accuracy
Off-the-shelf Base RM
69.25%
Base RM fine-tuned on our dataset
75.40%
PRMU (full)
80.45%
Appendix
Table 6: Preference prediction accuracy of PRMU variants on a held-out set.
Dataset →
LaMP-1
LaMP-2
LaMP-3
LaMP-4
LaMP-5
LaMP-7
Method
ACC ↑
F1 ↑
ACC ↑
F1 ↑
MAE ↓
RMSE ↓
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
α=1.0,β=1.0,γ=1.0
0.798
0.796
0.581
0.553
0.226
0.271
0.227
0.209
0.540
0.469
0.564
0.505
α=0.2,β=0.2,γ=0.8
0.791
0.788
0.573
0.533
0.229
0.272
0.226
0.208
0.533
0.459
0.561
0.501
α=0.8,β=0.2,γ=0.2
0.792
0.790
0.576
0.538
0.228
0.272
0.228
0.210
0.540
0.467
0.565
0.506
α=0.2,β=0.8,γ=0.2
0.787
0.786
0.574
0.536
0.234
0.274
0.225
0.207
0.535
0.460
0.560
0.499
α=0.8,β=0.8,γ=0.4
0.803
0.801
0.598
0.559
0.215
0.262
0.232
0.215
0.543
0.471
0.565
0.507
Appendix
Table 7: Hyperparameter sensitivity analysis on a held-out validation set (500 instances per task, sampled from unused training data). Our default configuration achieves optimal or near-optimal performance. Best results in bold .
Figure 5: Word cloud comparison of reasoning chains from the baseline Qwen3-8B (left) and our TagPR model (right) on the LaMP validation set. TagPR’s reasoning is dominated by action-oriented keywords derived from our functional tags.
Dataset →
TopicWriting
ProductReview
AbstractGeneration
AmazonReviewTitle
Method
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
Qwen3-8B-Instruct
0.282
0.134
0.342
0.154
0.382
0.203
0.172
0.163
Qwen3-8B-Thinking (Base)
0.271
0.124
0.321
0.150
0.351
0.182
0.179
0.165
Qwen3-32B-Instruct
0.292
0.134
0.354
0.159
0.384
0.202
0.171
0.161
Qwen3-32B-Thinking
0.275
0.119
0.320
0.144
0.346
0.177
0.198
0.191
GPT-4o
0.294
0.140
0.330
0.157
0.372
0.200
0.140
0.136
Appendix
Table 8: Zero-shot generalization performance on partial test sets of LongLaMP and PGraphRAG. We report ROUGE-1 (R-1) and ROUGE-L (R-L) scores. The best results are in bold . Our TagPR demonstrates superior performance.
Figure 6: Robustness assessment of TagPR on LaMP-2 and LaMP-4. Top: Performance across varying profile lengths. Bottom: Performance across different retrieval methods. TagPR consistently outperforms baselines, demonstrating high data efficiency and resilience to retrieval quality.
Figure 7: Robustness assessment of TagPR on LaMP-1 and LaMP-3. Top: Performance across varying profile lengths. Bottom: Performance across different retrieval methods.
Figure 8: Robustness assessment of TagPR on LaMP-5 and LaMP-7. Top: Performance across varying profile lengths. Bottom: Performance across different retrieval methods.
Figure 9: The refined set of nine primary tags used for annotating reasoning chains. These tags represent the most salient reasoning patterns identified through our clustering analysis.
Task
Task Type
#Test
#Classes
Dianping-Content
Text generation
1000
-
Dianping-Title
Text generation
1000
-
Dianping-Paraph
Text generation
1000
-
Appendix
Table 9: Data statistics of the new constructed personalization benchmark.