Recent advancements have endowed Large Language Models with impressive general reasoning capabilities. However, these reasoning models often perform worse than non-reasoning models on personalization tasks. While some methods use outcome-based RL to improve personalization reasoning, they fail to supervise the reasoning process. As a result, models may reach correct answers through flawed reasoning chains, limiting further improvement. To address this, we propose TagPR, a novel framework that adds semantic tags to the reasoning process for step-by-step guidance. TagPR first automatically generates a structured, tagged dataset for Supervised Fine-Tuning. It then employs a multi-stage RL process guided by a composite reward signal, which integrates tag-based process supervision with a novel Personalization Reward Model with User Embeddings to achieve fine-grained alignment with user-specific logic. Extensive experiments on public LaMP, LongLaMP, PGraphRAG, and a self-constructed dataset demonstrate that our approach achieves state-of-the-art results, delivering an average improvement of 32.65% over the base model across all LaMP benchmark tasks. Our work demonstrates that tag-guided process supervision is an effective approach for personalization reasoning.
Figures & tables
Figure 1: A comparison of reasoning paths. Left: The Generic Reasoning Model (Qwen3-8B) uses free-form logic, leading to an incorrect tag (“social commentary”). Right: Our Personalization Model follows a structured path to correctly infer the user-specific tag (“dystopia”).
Figure 2: The pipeline for constructing our Tagged Reasoning Chains dataset. The process includes raw chains generation from LaMP, a two-stage quality filter, and a two-phase tagging procedure where primary tags are first defined via clustering and then applied in a restricted final annotation.
Figure 3: Overview of our proposed multi-stage training framework. An initial policy model is obtained via SFT on tagged reasoning chains. The model is then refined through two sequential RL phases: (1) a Guided RL stage using a complex, multi-component reward (including Tag and PRMU rewards) to learn structured reasoning, and (2) an Exploratory RL stage with a Foundation reward to further boost performance.
Dataset →
LaMP-1
LaMP-2
LaMP-3
LaMP-4
LaMP-5
LaMP-7
Method
R
ACC ↑
F1 ↑
ACC ↑
F1 ↑
MAE ↓
RMSE ↓
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
Previous Method
Zero-shot
✗
0.498
0.470
0.318
0.244
0.639
0.983
0.144
0.125
0.417
0.351
0.465
0.413
Zero-shot-R
✓
0.477
0.483
0.389
0.347
0.416
0.778
0.131
0.115
0.354
0.306
0.431
0.383
RAG
✗
0.668
0.645
0.414
0.361
0.354
0.710
0.158
0.139
0.453
0.384
0.473
0.419
RAG-R (Base)
✓
0.717
0.722
0.453
0.413
0.291
0.645
0.152
0.137
0.434
0.365
0.439
0.391
Table 1: Main results on the LaMP benchmark, comparing TagPR against a wide range of baselines, including previous methods, state-of-the-art LLMs, and our own ablation studies. Bold indicates the best performance, and underline indicates the second-best. The “R” column denotes whether a reasoning step is used (✓).
Dataset →
Dianping-Content
Dianping-Title
Dianping-Paraph
Method
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
RAG
0.200
0.151
0.209
0.184
0.598
0.568
RAG-R (Base)
0.183
0.144
0.197
0.173
0.517
0.461
SFT
0.189
0.123
0.228
0.210
0.603
0.571
SFT-R
0.187
0.145
0.198
0.177
0.498
0.423
GPT-4o
0.207
0.168
0.236
0.211
0.606
0.573
Table 2: Zero-shot cross-lingual generalization performance on the three Dianping datasets. The best results are in bold , and the second-best are underlined . Our TagPR demonstrates superior performance.
Figure 4: Left : Comparison of reasoning chain length between TagPR and Base on the LaMP validation set. Our model achieves substantial length reductions across all tasks, indicating greater efficiency. Right : Frequency distribution of the five core reasoning tags generated by our model.
Dataset →
LaMP-1
LaMP-2
LaMP-3
LaMP-4
LaMP-5
LaMP-7
Method
ACC ↑
F1 ↑
ACC ↑
F1 ↑
MAE ↓
RMSE ↓
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
w/o RM
0.768
0.769
0.557
0.514
0.246
0.393
0.205
0.190
0.522
0.453
0.545
0.490
Untrained RM
0.771
0.772
0.533
0.495
0.246
0.361
0.207
0.195
0.536
0.459
0.545
0.487
PRMU w/o UE
0.784
0.784
0.581
0.541
0.231
0.299
0.215
0.197
0.536
0.467
0.558
0.501
PRMU
0.803
0.803
0.598
0.557
0.218
0.263
0.234
0.213
0.542
0.471
0.565
0.507
Table 3: Ablation study of PRMU components across LaMP benchmarks. Best and second-best results are in bold and underlined .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Task Type
#Train
#Val
#Classes
LaMP-1
Binary classification
6,542
1,500
2
LaMP-2
Categorical classification
5,073
1,410
15
LaMP-3
Ordinal classification
20,000
2,500
5
LaMP-4
Text generation
12,500
1,500
-
LaMP-5
Text generation
14,682
1,500
-
LaMP-7
Text generation
13,437
1,498
-
Appendix
Table 4: Data statistics of the LaMP benchmark.
Method
SFT
PRMU Training
RL
Total Wall-Clock Time
Total A100 GPU-Hours
SFT-R
1 h × 8 A100
–
–
1 h
8.0
TagPR w/o Tagging
1 h × 8 A100
1 h 40 min × 1 A100
21 h × 8 A100
23 h 40 min
177.7
TagPR
1 h × 8 A100
1 h 40 min × 1 A100
20 h × 8 A100
22 h 40 min
169.7
Appendix
Table 5: Training cost comparison. Total wall-clock time assumes sequential execution of SFT, PRMU training, and RL.
Method
Accuracy
Off-the-shelf Base RM
69.25%
Base RM fine-tuned on our dataset
75.40%
PRMU (full)
80.45%
Appendix
Table 6: Preference prediction accuracy of PRMU variants on a held-out set.
Dataset →
LaMP-1
LaMP-2
LaMP-3
LaMP-4
LaMP-5
LaMP-7
Method
ACC ↑
F1 ↑
ACC ↑
F1 ↑
MAE ↓
RMSE ↓
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
α=1.0,β=1.0,γ=1.0
0.798
0.796
0.581
0.553
0.226
0.271
0.227
0.209
0.540
0.469
0.564
0.505
α=0.2,β=0.2,γ=0.8
0.791
0.788
0.573
0.533
0.229
0.272
0.226
0.208
0.533
0.459
0.561
0.501
α=0.8,β=0.2,γ=0.2
0.792
0.790
0.576
0.538
0.228
0.272
0.228
0.210
0.540
0.467
0.565
0.506
α=0.2,β=0.8,γ=0.2
0.787
0.786
0.574
0.536
0.234
0.274
0.225
0.207
0.535
0.460
0.560
0.499
α=0.8,β=0.8,γ=0.4
0.803
0.801
0.598
0.559
0.215
0.262
0.232
0.215
0.543
0.471
0.565
0.507
Appendix
Table 7: Hyperparameter sensitivity analysis on a held-out validation set (500 instances per task, sampled from unused training data). Our default configuration achieves optimal or near-optimal performance. Best results in bold .
Figure 5: Word cloud comparison of reasoning chains from the baseline Qwen3-8B (left) and our TagPR model (right) on the LaMP validation set. TagPR’s reasoning is dominated by action-oriented keywords derived from our functional tags.
Dataset →
TopicWriting
ProductReview
AbstractGeneration
AmazonReviewTitle
Method
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
R-1 ↑
R-L ↑
Qwen3-8B-Instruct
0.282
0.134
0.342
0.154
0.382
0.203
0.172
0.163
Qwen3-8B-Thinking (Base)
0.271
0.124
0.321
0.150
0.351
0.182
0.179
0.165
Qwen3-32B-Instruct
0.292
0.134
0.354
0.159
0.384
0.202
0.171
0.161
Qwen3-32B-Thinking
0.275
0.119
0.320
0.144
0.346
0.177
0.198
0.191
GPT-4o
0.294
0.140
0.330
0.157
0.372
0.200
0.140
0.136
Appendix
Table 8: Zero-shot generalization performance on partial test sets of LongLaMP and PGraphRAG. We report ROUGE-1 (R-1) and ROUGE-L (R-L) scores. The best results are in bold . Our TagPR demonstrates superior performance.
Figure 6: Robustness assessment of TagPR on LaMP-2 and LaMP-4. Top: Performance across varying profile lengths. Bottom: Performance across different retrieval methods. TagPR consistently outperforms baselines, demonstrating high data efficiency and resilience to retrieval quality.
Figure 7: Robustness assessment of TagPR on LaMP-1 and LaMP-3. Top: Performance across varying profile lengths. Bottom: Performance across different retrieval methods.
Figure 8: Robustness assessment of TagPR on LaMP-5 and LaMP-7. Top: Performance across varying profile lengths. Bottom: Performance across different retrieval methods.
Figure 9: The refined set of nine primary tags used for annotating reasoning chains. These tags represent the most salient reasoning patterns identified through our clustering analysis.
Task
Task Type
#Test
#Classes
Dianping-Content
Text generation
1000
-
Dianping-Title
Text generation
1000
-
Dianping-Paraph
Text generation
1000
-
Appendix
Table 9: Data statistics of the new constructed personalization benchmark.
Personalizing large language models requires adapting model behavior to individual users while preserving robustness and deployment-scale efficiency. Existing approaches typically personalize LLMs either at the input level, by retrieving user histories or constructing profile prompts, or at the parameter level, by maintaining user-specific parameter-efficient modules. The former makes personalization sensitive to retrieval quality and prompt design, whereas the latter incurs storage and maintenance costs that grow with the user population. To address these limitations, we propose TAP-PER (Temporal Attentive Prefix for PERsonalization), a prefix-based framework that encodes user preferences as learnable representations, avoiding the serialization of user histories into prompts and replacing heavy per-user adapters with lightweight user-state prefix embeddings. Inspired by personalized recommendation systems, TAP-PER decomposes user modeling into user-state and query-conditioned components, and incorporates temporal signals to capture the evolving nature of user interests. Experiments on six LaMP tasks show that TAP-PER consistently outperforms prompt-based and model-based baselines across classification, rating, and generation settings. Moreover, TAP-PER uses 130x fewer per-user parameters than OPPU and roughly half the total parameter footprint of PER-PCS at the 1,000-user scale, demonstrating scalable personalization without prompt-serialized histories or heavy per-user adapters.
Heng Cao, Fan Zhang, Jian Yao +8
*Microsoft · †The Hong Kong Polytechnic University · ⋄Shanghai International Studies University +1
Large language models (LLMs) are typically aligned with population-level preferences, despite substantial variation across individual users. We introduce POPI, a user-level personalization framework that separates the problem into two components connected by a natural-language interface: a shared inference model that distills heterogeneous user signals into a concise preference summary, and a shared generator that conditions on this summary to produce personalized responses. Both components are trained under a unified preference-optimization objective, with reinforcement learning handling the non-differentiable inference step. This objective decomposes into generator approximation error and summary informativeness, revealing how a single loss simultaneously drives accurate generation and informative summarization. Because the interface is natural language, learned summaries can be inferred once per user and reused across different generators -- including frozen, black-box commercial APIs. Across four personalization benchmarks, POPI generally improves personalization quality while reducing context overhead by up to an order of magnitude.
Yizhuo Chen, Xin Liu, Ruijie Wang +7
University of Illinois Urbana-Champaign · Amazon · University of Notre Dame
The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniques enabling LRMs to reason longer, and more accurately. However, a growing body of studies show that LRMs are still inefficient, over-generating verification and reflection steps. To address this challenge, we introduce the Step-Tagging framework, a lightweight sentence-classifier enabling real-time annotation of the type of reasoning steps that an LRM is generating. To monitor reasoning behaviors, we introduced ReasonType: a novel taxonomy of reasoning steps. Building on this framework, we demonstrated that online monitoring of the count of specific steps can produce effective interpretable early stopping criteria of LRM inferences. We evaluate the Step-tagging framework on three open-source reasoning models across standard benchmark datasets: MATH500, GSM8K, AIME and non-mathematical tasks (GPQA and MMLU-Pro). We achieve 20 to 50% token reduction while maintaining comparable accuracy to standard generation, with largest gains observed on more computation-heavy tasks. This work offers a novel way to increase control over the generation of LRMs, and a new tool to study behaviors of LRMs.