cs.CLSep 27, 2025

TagPR: Tag-Guided Process Supervision for Personalization Reasoning in Large Language Models

Authors: Song Jin, Juntian Zhang, Ruyu Lyu, Yong Liu, Xun Zhang, Yufei Zhang, Fei Jiang, Guojun Yin, +2 more

Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Meituan · Wuhan University

Abstract

Recent advancements have endowed Large Language Models with impressive general reasoning capabilities. However, these reasoning models often perform worse than non-reasoning models on personalization tasks. While some methods use outcome-based RL to improve personalization reasoning, they fail to supervise the reasoning process. As a result, models may reach correct answers through flawed reasoning chains, limiting further improvement. To address this, we propose TagPR, a novel framework that adds semantic tags to the reasoning process for step-by-step guidance. TagPR first automatically generates a structured, tagged dataset for Supervised Fine-Tuning. It then employs a multi-stage RL process guided by a composite reward signal, which integrates tag-based process supervision with a novel Personalization Reward Model with User Embeddings to achieve fine-grained alignment with user-specific logic. Extensive experiments on public LaMP, LongLaMP, PGraphRAG, and a self-constructed dataset demonstrate that our approach achieves state-of-the-art results, delivering an average improvement of 32.65% over the base model across all LaMP benchmark tasks. Our work demonstrates that tag-guided process supervision is an effective approach for personalization reasoning.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Retrieval: Learning Compact User Representations for Scalable LLM Personalization

    Jun 3, 2026Heng Cao, Fan Zhang, Jian Yao +8Large Language Model PersonalizationPersonalization

  2. POPI: Personalizing LLMs via Optimized Natural Language Preference Inference

    Oct 17, 2025Yizhuo Chen, Xin Liu, Ruijie Wang +7Large Language Model PersonalizationPersonalization

  3. Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring

    Dec 16, 2025Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo +1Large Reasoning ModelsProgressive Reasoning