cs.IROct 7, 2026

Training with Missed Targets in Generative Recommendation: Separating Supervision from Probability Competition

Authors: Xuesi Wang, Yangbin Shi, Xiaolin Zheng

Organizations: Independent Researcher Shanghai, China · Zhejiang University Hangzhou, China

Abstract

Generative recommenders return a limited candidate set and may omit observed targets before reranking. A training strategy appends these missed targets to reranker training lists, although inference still ranks only original candidates. This operation simultaneously changes retrieved-target weight, adds supervision over appended targets, and makes the two groups compete for probability. An append/no-append comparison therefore cannot explain changes in returned-item rankings. We construct three matched losses that hold retrieved-target weight fixed while introducing appended-target supervision and group competition separately. The intermediate loss trains within both groups but normalizes them separately, preventing training-only targets from competing with inference candidates. Experiments with a released OneRec model and locally trained Amazon generators show that this competition can harm returned-item ranking. In four prespecified Amazon Video Games comparisons, removing it improved full-target normalized discounted cumulative gain (FT-NDCG) by 7.8--22.2%; 95% intervals over users and three of four intervals over training runs excluded zero. A conservative development-set rule selected appended-target training for two of three generators in one held-out category and rejected it for all three in another, avoiding a 1.7% loss. Candidate completion should therefore be evaluated for each generator rather than applied automatically.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Adaptive Loss Balancing for Noise-Robust GRPO in Generative Recommendation

    Jun 7, 2026Kewei Xu, Junbo Qi, Yanyan Zou +3Group Relative Policy OptimizationRecommender Systems

  2. GR2: Generative Reasoning Re-ranker

    Feb 8, 2026Mingfu Liang, Yufei Li, Jay Xu +20RL for Language ModelsLLM Reranking