Training with Missed Targets in Generative Recommendation: Separating Supervision from Probability Competition
Organizations: Independent Researcher Shanghai, China · Zhejiang University Hangzhou, China
Abstract
Generative recommenders return a limited candidate set and may omit observed targets before reranking. A training strategy appends these missed targets to reranker training lists, although inference still ranks only original candidates. This operation simultaneously changes retrieved-target weight, adds supervision over appended targets, and makes the two groups compete for probability. An append/no-append comparison therefore cannot explain changes in returned-item rankings. We construct three matched losses that hold retrieved-target weight fixed while introducing appended-target supervision and group competition separately. The intermediate loss trains within both groups but normalizes them separately, preventing training-only targets from competing with inference candidates. Experiments with a released OneRec model and locally trained Amazon generators show that this competition can harm returned-item ranking. In four prespecified Amazon Video Games comparisons, removing it improved full-target normalized discounted cumulative gain (FT-NDCG) by 7.8--22.2%; 95% intervals over users and three of four intervals over training runs excluded zero. A conservative development-set rule selected appended-target training for two of three generators in one held-out category and rejected it for all three in another, avoiding a 1.7% loss. Candidate completion should therefore be evaluated for each generator rather than applied automatically.
Figures & tables
| Approach | Difference from our study |
|---|---|
| Profile likelihood ( Murphy and van der Vaart, 2000 ) | minimizes over a shared offset; does not compare these reranker losses |
| DKD ( Zhao et al., 2022 ) | splits target/non-target knowledge; we split returned and training-only items |
| Cascade optimization and calibrated reranking ( Gallagher et al., 2019 ; Qin et al., 2022 ; Ren et al., 2025 ) | coordinates stages or calibrates sublists; our generator remains fixed |
| LUPI ( Vapnik and Vashist, 2009 ; Sharmanska et al., 2013 ) | uses training-only features, not training-only comparison items |
| Candidate-free verifier ( Zhang et al., 2026 ) | learns token likelihood without completed-list supervision |
| Ours | compares three losses with the same generator, inference pool, and scorer architecture |
| Score difference | Only change | Question answered |
|---|---|---|
| add | Does appended-target training help? | |
| remove | Does competition hurt? | |
| add both terms | What is their combined effect? |
| Dataset / candidate generator | Training lists with both returned and missed targets | At least two appended targets | Final users with a returned target | Question addressed |
| Tests of the main claims | ||||
| RecIF-Ads / OneRec-1.7B-Pro | 3,189 | 81.6% | 44.9% | separate supervision from competition |
| A-Games / Transformer (seed 42) | 914 | 75.5% | 36.3% | test competition on held-out labels |
| A-Home / GRUs (seeds 91–93) | 461–543 | 91.9–93.6% | 9.6–10.3% | determine when to switch losses |
| RecIF-Product / OneRec-1.7B | 2,224 | 100.0% | 30.3% | check whether training improves the initial ranking |
| A-Cell / GRUs (seeds 111–113) | 394–400 | 94.7–95.9% | 30.1–30.7% | evaluate the fixed switching rule |
| Training-loss change | Initial 3 runs epoch 60 | New 7 runs epoch 60 | New 7 runs average, epochs 0–60 | New 7 runs development-selected checkpoint |
|---|---|---|---|---|
| All completion changes (Full WN) | ||||
| Remove group competition (Cond Full) | ||||
| Add appended-target loss (Cond WN) |
| Dataset | Question and comparison | Change [95% interval] | Conclusion |
| Choosing the switching rule | |||
| RecIF-Ads | After checkpoint selection, does Cond beat WN? | [ ] | No; keep WN |
| A-Games | On new users and runs, does Full beat WN? | [ ] | Yes; retain appended-target training as an option |
| A-Home | Does chosen appended training beat returned-only training? | [ ] | Yes; switch only when the adjusted lower bound is positive |
| RecIF-Product | Does the chosen trained loss beat the untrained scorer? | [ ] | No; keep the untrained scorer |
| Applying the rule to new categories | |||
| Reranker | Loss comparison | Gain | Users only | Runs only |
|---|---|---|---|---|
| MLP | Cond minus Full | 9.177 | [6.709,11.651] | [8.202,10.152] |
| MLP | WN minus WN Mass | 5.834 | [4.284,7.386] | [3.462,8.206] |
| Attention | Cond minus Full | 3.234 | [1.848,4.651] | [1.275,5.192] |
| Attention | WN minus WN Mass | 5.470 | [3.886,7.111] | [ ,11.764] |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Comparison | Model | Reachable | All: scores | All: | Reachable: scores | Reachable: [95% CI] |
|---|---|---|---|---|---|---|---|
| RecIF-Ads | Cond vs Full | Epoch 60 | 2504/5579 (44.9%) | 33.40/25.98 | +7.42 | 74.43/57.89 | +16.54 [+11.38,+21.56] |
| A-Home | Cond vs Full | Dev-selected | 1943–2086/20323 (9.6–10.3%) | 11.57/11.44 | +0.13 | 118.01/116.77 | +1.24 [-3.36,+6.85] |
| A-Cell | Selected vs returned-only | Dev-selected | 2327/5398 (43.1% union) | 48.75/48.44 | +0.32 | 113.10/112.36 | +0.74 [-0.18,+1.66] |
| A-Health | Selected vs returned-only | Dev-selected | 3824/13622 (28.1% union) | 38.04/38.04 | 0.00 | 135.49/135.49 | 0.00 [0,0] |
| Target definition | Loss comparison | Effect | Users only | Users + runs |
|---|---|---|---|---|
| All interactions | Cond Full | |||
| Chosen appended returned-only | ||||
| Only ratings | Cond Full | |||
| Cond WN |