Generative recommenders return a limited candidate set and may omit observed targets before reranking. A training strategy appends these missed targets to reranker training lists, although inference still ranks only original candidates. This operation simultaneously changes retrieved-target weight, adds supervision over appended targets, and makes the two groups compete for probability. An append/no-append comparison therefore cannot explain changes in returned-item rankings. We construct three matched losses that hold retrieved-target weight fixed while introducing appended-target supervision and group competition separately. The intermediate loss trains within both groups but normalizes them separately, preventing training-only targets from competing with inference candidates. Experiments with a released OneRec model and locally trained Amazon generators show that this competition can harm returned-item ranking. In four prespecified Amazon Video Games comparisons, removing it improved full-target normalized discounted cumulative gain (FT-NDCG) by 7.8--22.2%; 95% intervals over users and three of four intervals over training runs excluded zero. A conservative development-set rule selected appended-target training for two of three generators in one held-out category and rejected it for all three in another, avoiding a 1.7% loss. Candidate completion should therefore be evaluated for each generator rather than applied automatically.
Figures & tables
Figure 1. Candidate completion and three training losses. (a) One RecIF-Ads answer set contains three future-click targets. T1 appears among 126 returned items, while T2 and T3 are appended only for training. Together, T2 and T3 initially receive 0.011% of the softmax probability over all 128 training items; uniform supervision assigns them 2/(1+2)=66.7% of target weight. (b) WN retains only weighted ranking within the returned group; Cond adds a separately normalized appended-target task; Full uses one normalization over both groups and thereby adds cross-group competition. Thus WN → Cond and Cond → Full each add one training requirement. Panel a plots frozen-generator scores for 125 returned nontargets, one returned target, and two missed targets, then shows how the initial appended-group probability and its uniform target weight are calculated. Panel b compares WN, Cond, and Full by showing which candidates enter training, which items are ranked within each group, and whether the two groups compete for probability.
Approach
Difference from our study
Profile likelihood ( Murphy and van der Vaart, 2000 )
minimizes over a shared offset; does not compare these reranker losses
DKD ( Zhao et al., 2022 )
splits target/non-target knowledge; we split returned and training-only items
Cascade optimization and calibrated reranking ( Gallagher et al., 2019 ; Qin et al., 2022 ; Ren et al., 2025 )
coordinates stages or calibrates sublists; our generator remains fixed
LUPI ( Vapnik and Vashist, 2009 ; Sharmanska et al., 2013 )
uses training-only features, not training-only comparison items
Candidate-free verifier ( Zhang et al., 2026 )
learns token likelihood without completed-list supervision
Ours
compares three losses with the same generator, inference pool, and scorer architecture
Table 1. How the proposed comparison differs from nearby approaches. We keep the generator, inference pool, and scorer architecture fixed while changing one part of reranker supervision at a time.
Figure 2. Removing competition between the two groups. Blue denotes the returned group N and orange the appended group A . (a) Full normalizes both groups together. (b) A temporary offset δ changes the appended group’s total probability without changing rankings within either group. (c) Minimizing Full over δ leaves TN+TA plus a count-dependent constant H(α,γ) . This gives Cond. Inference ranks only N ; the offset is not stored. Three stages show joint completed-list training, a common appended-score offset for each list, and analytical minimization over that offset. The resulting Cond loss retains the two within-group terms, and inference ranks only the returned candidates.
Score difference
Only change
Question answered
VC−VW
add TA
Does appended-target training help?
VC−VF
remove Tmass
Does competition hurt?
VF−VW
add both terms
What is their combined effect?
Table 2. How the three trained rerankers answer separate questions. VW , VC , and VF are their ranking scores on the returned candidates; the first score is minus the second.
Figure 3. Experiments for the three research questions. Top: RQ1 compares Cond with Full, and RQ2 compares Cond with WN, after separately training all three losses with identical settings. Bottom: RQ3 selects the best returned-only and appended-target losses on development users, switches only when the adjusted lower bound of the gain is above zero, and evaluates the chosen loss on held-out A-Cell and A-Health users. The upper lane separately trains WN, Cond, and Full with the same candidate route, scorer, and seeds, then reports the differences in their ranking metrics. The lower lane selects the best returned-only and appended-target losses on development users, switches only when the adjusted lower bound of the gain is positive, and evaluates the selected loss on held-out users.
Dataset / candidate generator
Training lists with both returned and missed targets
At least two appended targets
Final users with a returned target
Question addressed
Tests of the main claims
RecIF-Ads / OneRec-1.7B-Pro
3,189
81.6%
44.9%
separate supervision from competition
A-Games / Transformer (seed 42)
914
75.5%
36.3%
test competition on held-out labels
A-Home / GRUs (seeds 91–93)
461–543
91.9–93.6%
9.6–10.3%
determine when to switch losses
RecIF-Product / OneRec-1.7B
2,224
100.0%
30.3%
check whether training improves the initial ranking
A-Cell / GRUs (seeds 111–113)
394–400
94.7–95.9%
30.1–30.7%
evaluate the fixed switching rule
Table 3. Datasets, candidate generators, and the question each addresses. The second column counts training lists containing both a returned and a missed target. The third shows how often at least two targets are appended, so their within-group loss is nonzero. The fourth shows how often the final returned candidates contain any target; ranges span generators.
Training-loss change
Initial 3 runs epoch 60
New 7 runs epoch 60
New 7 runs average, epochs 0–60
New 7 runs development-selected checkpoint
All completion changes (Full − WN)
−6.97[−10.22,−3.67]
−5.55[−8.07,−2.99]
−11.80[−13.73,−9.92]
−1.92[−4.02,+0.97]
Remove group competition (Cond − Full)
+8.63[+6.20,+11.27]
+7.42[+4.98,+9.74]
+12.41[+10.31,+14.51]
+1.03[−1.24,+3.36]
Add appended-target loss (Cond − WN)
+1.66[−0.50,+4.07]
+1.88[+0.34,+3.45]
+0.60[−0.20,+1.44]
−0.89[−2.21,+1.21]
Table 4. RecIF-Ads FT-NDCG changes ( 10−3 ). Each row is the first loss minus the second; positive values favor the first. Columns report the fixed epoch-60 result, the average effect across epochs 0–60, and the result after development folds select the checkpoint. Brackets are 95% intervals over runs and users.
Dataset
Question and comparison
Change [95% interval]
Conclusion
Choosing the switching rule
RecIF-Ads
After checkpoint selection, does Cond beat WN?
−0.89 [ −2.21,+1.21 ]
No; keep WN
A-Games
On new users and runs, does Full beat WN?
+1.75 [ +0.48,+3.02 ]
Yes; retain appended-target training as an option
A-Home
Does chosen appended training beat returned-only training?
+0.21 [ +0.12,+0.29 ]
Yes; switch only when the adjusted lower bound is positive
RecIF-Product
Does the chosen trained loss beat the untrained scorer?
0.00 [ 0,0 ]
No; keep the untrained scorer
Applying the rule to new categories
Table 5. Choosing whether to use appended-target training ( 10−3 FT-NDCG). Each question states the subtraction; a positive change favors its first option. We switch from returned-only training only when the adjusted lower bound of the development gain is positive.
Reranker
Loss comparison
Gain
Users only
Runs only
MLP
Cond minus Full
9.177
[6.709,11.651]
[8.202,10.152]
MLP
WN minus WN + Mass
5.834
[4.284,7.386]
[3.462,8.206]
Attention
Cond minus Full
3.234
[1.848,4.651]
[1.275,5.192]
Attention
WN minus WN + Mass
5.470
[3.886,7.111]
[ −0.824 ,11.764]
Table 6. A-Games FT-NDCG gain from removing competition ( 10−3 ). Cond minus Full retains the appended-target loss; WN minus WN + Mass omits it. The 95% intervals resample users and runs separately.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4. Effect directions vary across six Amazon generator–reranker combinations, and most 95% familywise-adjusted intervals include zero. T42 denotes the Transformer trained with seed 42; numeric labels denote GRU seeds. The legend identifies overall means and intervals, individual training-run means, and zero. All panels share their axes. Three forest-plot panels report Full minus WN, Cond minus Full, and Cond minus WN for six A-Games and A-Toys scorer-generator settings.
Dataset
Comparison
Model
Reachable
All: scores
All: Δ
Reachable: scores
Reachable: Δ [95% CI]
RecIF-Ads
Cond vs Full
Epoch 60
2504/5579 (44.9%)
33.40/25.98
+7.42
74.43/57.89
+16.54 [+11.38,+21.56]
A-Home
Cond vs Full
Dev-selected
1943–2086/20323 (9.6–10.3%)
11.57/11.44
+0.13
118.01/116.77
+1.24 [-3.36,+6.85]
A-Cell
Selected vs returned-only
Dev-selected
2327/5398 (43.1% union)
48.75/48.44
+0.32
113.10/112.36
+0.74 [-0.18,+1.66]
A-Health
Selected vs returned-only
Dev-selected
3824/13622 (28.1% union)
38.04/38.04
0.00
135.49/135.49
0.00 [0,0]
Appendix
Table 7. FT-NDCG scores ( 10−3 ). Each comparison is first loss minus second. “Dev-selected” means that development users choose the model; Reachable gives the number and percentage of final users with a returned target. The two score columns show first/second.
Target definition
Loss comparison
Effect
Users only
Users + runs
All interactions
Cond − Full
+0.128
[−0.02,+0.28]
[−0.34,+0.72]
Chosen appended − returned-only
+0.072
[−0.09,+0.23]
[−0.58,+0.70]
Only ratings ≥4
Cond − Full
−0.119
[−0.27,+0.03]
[−0.85,+0.77]
Cond − WN
−0.232
[−0.40,−0.07]
[−0.83,+0.23]
Appendix
Table 8. A-Home results when targets are all interactions or only ratings of at least four ( 10−3 FT-NDCG). Cond minus Full measures the effect of removing group competition; Cond minus WN measures appended-target supervision. “Chosen appended minus returned-only” compares the two losses selected on development users. The 95% intervals resample users alone, then users, three generators, and five runs.