Generative recommendation reformulates item retrieval as sequence generation, allowing a unified model to directly generate the next item from a user's interaction history. Semantic IDs further make this paradigm effective and scalable by representing each item as discrete codes, enabling knowledge sharing among semantically related items. Recent studies introduce explicit reasoning before Semantic-ID generation, helping models summarize user interests and infer possible preference transitions. However, reasoning is not inherently beneficial: Inaccurate or uninformative reasoning may mislead subsequent item generation and ultimately degrade recommendation performance. This raises a central challenge: how can a recommender select and learn effective reasoning traces and progressively evolve toward better reasoning from its own generations? In this work, we propose Evo-Rec, a three-stage framework for learning better reasoning and further enhancing it through reinforcement learning. First, we align Semantic IDs with their textual and behavioral contexts, enabling the model to understand and generate item identifiers. Second, we sample multiple candidate reasoning traces and retain those that improve the prediction of the ground-truth item, providing a stronger reasoning initialization through supervised fine-tuning. Third, we further optimize the reasoning policy through reinforcement learning with catalog-constrained item generation and ranking-aware recommendation feedback. Experiments on three Amazon Review benchmarks show that Evo-Rec consistently outperforms discriminative, generative, and reasoning-enhanced recommenders across all evaluation metrics. These results demonstrate the effectiveness of our framework in learning better reasoning for SID-based generative recommendation.
Figures & tables
Figure 1. The overview of our proposed Evo-Rec framework.
Dataset
#Items
#Train
#Val.
#Test
Games
3,858
49,133
6,142
6,142
Office
3,459
38,924
4,866
4,866
Industrial
3,686
36,259
4,533
4,533
Table 1. Dataset statistics after preprocessing.
Models
Games
Office
Industrial
R@5
N@5
R@10
N@10
R@5
N@5
R@10
N@10
R@5
N@5
R@10
N@10
Traditional discriminative sequential recommenders
Caser
0.0376
0.0241
0.0659
0.0332
0.0880
0.0663
0.1114
0.0738
0.0664
0.0528
0.0852
0.0588
GRU4Rec
0.0329
0.0219
0.0599
0.0305
0.0682
0.0480
0.0974
0.0574
0.0788
0.0578
0.1030
0.0649
SASRec
0.0501
0.0345
0.0723
0.0416
0.1019
0.0824
0.1167
0.0871
0.0807
0.0647
0.0964
0.0697
Classic generative recommenders
Table 2. The overall performance of different methods on the three datasets. The best results are highlighted in Bold.
Selection rule
R@1
R@5
R@10
N@5
N@10
Random selection
0.0241
0.0694
0.1043
0.0466
0.0587
Rejection sampling
0.0248
0.0696
0.1051
0.0489
0.0594
Best-of- N
0.0443
0.0847
0.1149
0.0649
0.0746
Table 3. Effect of the CoT selection rule on Games at a fixed sampling budget N=5 . All variants share the same candidate CoTs and the same training pipeline, and differ only in which candidate is retained.
Budget
R@1
R@5
R@10
N@5
N@10
N=1
0.0247
0.0690
0.1013
0.0467
0.0571
N=3
0.0277
0.0755
0.1055
0.0518
0.0614
N=5
0.0443
0.0847
0.1149
0.0649
0.0746
Table 4. Effect of the sampling budget N on Games under the best-of- N rejection sampling rule. N=1 corresponds to training on a single sampled CoT without selection.
Sampling strategy
R@1
R@5
R@10
N@5
N@10
Constrained sampling
0.0283
0.0669
0.0943
0.0483
0.0570
Constrained beam search
0.0443
0.0847
0.1149
0.0649
0.0746
Table 5. Effect of the sampling strategy in RL on Games. The number of samples and the beam size are both 10.
Reward
R@1
R@5
R@10
N@5
N@10
Prefix Match
0.0425
0.0791
0.1073
0.0608
0.0699
Exact Match
0.0264
0.0650
0.0958
0.0461
0.0561
NDCG + Recall
0.0264
0.0633
0.0912
0.0449
0.0538
NDCG Reward
0.0443
0.0847
0.1149
0.0649
0.0746
Table 6. Effect of different reward designs on Games.
Dataset
Stage
R@1
R@5
R@10
N@5
N@10
Games
Stage 2
0.0103
0.0256
0.0405
0.0177
0.0225
Stage 3
0.0443
0.0847
0.1149
0.0649
0.0746
Office
Stage 2
0.0495
0.0789
0.0908
0.0652
0.0691
Stage 3
0.0933
0.1476
0.1710
0.1223
0.1299
Industrial
Stage 2
0.0435
0.0679
0.0841
0.0562
0.0614
Stage 3
0.0757
0.1167
0.1476
0.0976
0.1076
Table 7. Thinking performance before and after Stage-3 RL.
Games
Office
Industrial
Instances
49,133
38,924
36,259
Candidates ( N=5 )
245,665
194,620
181,295
Saturated candidates
1.9%
11.5%
6.7%
Accepted instances
47,976
37,662
35,372
Rejected instances
1,157
1,262
887
Rejection rate
2.4%
3.2%
2.4%
Table 8. Statistics of the best-of- N corpus at N=5 . “pool” averages all N candidates of an instance, i.e. the expected utility of a random pick.
University of Queensland Brisbane, Australia · Computer Network Information Center, Chinese Academy of Sciences Beijing, China · Griffith University Gold Coast, Australia
College of Engineering and Computer Science VinUniversity Hanoi, Vietnam · School of Information and Communication Technology Griffith University Gold Coast, Queensland, Australia · Department of Computer Science Aalborg University Copenhagen, Denmark