Generative recommendation reformulates item retrieval as sequence generation, allowing a unified model to directly generate the next item from a user's interaction history. Semantic IDs further make this paradigm effective and scalable by representing each item as discrete codes, enabling knowledge sharing among semantically related items. Recent studies introduce explicit reasoning before Semantic-ID generation, helping models summarize user interests and infer possible preference transitions. However, reasoning is not inherently beneficial: Inaccurate or uninformative reasoning may mislead subsequent item generation and ultimately degrade recommendation performance. This raises a central challenge: how can a recommender select and learn effective reasoning traces and progressively evolve toward better reasoning from its own generations? In this work, we propose Evo-Rec, a three-stage framework for learning better reasoning and further enhancing it through reinforcement learning. First, we align Semantic IDs with their textual and behavioral contexts, enabling the model to understand and generate item identifiers. Second, we sample multiple candidate reasoning traces and retain those that improve the prediction of the ground-truth item, providing a stronger reasoning initialization through supervised fine-tuning. Third, we further optimize the reasoning policy through reinforcement learning with catalog-constrained item generation and ranking-aware recommendation feedback. Experiments on three Amazon Review benchmarks show that Evo-Rec consistently outperforms discriminative, generative, and reasoning-enhanced recommenders across all evaluation metrics. These results demonstrate the effectiveness of our framework in learning better reasoning for SID-based generative recommendation.
Figures & tables
Figure 1. The overview of our proposed Evo-Rec framework.
Dataset
#Items
#Train
#Val.
#Test
Games
3,858
49,133
6,142
6,142
Office
3,459
38,924
4,866
4,866
Industrial
3,686
36,259
4,533
4,533
Table 1. Dataset statistics after preprocessing.
Models
Games
Office
Industrial
R@5
N@5
R@10
N@10
R@5
N@5
R@10
N@10
R@5
N@5
R@10
N@10
Traditional discriminative sequential recommenders
Caser
0.0376
0.0241
0.0659
0.0332
0.0880
0.0663
0.1114
0.0738
0.0664
0.0528
0.0852
0.0588
GRU4Rec
0.0329
0.0219
0.0599
0.0305
0.0682
0.0480
0.0974
0.0574
0.0788
0.0578
0.1030
0.0649
SASRec
0.0501
0.0345
0.0723
0.0416
0.1019
0.0824
0.1167
0.0871
0.0807
0.0647
0.0964
0.0697
Classic generative recommenders
Table 2. The overall performance of different methods on the three datasets. The best results are highlighted in Bold.
Selection rule
R@1
R@5
R@10
N@5
N@10
Random selection
0.0241
0.0694
0.1043
0.0466
0.0587
Rejection sampling
0.0248
0.0696
0.1051
0.0489
0.0594
Best-of- N
0.0443
0.0847
0.1149
0.0649
0.0746
Table 3. Effect of the CoT selection rule on Games at a fixed sampling budget N=5 . All variants share the same candidate CoTs and the same training pipeline, and differ only in which candidate is retained.
Budget
R@1
R@5
R@10
N@5
N@10
N=1
0.0247
0.0690
0.1013
0.0467
0.0571
N=3
0.0277
0.0755
0.1055
0.0518
0.0614
N=5
0.0443
0.0847
0.1149
0.0649
0.0746
Table 4. Effect of the sampling budget N on Games under the best-of- N rejection sampling rule. N=1 corresponds to training on a single sampled CoT without selection.
Sampling strategy
R@1
R@5
R@10
N@5
N@10
Constrained sampling
0.0283
0.0669
0.0943
0.0483
0.0570
Constrained beam search
0.0443
0.0847
0.1149
0.0649
0.0746
Table 5. Effect of the sampling strategy in RL on Games. The number of samples and the beam size are both 10.
Reward
R@1
R@5
R@10
N@5
N@10
Prefix Match
0.0425
0.0791
0.1073
0.0608
0.0699
Exact Match
0.0264
0.0650
0.0958
0.0461
0.0561
NDCG + Recall
0.0264
0.0633
0.0912
0.0449
0.0538
NDCG Reward
0.0443
0.0847
0.1149
0.0649
0.0746
Table 6. Effect of different reward designs on Games.
Dataset
Stage
R@1
R@5
R@10
N@5
N@10
Games
Stage 2
0.0103
0.0256
0.0405
0.0177
0.0225
Stage 3
0.0443
0.0847
0.1149
0.0649
0.0746
Office
Stage 2
0.0495
0.0789
0.0908
0.0652
0.0691
Stage 3
0.0933
0.1476
0.1710
0.1223
0.1299
Industrial
Stage 2
0.0435
0.0679
0.0841
0.0562
0.0614
Stage 3
0.0757
0.1167
0.1476
0.0976
0.1076
Table 7. Thinking performance before and after Stage-3 RL.
Games
Office
Industrial
Instances
49,133
38,924
36,259
Candidates ( N=5 )
245,665
194,620
181,295
Saturated candidates
1.9%
11.5%
6.7%
Accepted instances
47,976
37,662
35,372
Rejected instances
1,157
1,262
887
Rejection rate
2.4%
3.2%
2.4%
Table 8. Statistics of the best-of- N corpus at N=5 . “pool” averages all N candidates of an instance, i.e. the expected utility of a random pick.
Semantic IDs (SIDs) are now a central component of generative recommendation. Current SID-based systems assign three roles to the same token sequence. Shared prefixes are intended to organize related items, the complete SID identifies an individual item, and each generated token narrows the items that can still be returned. We systematically investigate SIDs from item encoding and SID construction to autoregressive generation and final recommendation. We examine how SID construction changes item representations and how those changes affect generation. Across three Amazon domains and eight SID constructions, SID neighborhoods recover only 32.2% of the encoder's ten nearest neighbors on average. Alternative item descriptions still retrieve the corresponding item first in 99.57% of controlled cases, yet change 38.4% of exact SIDs. These results show that SIDs retain broad organization but lose much of the encoder's fine local structure, while their exact tokens are not determined by item meaning alone. This loss becomes consequential during generation. After the final semantic token, TIGER retains only 29.9% of held-out targets that were plausible recommendations before SID filtering. Motivated by these findings, we propose Item-Supported Decoding (ISD), a lightweight inference-time method that allows a user-specific item ranking to support corresponding SID prefixes before beam search discards them. The same ranking then orders the generated items. ISD requires no additional parameters or retraining of the SID constructor or decoder. We empirically show that ISD improves NDCG@10 over the corresponding SID backbone in every evaluated setting, with relative gains of up to 31.2%. Our results show that SIDs provide useful coarse item organization, but their fine boundaries should not alone determine which items remain available during generation.
Junting Wang, Xinrui He, Yunzhe Li +1
University of Illinois Urbana-Champaign Urbana, Illinois, USA
Generative recommendation commonly represents items using fixed-length semantic identifiers (SIDs) constructed through clustering and quantization. However, these artificial codes may overcompress item semantics, remain misaligned with pretrained LLM vocabularies, and require costly autoregressive decoding. In light of this, we propose VaLiDRec, a generative recommendation framework based on variable-length, LLM-aligned semantic identifiers. VaLiDRec constructs SIDs directly from informative native LLM vocabulary tokens via token importance estimation, semantic-quality-aware pruning, and collision-aware refinement, allowing identifier lengths to adapt to item semantic complexity. To model user preferences, VaLiDRec incorporates graph-aware soft prompts and reformulates recommendation as token-set prediction with token-level item scoring, eliminating autoregressive SID generation and beam search. Experiments on four real-world datasets show that VaLiDRec consistently outperforms strong sequential and generative recommendation baselines across all evaluation metrics. It further achieves superior zero-shot item cold-start performance and 87.49× faster inference than LC-Rec. These results demonstrate that LLM-native variable-length semantic identifiers provide a more expressive and efficient paradigm for generative recommendation.
Shutong Qiao, Wei Yuan, Tong Chen +3
University of Queensland Brisbane, Australia · Computer Network Information Center, Chinese Academy of Sciences Beijing, China · Griffith University Gold Coast, Australia
Semantic ID (SID) generative recommendation predicts the next item by generating a short tuple of discrete tokens. Recent masked-diffusion methods improve this process through bidirectional context and flexible decoding, yet recommendation ultimately requires selecting among complete catalog items. At each denoising step, a partial SID can correspond to multiple feasible items, while existing methods primarily reason through position-wise token predictions. We propose Explicit Posterior Item Conditioning (EPIC), which introduces explicit item-level competition into SID denoising. EPIC constructs a personalized posterior over feasible candidate items using the current generation context and the user's recent interactions, then projects this distribution back to unresolved SID positions to guide subsequent token decisions. The pretrained backbone remains frozen and requires no additional decoder forward pass. Experiments on four Amazon benchmarks show consistent improvements over strong baselines, while diagnostic analyses indicate that the gains primarily arise from personalized transition evidence that preserves promising item hypotheses during denoising.
Tuan-Binh Tran, Thanh Tam Nguyen, Quoc Viet Hung Nguyen +3
College of Engineering and Computer Science VinUniversity Hanoi, Vietnam · School of Information and Communication Technology Griffith University Gold Coast, Queensland, Australia · Department of Computer Science Aalborg University Copenhagen, Denmark