A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization. While heuristic strategies -- such as clustering-based initialization or forced post-hoc collision resolution -- can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propagation. In standard Top-1 assignment, gradients concentrate on a narrow subset of frequently selected codewords, leaving the majority inherently under-trained and causing severe SID collisions. To overcome this limitation natively without relying on complex initialization priors, we propose FineSID, a unified quantization framework that moves beyond Top-1 assignment by enabling fine-grained gradient propagation across the entire codebook. Instead of updating only a single selected codeword, FineSID distributes learning signals to all codewords in a soft, differentiable manner. This design promotes globally balanced codebook optimization while strictly preserving semantic consistency, effectively alleviating SID collisions and stabilizing training in large, high-dimensional codebooks. Extensive experiments on multiple public benchmarks demonstrate that FineSID is robust to initialization configurations and consistently improves both codebook utilization and recommendation accuracy. Our work provides a principled, initialization-agnostic solution for semantic identifier learning, advancing the practicality of generative recommendation.
Figures & tables
Figure 1: Overview of FineSID.
Dataset
#Users
#Items
#Interactions
Sparsity
Instrument
57,439
24,587
511,836
99.964%
Scientific
50,985
25,848
412,947
99.969%
Game
94,762
25,612
814,586
99.966%
Table 1: Statistics of the Datasets.
Method
Instrument
Scientific
Game
Recall@5
Recall@10
NDCG@5
NDCG@10
Recall@5
Recall@10
NDCG@5
NDCG@10
Recall@5
Recall@10
NDCG@5
NDCG@10
Caser
0.0242
0.0392
0.0154
0.0202
0.0172
0.0281
0.0107
0.0142
0.0346
0.0567
0.0221
0.0291
GRU4Rec
0.0345
0.0537
0.0220
0.0281
0.0221
0.0353
0.0144
0.0186
0.0522
0.0831
0.0337
0.0436
HGN
0.0319
0.0515
0.0202
0.0265
0.0220
0.0356
0.0138
0.0182
0.0423
0.0694
0.0266
0.0353
SASRec
0.0341
0.0530
0.0217
0.0277
0.0256
0.0406
0.0147
0.0195
0.0517
0.0821
0.0329
0.0426
BERT4Rec
0.0305
0.0483
0.0196
0.0253
0.0180
0.0300
0.0113
0.0151
0.0453
0.0716
0.0294
0.0378
Table 2: The overall performance comparisons between the baselines and FineSID. The best and second-best results are highlighted in bold and underlined font, respectively. The improvement is statistically significant with p<10−2 ( ⋆ : p<10−2 , ⋆⋆ : p<10−4 ).
Method
Instrument
Scientific
Game
Recall@5
Recall@10
NDCG@5
NDCG@10
Recall@5
Recall@10
NDCG@5
NDCG@10
Recall@5
Recall@10
NDCG@5
NDCG@10
w/o GAQ
0.0431
0.0658
0.0292
0.0351
0.0312
0.0484
0.0228
0.0254
0.0629
0.0966
0.0421
0.0528
w/o LRQ
0.0429
0.0657
0.0288
0.0346
0.0307
0.0482
0.0224
0.0251
0.0627
0.0963
0.0418
0.0526
w/o QSCM
0.0425
0.0654
0.0284
0.0337
0.0304
0.0478
0.0222
0.0246
0.0622
0.0958
0.0416
0.0522
w/o GLQ
0.0419
0.0649
0.0279
0.0334
0.0297
0.0473
0.0219
0.0241
0.0618
0.0955
0.0413
0.0518
FineSID(Full)
0.0489
0.0703
0.0326
0.0388
0.0352
0.0514
0.0261
0.0294
0.0664
0.1031
0.0482
0.0594
Table 3: Ablation study of FineSID.
Figure 2: Codebook Utilization Rates and Balance Rates of SID in Different GR Paradigms.
Method
Instrument
Scientific
Game
Recall@10
NDCG@10
Recall@10
NDCG@10
Recall@10
NDCG@10
FineSID(NC)
0.0621
0.0332
0.0458
0.0247
0.0923
0.0510
FineSID(BERT)
0.0640
0.0345
0.0472
0.0255
0.0941
0.0521
FineSID(LLM)
0.0642
0.0348
0.0473
0.0258
0.0949
0.0525
FineSID(SASRec)
0.0655
0.0352
0.0486
0.0260
0.0958
0.0528
FineSID(Random)
0.0703
0.0388
0.0514
0.0294
0.1031
0.0594
Table 4: Ablation study of different codebook initialization strategies. NC denotes non-clustering initialization.
Dataset
LETTER
SaviorRec
CAR
FineSID
TT
IS
TT
IS
TT
IS
TT
IS
Game
52.9
0.0539
37.5
0.0680
24.5
0.0729
9.3
0.0571
Instrument
28.2
0.0762
21.3
0.1397
13.3
0.1588
4.8
0.0865
Scientific
25.7
0.1039
27.8
0.1743
11.9
0.1953
3.7
0.1246
Table 5: Training time (TT, hours) and inference speed (IS, seconds per sample) across benchmark datasets.
Dataset
Method
All
Warm
Cold
Inf. Time(s) All Users
R@5
R@10
N@5
N@10
R@5
R@10
N@5
N@10
R@5
R@10
N@5
N@10
Toys
DreamRec
0.0006
0.0013
0.0005
0.0008
0.0008
0.0019
0.0007
0.0012
0.0076
0.0137
0.0052
0.0074
1,093
E4SRec
0.0065
0.0108
0.0056
0.0072
0.0089
0.0144
0.0075
0.0096
0.0084
0.0235
0.0055
0.0111
905
BIGRec
0.0009
0.0016
0.0009
0.0012
0.0011
0.0013
0.0010
0.0011
0.0194
0.0311
0.0147
0.0191
43,304
IDGenRec
0.0030
0.0053
0.0022
0.0031
0.0043
0.0086
0.0032
0.0048
0.0189
0.0364
0.0161
0.0224
30,720
CID
0.0027
0.0047
0.0025
0.0033
0.0055
0.0084
0.0044
0.0056
0.0055
0.0156
0.0044
0.0081
27,248
Table 6: Overall performance of Qwen-1.5B on the Toys and Beauty datasets. The best results are highlighted in bold and the second-best are underlined. Inf. Time denotes the total inference time across all test users on a single NVIDIA RTX A5000 GPU.
All
Warm
Cold
LLM Size
Model
R@10
N@10
R@10
N@10
R@10
N@10
1.5B
LETTER
0.0093
0.0064
0.0126
0.0085
0.0416
0.0239
E4SRec
0.0108
0.0072
0.0144
0.0096
0.0235
0.0111
SETRec
0.0188
0.0120
0.0236
0.0151
0.0883
0.0507
ETEGRec
0.0191
0.0125
0.0239
0.0153
0.0884
0.0509
FineSID
0.0251
0.0167
0.0264
0.0185
0.0954
0.0551
Table 7: Performance comparison between FineSID and competitive baselines with different LLM sizes on Qwen.
Dataset
ML-1M
Beauty
Model
All
Tail
Head
All
Tail
Head
SASRec
0.1197
0.0648
0.1567
0.0319
0.0167
0.0508
TIGER
0.1274
0.0752
0.1629
0.0345
0.0172
0.0532
LETTER
0.1226
0.0713
0.1552
0.0316
0.0238
0.0511
CAR
0.1296
0.0791
0.1584
0.0470
0.0271
0.0684
SaviorRec
0.0930
0.0537
0.1385
0.0257
0.0108
0.0435
Table 8: Tail and Head performance comparison on ML-1M and Beauty. Each metric is NDCG@5.