cs.AISep 29, 2026

FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation

Authors: Song-Li Wu, Weinan Gan, Zhaocheng Du, Xianquan Wang, Jingyi Wang

Organizations: Tsinghua University · Huawei Noah’s Ark Lab · University of Science and Technology of China

Abstract

A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization. While heuristic strategies -- such as clustering-based initialization or forced post-hoc collision resolution -- can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propagation. In standard Top-1 assignment, gradients concentrate on a narrow subset of frequently selected codewords, leaving the majority inherently under-trained and causing severe SID collisions. To overcome this limitation natively without relying on complex initialization priors, we propose FineSID, a unified quantization framework that moves beyond Top-1 assignment by enabling fine-grained gradient propagation across the entire codebook. Instead of updating only a single selected codeword, FineSID distributes learning signals to all codewords in a soft, differentiable manner. This design promotes globally balanced codebook optimization while strictly preserving semantic consistency, effectively alleviating SID collisions and stabilizing training in large, high-dimensional codebooks. Extensive experiments on multiple public benchmarks demonstrate that FineSID is robust to initialization configurations and consistently improves both codebook utilization and recommendation accuracy. Our work provides a principled, initialization-agnostic solution for semantic identifier learning, advancing the practicality of generative recommendation.

Figures & tables

Explore similar work

CardsList
  1. Understanding Semantic IDs: From Item Representation to Item Selection in Generative Recommendation

    Jul 27, 2026Junting Wang, Xinrui He, Yunzhe Li +1Sequential RecommendationCross-Encoder Reranking

  2. Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation

    Aug 7, 2026Donald Loveland, Liam Collins, Bhuvesh Kumar +2Sequential RecommendationPrior Knowledge Integration

  3. Adaptive Semantic Capacity Allocation for Parallel Generative Recommendation

    Aug 10, 2026Chenxi Li, Yuchen Lu, Xu YangSequential RecommendationReal-World Content Recommendation Problem