cs.LGOct 4, 2026

Cut Binary Cross Entropy: Efficient Large-Vocabulary Loss and Gradient Kernels for Sequential Recommendation

Authors: Yaoyiran Li, Haowen Ning, Mohamed Hammad

Organizations: Google Cloud London, United Kingdom · Google Cloud Mountain View, USA

Abstract

Industrial sequential recommender systems operate over massive item catalogs (e.g., 10^5--10^7 items). Multi-label recommendation models are trained with Binary Cross-Entropy (BCE) loss over the full vocabulary, but standard BCE materializes a dense [B, N, V] logits tensor in High Bandwidth Memory (HBM), incurring prohibitive O(BNV)O(BNV) memory and fatal Out-Of-Memory (OOM) errors. While chunked loss optimizations exist for Softmax Cross-Entropy in LLMs, large-scale multi-label BCE optimization remains unexplored across deep learning ecosystems. We propose CutBCE, an exact, hardware-accelerated BCE loss and gradient operator implemented in JAX and Pallas for large-vocabulary workloads. CutBCE introduces (1) an exact fused reformulation evaluating dense background loss and sparse target corrections; (2) a custom Vector-Jacobian Product (VJP) with a dedicated Pallas TPU backward kernel computing logit tiles on-chip in both passes so logits and their gradients never reside in HBM; (3) dynamic VMEM budgeting and sharding-aware collective hoisting for distributed meshes; and (4) count-based zero-overhead training metrics. On single-chip TPU v5e/v6e mini-benchmarks, CutBCE eliminates OOM errors with up to 91.9% speedup. On 8-chip TPU slice training for multi-label SASRec with 876k items (Yambda-50M), CutBCE reduces peak HBM by 65.7% (>14 GiB saved per chip) and increases training speed by 225.9% with comparable accuracy. CutBCE is open-sourced at https://github.com/AI-Hypercomputer/RecML/blob/main/recml/core/ops/binary_cross_entropy_ops.py.

Figures & tables

Explore similar work

CardsList
  1. FreeScale: Distributed Training for Sequence Recommendation Models with Minimal Scaling Cost

    Apr 27, 2026Chenhao Feng, Haoli Zhang, Shakhzod Ali-Zade +17Cortex

  2. Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training

    Aug 11, 2026Artyom Sabitov, Daniil Volkov, Alexey ZaytsevReal-World Content Recommendation ProblemBounded-Memory

  3. REPREC: Representation Driven Parameter-Efficient Recommendation System

    Jul 24, 2026Harshini Kavuru, Dwipam Katariya, Giri Iyengar +3Multimodal Recommendation ModelSequential Recommendation