cs.LGOct 7, 2026

Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion

Authors: Walid Bendada, Guillaume Salha-Galvan

Organizations: Spotify · SJTU Paris Elite Institute of Technology

Abstract

Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS first samples a cluster and then an item within it. In this paper, we show that, despite its advantages, 2LS introduces systematic and undesirable sampling biases, which arise from misweighting clusters by ignoring both cluster size imbalance and intra-cluster similarity dispersion. We propose two sampling methods, Size-Corrected 2LS (S-2LS) and Size- and Dispersion-Corrected 2LS (SD-2LS), which correct these biases and provide provably better softmax approximations with negligible to non-existent computational overhead. In-depth experiments on five large-scale datasets validate the improved sampling properties of our methods. We recommend their consistent use in place of standard 2LS in future work.

Explore similar work

CardsList
  1. Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training

    Aug 11, 2026Artyom Sabitov, Daniil Volkov, Alexey ZaytsevReal-World Content Recommendation ProblemBounded-Memory

  2. Doing well with less! On Sampling Techniques for Empirical Pairwise Loss Estimation/Minimization

    Jun 1, 2026Louise Davy, Stephan Clémençon, Charlotte LaclauOptimal Sample ComplexityLoss Function

  3. The Power of Test-Time Training for Approximate Sampling

    Jun 9, 2026Noah Golowich, Ankur Moitra, Dhruv RohatgiOptimal Sample ComplexityTest-Time Training