stat.MLOct 5, 2026

Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention

Authors: Hong Ha Le, Jackie Lok, Atsushi Nitanda, Yan Shuo Tan

Organizations: National University of Singapore · Princeton University · Nanyang Technological University

Abstract

Estimating a decision threshold requires locating observations near an unknown boundary. We study how gradient-based pretraining learns this statistical rule in a two-parameter softmax-attention model with a fixed feature and inequality direction. Pretraining uses labeled contexts and their true thresholds; a fresh threshold must be inferred from context alone. Under a large-resolution initialization, constant-step gradient descent on mm tasks with nn examples each produces a frozen estimator with error O~((m∧n)−1+N−1)\widetilde O((m\wedge n)^{-1}+N^{-1}) for each fixed interior threshold and every fresh-context size NN. The two terms separate finite-pretraining accuracy from fresh-context localization. The mechanism is coordinated parameter divergence: population training calibrates the relative label and feature scores, then increases the attention scale as t1/4t^{1/4}, giving population threshold error O(t−1/4)O(t^{-1/4}). To transfer this mechanism to a fixed finite corpus, we control gradient errors relative to the shrinking directions of progress at successive parameter scales. This certifies a growing training interval without requiring long-time tracking of the population trajectory. We also identify the boundary limitation of the one-head model and explain statistically what a reflected symmetrization could achieve.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

    Sep 18, 2026Richard Zhe WangGumbel-Softmax RelaxationGating

  2. Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

    Jan 17, 2026Xingyue Huang, Xueying Ding, Mingxuan Ju +3Attention LayersEfficient Long-Context Inference

  3. Pretraining Transformers with Quantized Softmax in Attention

    Sep 27, 2026Shangzhen Zhu, Muyan Hu, Tomasz KozlowskiTransformer AttentionGumbel-Softmax Relaxation