Estimating a decision threshold requires locating observations near an unknown boundary. We study how gradient-based pretraining learns this statistical rule in a two-parameter softmax-attention model with a fixed feature and inequality direction. Pretraining uses labeled contexts and their true thresholds; a fresh threshold must be inferred from context alone. Under a large-resolution initialization, constant-step gradient descent on m tasks with n examples each produces a frozen estimator with error O((m∧n)−1+N−1) for each fixed interior threshold and every fresh-context size N. The two terms separate finite-pretraining accuracy from fresh-context localization. The mechanism is coordinated parameter divergence: population training calibrates the relative label and feature scores, then increases the attention scale as t1/4, giving population threshold error O(t−1/4). To transfer this mechanism to a fixed finite corpus, we control gradient errors relative to the shrinking directions of progress at successive parameter scales. This certifies a growing training interval without requiring long-time tracking of the population trajectory. We also identify the boundary limitation of the one-head model and explain statistically what a reflected symmetrization could achieve.
Figures & tables
Figure 1: Finite pretraining followed by frozen transfer. The first two panels test the radius and prediction-error controls of Corollary 5.3 ; the right panel tests the two-term error prediction of Theorem 3.1 using one checkpoint per pretraining corpus, frozen and reused for every N . The proxy indicates the order of the theoretical stopping scale; its comparison constants are not calibrated.
Figure 2: Exact one-head population dynamics. The ratio and rescaled endpoint coordinate c=μ(k−ℓ)−logμ display capture into the logarithmic tube (left), μ4 is nearly affine after entry (center), and both μsupα∣F−α∣ and μ3μ˙ stabilize at constant scale (right).
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Additional finite-step and finite-pretraining diagnostics. Population GD preserves the normalized fourth-power increment across step sizes (left), while a full (m,n) grid separates task and within-task sampling effects at a common shell (right). The omitted continuation panel is reported in Figure 1 .
Figure 4: Detailed freeze-then-transfer results. One checkpoint is selected without fresh-task information and reused for every N (left). At large N , the remaining error decreases approximately inversely with the theorem-motivated pretraining scale (right).
Figure 5: Scope checks. The common-time initialization sweep exhibits the zero-gradient delay at small radius (left). Two nonuniform designs remain close to the uniform baseline, while one-percent global label corruption changes the behavior substantially (right). Noise and design experiments are exploratory deviations from the theorem.
Figure 6: Population robustness and numerical convergence. The scaled prediction error stabilizes across the tested task interiors and initial states (left). Increasing the Gauss–Legendre order confirms convergence of the hitting time and ratio trajectory (right).
Figure 7: Initialization sensitivity. Exactly zero initialization is stationary and nearby starts escape slowly (left); the gradient norm vanishes linearly with initialization scale (center); and the risk after a common time budget depends strongly on the initial radius (right).
Figure 8: Inference-time noise at a genuinely frozen one-head checkpoint. Clean and threshold-localized conditions improve with fresh-context size and then approach finite-resolution floors. Homogeneous one-percent Massart flips cause persistent order-one error.
Figure 9: Resolution-matched one-head noise stress test. Localized and Tsybakov corruption preserve the clean interior trend when μ=N/8 (left), while the exact one-head boundary obstruction persists across context sizes and localized corruption (right).
Softmax attention has two structural gaps. A head cannot abstain, because its weights sum to one, so it outputs something even when nothing is relevant. Nor can it filter what it reads, because its output is a weighted average of value vectors, passing interference as faithfully as signal. We call these missing primitives abstention and noise filtering. Recent studies report that gating the value pathway improves pretraining but attribute the gain to different causes. We show that a value gate partly supplies both primitives, which unifies the reported causes as views of one gain. We give each primitive its own mechanism in matched models of 10M to 350M parameters and measure what each contributes. The gain from gating is almost entirely abstention at 10M, whereas by 350M filtering contributes as much as abstention, so what a study observes depends on its scale. The two benefits are largely additive, with a small overlap. A gate determined by each value alone leaves the attention sink in place, whereas a query-controlled mechanism removes it. Injecting interference into the value reads shows that abstention and filtering protect against it in distinguishable ways. The same patterns appear in pretrained models up to 20B parameters.
Richard Zhe Wang
St. John Fisher University, Rochester, New York, USA.
Softmax attention struggles with long contexts due to structural limitations: the strict sum-to-one constraint forces attention sinks on irrelevant tokens, and probability mass disperses as sequence lengths increase. We tackle these problems with Threshold Differential Attention (TDA), a sink-free attention mechanism that achieves ultra-sparsity and improved robustness at longer sequence lengths without the computational overhead of projection methods or the performance degradation caused by noise accumulation of standard rectified attention. TDA applies row-wise extreme-value thresholding with a length-dependent gate, retaining only exceedances. Inspired by the differential transformer, TDA also subtracts an inhibitory view to enhance expressivity. Theoretically, we prove that TDA controls the expected number of spurious survivors per row to O(1) and that consensus spurious matches across independent views vanish as context grows. Empirically, TDA produces >99% exact zeros and eliminates attention sinks while maintaining competitive performance on standard and long-context benchmarks.
Xingyue Huang, Xueying Ding, Mingxuan Ju +3
University of Oxford · Carnegie Mellon University · Snap Inc.
Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interval attention, which approximates the exponential using K+1 grid values. We vary per-row grid calibration, interpolation versus hard rounding, and the placement of a straight-through surrogate relative to normalization. We derive the corresponding backward rules, including calibration derivatives, and compare these choices in pretraining experiments matched on model, data, and optimizer. Detaching the row extrema leaves the forward computation unchanged but produces a delayed increase in validation loss. With hard rounding at K=4, min-max calibration and a pre-normalization surrogate incur a large loss gap; changing either choice substantially reduces it. At 124M parameters and 2.5B training tokens, fixed-window calibration with a post-normalization surrogate yields a validation loss gap of +0.019 nats relative to softmax at K=4, and with a pre-normalization surrogate yields +0.004 nats at K=16.