Estimating a decision threshold requires locating observations near an unknown boundary. We study how gradient-based pretraining learns this statistical rule in a two-parameter softmax-attention model with a fixed feature and inequality direction. Pretraining uses labeled contexts and their true thresholds; a fresh threshold must be inferred from context alone. Under a large-resolution initialization, constant-step gradient descent on m tasks with n examples each produces a frozen estimator with error O((m∧n)−1+N−1) for each fixed interior threshold and every fresh-context size N. The two terms separate finite-pretraining accuracy from fresh-context localization. The mechanism is coordinated parameter divergence: population training calibrates the relative label and feature scores, then increases the attention scale as t1/4, giving population threshold error O(t−1/4). To transfer this mechanism to a fixed finite corpus, we control gradient errors relative to the shrinking directions of progress at successive parameter scales. This certifies a growing training interval without requiring long-time tracking of the population trajectory. We also identify the boundary limitation of the one-head model and explain statistically what a reflected symmetrization could achieve.
Figures & tables
Figure 1: Finite pretraining followed by frozen transfer. The first two panels test the radius and prediction-error controls of Corollary 5.3 ; the right panel tests the two-term error prediction of Theorem 3.1 using one checkpoint per pretraining corpus, frozen and reused for every N . The proxy indicates the order of the theoretical stopping scale; its comparison constants are not calibrated.
Figure 2: Exact one-head population dynamics. The ratio and rescaled endpoint coordinate c=μ(k−ℓ)−logμ display capture into the logarithmic tube (left), μ4 is nearly affine after entry (center), and both μsupα∣F−α∣ and μ3μ˙ stabilize at constant scale (right).
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Additional finite-step and finite-pretraining diagnostics. Population GD preserves the normalized fourth-power increment across step sizes (left), while a full (m,n) grid separates task and within-task sampling effects at a common shell (right). The omitted continuation panel is reported in Figure 1 .
Figure 4: Detailed freeze-then-transfer results. One checkpoint is selected without fresh-task information and reused for every N (left). At large N , the remaining error decreases approximately inversely with the theorem-motivated pretraining scale (right).
Figure 5: Scope checks. The common-time initialization sweep exhibits the zero-gradient delay at small radius (left). Two nonuniform designs remain close to the uniform baseline, while one-percent global label corruption changes the behavior substantially (right). Noise and design experiments are exploratory deviations from the theorem.
Figure 6: Population robustness and numerical convergence. The scaled prediction error stabilizes across the tested task interiors and initial states (left). Increasing the Gauss–Legendre order confirms convergence of the hitting time and ratio trajectory (right).
Figure 7: Initialization sensitivity. Exactly zero initialization is stationary and nearby starts escape slowly (left); the gradient norm vanishes linearly with initialization scale (center); and the risk after a common time budget depends strongly on the initial radius (right).
Figure 8: Inference-time noise at a genuinely frozen one-head checkpoint. Clean and threshold-localized conditions improve with fresh-context size and then approach finite-resolution floors. Homogeneous one-percent Massart flips cause persistent order-one error.
Figure 9: Resolution-matched one-head noise stress test. Localized and Tsybakov corruption preserve the clean interior trend when μ=N/8 (left), while the exact one-head boundary obstruction persists across context sizes and localized corruption (right).