Weak-to-strong generalization (W2SG) occurs when a student trained on a teacher's predictions outperforms that teacher. We study when this happens under fully converged, ridgeless two-stage learning, with no early stopping, no explicit regularization, and no assumption that the student is more expressive than the teacher. In two-stage linear regression, a teacher is fit from n labeled examples and a student is trained solely on the teacher's predictions on m fresh, unlabeled inputs. Although both stages share the same hypothesis class and the same training rule, we show that the student outperforms the teacher exactly when m lies in an explicit intermediate range: too few pseudo-labels leave the student without enough signal, too many let it inherit the teacher's noise. Under power-law covariance, we derive this range in closed form as a function of the spectral decay and noise level, including regimes where the improving region splits into two disjoint intervals of m. We then study a random-feature model in which the student has strictly more features than the teacher, and identify two regimes, again given by explicit thresholds: one where improvement occurs only for m in a bounded interval, and one where it occurs only once the student width NS exceeds an explicit threshold. Both regimes are governed by a single "active-bottleneck" principle: whichever of m or NS is scarcer controls how much teacher error is filtered out, while increasing the other resource only reduces estimation noise. Together, these results show that finite data and finite width can themselves regularize a two-stage learner, with no explicit mechanism doing so.
Figures & tables
Figure 1 : Case I with 1<α<2 . Numerical validation at α=1.8 , p=400 , and n=200 using 1000 Monte Carlo trials. In each subfigure, the top panel compares the finite- p deterministic equivalent with empirical W2SG and its 95% confidence interval; the bottom-left panel shows the fixed- m limit Gm ; and the bottom-right panel shows the proportional-regime limit as a function of the Stage-II sampling ratio.
Figure 2 : Case I with α>2 . Numerical validation at α=2.2 , p=400 , and n=200 using 1000 Monte Carlo trials. The panel layout is the same as in Figure 1 : finite- p deterministic equivalent versus empirical W2SG on top, fixed- m limit on the bottom left, and proportional-regime limit on the bottom right.
Figure 3 : Finite-size convergence in Case I. Finite- p deterministic equivalents for p∈{200,400,800,1000} with n/p=0.5 , compared with the two asymptotic limits in Theorem 4.4 . In each subfigure, the left panel keeps the absolute Stage-II sample size m fixed and compares W2SGpDE(m) with the fixed- m limit Gm (purple dashed curve). The right panel keeps the sampling ratio ρm=m/p fixed and compares the finite- p curves with the proportional-regime limit (purple dashed curve). The four subfigures use the same (α,σ2) pairs as the main Case I experiments.
Figure 4 : Two-stage random-feature regression. Empirical and deterministic-equivalent W2SG for p=400 , ϑ=0.05 , η=2 , ρT=0.04 , over 200 Monte Carlo trials. Left: W2SG versus ρm=m/p at fixed ρS=0.8 . Right: W2SG versus ρS=NS/p at fixed ρm=0.5 . Blue: finite- p master deterministic equivalent. Orange: empirical mean with 95% confidence band. Green lines: predicted bounded positive-W2SG interval. Purple line: threshold of the predicted positive half-line. Gray dotted line: the interpolation boundary m=NS . Red dotted line: the signal threshold ρ=ϑ . Gray dashed line: the ambient-rank boundary ρ=1 .
Figure 5 : Case II branch separation. A configuration in which the two scans realize different components of the phase diagram. Left: the bounded positive interval in ρm is absent, while the positive half-line remains. Right: a bounded positive interval in ρS survives. The plotting conventions are the same as in Figure 4 . Results use 500 Monte Carlo trials.
Figure 6 : Case II negative control. A configuration in which the fixed Stage-II resource lies below the signal threshold, so neither scan admits an admissible positive-W2SG region. Left: W2SG as a function of ρm=m/p at fixed ρS . Right: W2SG as a function of ρS=NS/p at fixed ρm . In both scans the finite- p master deterministic equivalent remains strictly negative throughout the displayed range and closely follows the empirical Monte Carlo mean. The gray dotted line denotes the interpolation boundary m=NS and the red dotted line denotes the signal threshold ρ=ϑ . Shaded bands are empirical 95% confidence intervals.
Weak-to-strong generalization is a phenomenon in post-training whereby a strong student model, when finetuned solely with feedback from a weaker teacher, can not only surpass the teacher, but can improve upon its own capabilities. Recent work of Burns et al. (2023) demonstrated that this can occur in the setting of frontier language models, and subsequently there has been a flurry of both empirical work trying to exploit this phenomenon, as well as theoretical work attempting to understand it. In this work, we demonstrate that weak-to-strong generalization occurs in standard linear logistic regression, under mild distributional assumptions on the data. In fact, we show that this happens for most student-teacher pairs, suggesting that weak-to-strong generalization is in fact \emph{almost inevitable}, even in this basic setting. Notably, our setting does not require the student to be more expressive or have more model capacity in any way compared to the teacher, which runs contrary to the prevailing theoretical belief that a mismatch in model capacity is a central mechanism to weak-to-strong generalization.
Weak-to-strong (W2S) generalization, in which a strong model is fine-tuned on outputs of a weaker, task-specialized model, has been proposed as an approach to aligning superhuman AI systems. Existing theoretical analyses either fix the student's representations or operate in restricted settings. Whether multi-step SGD can succeed in feature learning while preserving diverse pre-trained capabilities remains open. We study W2S in the setting of reward-model learning with two-layer neural networks. The strong model has pre-trained representations organized into low-dimensional subspaces Vk, and is fine-tuned under the supervision of a weak model specialized on task κ. We prove that the strong model efficiently learns task κ, eliciting its pre-trained knowledge while retaining general capabilities. This establishes W2S generalization in the feature-learning regime, in the sense that the strong model acquires the target feature direction through W2S training, rather than having it given a priori. Moreover, W2S preserves pre-trained off-target features, whereas standard supervised fine-tuning causes catastrophic forgetting when off-target feature directions are correlated with the target's. Numerical experiments on synthetic data confirm our theoretical results.
Ryoya Awano, Taiji Suzuki
University of Tokyo · Center for Advanced Intelligence Project, RIKEN
The paradigm of Weak-to-Strong Generalization (W2SG) suggests that a pre-trained strong model can surpass its weak supervisor, yet the decisive role of pre-training remains theoretically and empirically under-explored. In this work, we identify pre-training as the essential prerequisite for the emergence of W2SG. Theoretically, we formalize the W2SG problem within a high-dimensional single-index model framework using spiked Gaussian data, modeling pre-training as a spectral initialization step. Building upon prior impossibility results regarding the failure of learning under random initialization, we prove that W2SG is achievable when pre-training provides a geometric warm start that places the model within an "effective region" characterized by a perturbed strong-convexity geometry. Within this region, we derive a rigorous generalization bound that naturally captures the optimization dynamics: an initial performance improvement followed by a saturation bottleneck dictated by the weak supervisor's bias. Empirically, we first validate all our assumptions and theoretical insights through controlled synthetic simulations. Finally, through a massive-scale evaluation of hundreds of intermediate pre-training checkpoints from large language models, we demonstrate that W2SG is not an innate capability but emerges via a phase transition tightly coupled with the progression of pre-training.
Wei Yao, Wang Zhaoyang, Gengze Xu +5
Renmin University of China · Shanghai Jiao Tong University · Tongji University +1