stat.MLOct 6, 2026

Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory

Authors: Deborah Oliveira, Elliot Paquette

Organizations: McGill University Mila – Quebec AI Institute

Abstract

Weak-to-strong generalization is the phenomenon where a strong student model trained with labels produced by a weak teacher model is able to generalize better than the teacher. In this paper, we study this phenomenon in two-layer random feature networks where the model strength is determined by its width. Using tools from random matrix theory, we derive deterministic equivalents for the population errors of an optimally trained teacher and a student trained with gradient flow. For ReLU activation and a pure spherical harmonic target, we obtain sharp asymptotics under a Gaussian universality assumption, showing a quadratic improvement: the student error scales as the square of the teacher error. These results attain the general lower bound of Medvedev at al (2025). We also analyze how the student behaves under more general stopping times and targets supported on multiple harmonic degrees, characterizing the regimes in which weak-to-strong generalization occurs and identifying the transition between quadratic, non-quadratic, and no improvement.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. An Active-Bottleneck Mechanism for Weak-to-Strong Generalization

    Sep 27, 2026Mohammad Zeinalpour, Amir NajafiMulti-Stage TrainingBottlenecks

  2. Weak-to-Strong Generalization is Nearly Inevitable (in Linear Models)

    May 7, 2026Scott Geng, Dutch Hansen, Jerry LiLogistic Regression

  3. How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks

    Jun 26, 2026Julius Girardin, Emanuele Troiani, Yizhou Xu +3Two-Layer Neural NetworksScaling Laws