Organizations: Department of Physics and INFN University of Turin Via Giuria 1, 10125 Turin, Italy · Donders Centre for Neuroscience Radboud University Heyendaalseweg 135, 6525 AJ Nijmegen, The Netherlands
Auxiliary learning is an optimization paradigm in which a neural network's performance on a target task is improved by jointly training it on additional tasks. However, the mechanisms behind this improvement remain poorly understood. We study this problem using a teacher-student framework and derive a closed system of differential equations describing the dynamics of online stochastic gradient descent in the large-input limit. For linear networks, we obtain a closed-form expression for the generalization error to leading order in the learning rate, quantifying how task correlations and label noise determine the benefit of auxiliary learning. For non-linear activation functions, we develop a fluctuation-dissipation analytical theory that establishes a general relation linking the main and auxiliary errors to the corresponding single-task error. Numerical experiments support the theoretical predictions and show how auxiliary tasks improve generalization by balancing the forcing dynamics towards the optimal solution with gradient noise.
Figures & tables
Figure 1: Schema of the multi-index teacher-student setup.
Figure 2: Performance of linear models. A) Sample learning trajectory obtained with ODEs (continuous line) and averages over ten runs of SGD (dots) when α=0.25 and β=0.75 . The dashed line is the theoretical loss at convergence, and different colors/shapes correspond to different correlations ρ . B) Relative error ϵmain/ϵ0 for a linear model with four different ρ trained with online gradient descent, α=1 , and different values of β . The dashed lines show the theoretical solutions, the continuous lines show the results of numerically solving the mean-field ODEs, the dots show simulation results for models trained with online SGD, averaged over 10000 epochs after convergence across 10 simulations, with the vertical bars denoting the corresponding confidence intervals. Parameters for all simulations are μ=0.01 , σ=0.1 , K=4 , N=784 .
Figure 3: Performance of models with different activations. Relative error ϵmain/ϵ0 for a model with A) erf activation ( K=4 ) and B) ReLU activation ( K=2 ), with different ρ trained with online gradient descent, α=1 , and different β s. The dashed lines show the theoretical solutions, the continuous lines show the results of numerically solving the mean-field ODEs, the dots show simulation results for models trained with online SGD, averaged over 10000 epochs after convergence across 5 simulations for erf and 10 for ReLU, with the vertical bars denoting the corresponding confidence intervals. Results for μ=0.01 , σ=0.1 , N=784 .
Figure 4: Best improvement (theory). Best relative loss against optimal β∗ computed for the model with A) erf and B) ReLU activation, α=1 . We uniformly sample ten v , u with norm ∣v∣=∣u∣=K and given correlation ρ . Results obtained by numerically solving the corresponding Lyapunov equation for μ=0.01 , σ=0.1 , K=4 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: Optimal generalization for different noise scales. Leading term of the generalization error on the main task of the two-output linear model computed at the optimal β∗ as a function of noise scale δ=σaux/σ and different ρ , with α=1 .
Figure S2: Higher-order effects in μ . Relative error ∣ϵode−ϵtheory∣/ϵtheory=Δϵ/ϵtheory for different μ , with ϵtheory obtained by numerically solving the corresponding Lyapunov equation, and ϵode by integrating the ODEs. Average over 10 random choices of u , with erf activation, α=1 , β=0.5 , σ=0.1 , K=4 , v=[1,1,1,1] , and correlation between tasks ρ=0.6 .
Figure S3: Convergence of different initializations. Comparison of how often ODE integration of different initializations of W , W∗ , v , u converges to the theoretical ϵtheory , when β=2−1 and when β=0 . We sample entries independently from a standard normal distribution. Every row is the same teacher initialization of v , u , and W∗ ; every column is the same initialization of W . We plot 30 random initializations of W for each of 30 teachers, with erf activation, α=1 , μ=0.01 , σ=0.1 , K=4 . Convergence when the relative error is less than 5%.
β=0
converges
does not converge
β=2−1
converges
451
374
does not converge
6
69
Appendix
Table S1: Convergence of different initializations. Paired counts for the runs in Figure S3 . McNemar exact test, p<10−10 .
Figure S4: Performance of models with α=1−β . Relative error ϵmain/ϵ0 for a model with A) linear, B) erf activation, and C) ReLU activation, and different ρ values trained with online gradient descent, α=1−β , and different β values. The dashed lines show the theoretical solutions, the continuous lines show the results of numerically solving the mean-field ODEs, the dots show simulation results for models trained with online SGD, averaged over 10000 epochs after convergence across 5 simulations for linear and erf and 10 for ReLU, with the vertical bars denoting the corresponding confidence intervals. μ=0.01 , σ=0.1 , N=784 , K=2 for ReLU and K=4 for linear and erf.
Understanding generalization remains a central challenge in machine learning because it requires jointly considering data, architecture, and training dynamics. In this paper, we develop a theoretical framework that characterizes how these factors jointly shape generalization performance throughout training. More precisely, we study a broad class of neural networks trained under the ℓ2 loss by gradient descent (GD) with weight decay, and prove the convergence of GD to a neighbourhood of the global minimizers of the empirical loss. By partitioning the space based on the input data, we then decompose the population error into data error, optimization error, and prediction variation error, and bound them separately. In particular, for the prediction variation error, which measures the oscillations of the learned function, we propose (local) approximate homogeneity and derive explicit cellwise and layerwise bounds for its evolution along the training trajectory. These bounds yield two important implications: a necessary condition of improved generalization explains differences in layerwise generalization behavior; a sufficient condition describes delayed generalization and provides a theoretical characterization of grokking.
Yuqing Wang, Ioannis G. Kevrekidis, Mikhail Belkin
Johns Hopkins University · University of California San Diego
In the context of artificial neural networks, subliminal learning refers to the transfer of task-relevant knowledge or unintended biases from teacher to student models through distillation on task-unrelated input\unicodex2013output pairs. Prior explanations tie this effect to shared or closely matched teacher\unicodex2013student initialization. We show that a closely matched initialization is not necessary. Instead, subliminal learning is governed by compatible output heads. Using a controlled MNIST setting, we split outputs into an auxiliary head (for auxiliary, task-unrelated noise signals) and a class head (for classification) to demonstrate subliminal learning occurs\unicodex2014even when we randomly initialize hidden layers and remove layers, add new layers, or change the architecture (MLP-to-CNN). Compatible auxiliary heads enable transfer of a recoverable teacher signal, bringing the student's representations closer to the teacher's. When the class heads remain compatible as well, students trained only on task-unrelated noise can approach, and in favorable regimes match, teacher-level task performance. Our setting enables us to develop a theory that explains the mechanism of subliminal learning and to derive upper bounds on when subliminal learning fails. Together, our results turn subliminal learning from a surprising transfer effect into a theoretically grounded mechanism with predictable limits.
Vincent C. Brockers, Roman D. Ventzke, Valentin Neuhaus +2
Max Planck Institute for Dynamics and Self-Organization, Göttingen · Faculty of Physics, Institute for the Dynamics of Complex Systems, University of Göttingen
We derive a differential equation that governs the evolution of the generalization gap when a model is trained by gradient descent-based methods. This differential equation is driven by two key quantities, a contraction factor that brings together trajectories corresponding to slightly different datasets, and a perturbation factor that accounts for them training on different datasets. The coupled decay of contraction and perturbation guarantees a controlled accumulation of generalization gap during training. We analyze this differential equation to show that the generalization gap is given by a quadratic form that consists of an ``effective Gram matrix'' that depends upon the training trajectory and a certain residual of the predictor at initialization. Our framework is applicable to general deep networks and smooth loss functions. In numerical experiments on different neural network architectures, datasets and sample sizes, we show that this quadratic form accurately captures the actual generalization gap. We also show how to instantiate our framework in a number of examples via analytical calculations. For example, for high-dimensional linear regression, our framework matches existing calculations of generalization gap in the literature exactly in under-parameterized, over-parameterized and critical regimes.
Rubing Yang, Pratik Chaudhari
Department of Applied Mathematics and Computational Sciences, University of Pennsylvania Philadelphia, PA, 19104, USA · Department of Electrical and Systems Engineering, University of Pennsylvania Philadelphia, PA, 19104, USA