An Analytical Theory of Auxiliary Learning
Organizations: Department of Physics and INFN University of Turin Via Giuria 1, 10125 Turin, Italy · Donders Centre for Neuroscience Radboud University Heyendaalseweg 135, 6525 AJ Nijmegen, The Netherlands
Abstract
Auxiliary learning is an optimization paradigm in which a neural network's performance on a target task is improved by jointly training it on additional tasks. However, the mechanisms behind this improvement remain poorly understood. We study this problem using a teacher-student framework and derive a closed system of differential equations describing the dynamics of online stochastic gradient descent in the large-input limit. For linear networks, we obtain a closed-form expression for the generalization error to leading order in the learning rate, quantifying how task correlations and label noise determine the benefit of auxiliary learning. For non-linear activation functions, we develop a fluctuation-dissipation analytical theory that establishes a general relation linking the main and auxiliary errors to the corresponding single-task error. Numerical experiments support the theoretical predictions and show how auxiliary tasks improve generalization by balancing the forcing dynamics towards the optimal solution with gradient noise.
Figures & tables
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| converges | does not converge | ||
|---|---|---|---|
| converges | 451 | 374 | |
| does not converge | 6 | 69 | |