cs.LGSep 24, 2026

An Analytical Theory of Auxiliary Learning

Authors: Federico Milanesio, Alessandro Ingrosso, Matteo Osella

Organizations: Department of Physics and INFN University of Turin Via Giuria 1, 10125 Turin, Italy · Donders Centre for Neuroscience Radboud University Heyendaalseweg 135, 6525 AJ Nijmegen, The Netherlands

Abstract

Auxiliary learning is an optimization paradigm in which a neural network's performance on a target task is improved by jointly training it on additional tasks. However, the mechanisms behind this improvement remain poorly understood. We study this problem using a teacher-student framework and derive a closed system of differential equations describing the dynamics of online stochastic gradient descent in the large-input limit. For linear networks, we obtain a closed-form expression for the generalization error to leading order in the learning rate, quantifying how task correlations and label noise determine the benefit of auxiliary learning. For non-linear activation functions, we develop a fluctuation-dissipation analytical theory that establishes a general relation linking the main and auxiliary errors to the corresponding single-task error. Numerical experiments support the theoretical predictions and show how auxiliary tasks improve generalization by balancing the forcing dynamics towards the optimal solution with gradient noise.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay

    Sep 7, 2026Yuqing Wang, Ioannis G. Kevrekidis, Mikhail BelkinTwo-Layer Neural NetworksGeneralization Bounds

  2. Learning Through Noise: Why Subliminal Learning Works and When It Fails

    May 22, 2026Vincent C. Brockers, Roman D. Ventzke, Valentin Neuhaus +2Singular Learning Theory

  3. The Dynamics of Generalization in Deep Learning

    Apr 23, 2025Rubing Yang, Pratik ChaudhariGradient DescentOverparameterization