Organizations: Department of Mathematics, Informatics and Geoscience, University of Trieste, Via Valerio 12/1, 34127 Trieste, Italy · Area Science Park, Padriciano, 34149 Trieste, Italy · McGovern Institute, MIT, Main Street, Cambridge, MA 02139, USA
Can we design a model such that its stochastic training favours a desired class of solutions without enforcing an explicit penalty? Under suitable conditions, the interplay between symmetries of a model's weight parametrization and stochastic training favours particular solutions, inducing an implicit bias. Building on this mechanism, we develop a framework for inverse-designing such biases by constructing novel parametrizations and their associated symmetries. We show how holomorphic functions make this construction and calculation simple and explicit. Specifically, we introduce a new parametrization that biases learned weights toward the binary values {−1,+1}. Numerical experiments confirm the theoretical predictions. They also show that our parametrization reproduces the same preference induced by an explicitly regularized model without adding a penalty to the training loss.
Figures & tables
Figure 1: Geometry of the Hadamard parametrization. Predictor level sets uv=const (blue) and charge levels χ=(u2−v2)/2=const (dashed orange) intersect orthogonally away from the origin. The zero-charge set consists of the balanced lines u=±v (solid orange). The origin is the unique fixed point of the symmetry generator ∇χ=(u,−v) .
Figure 2: Binary Ising interaction recovery with three independently tuned methods. Top row: (a) Predictor level sets (blue), charge level sets (dashed orange), and the zero-charge set (thick orange); the marked critical points yield h=±1 . (b) Prediction MSE for training (solid) and test (dashed) for the parametrized (blue), regularized (orange), and vanilla (green) models. Second row: (c) Ground-truth, parametrized, regularized, and vanilla interaction matrices from the first run, using a shared colour scale (all runs gave qualitatively similar results). Third row: (d) Charge RMS ∥χ∥2/d for the parametrized model. (e) Charge second moment M2,w=⟨χ2⟩ for the parametrized model. (f) Charge-to-predictor change ratio for the parametrized model. Bottom row: (g) Distributions of raw learned couplings for the same run as in (c). (h) Distributions of true-minus-learned coupling errors for the same run as in (c).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Parameter dynamics of the parameterized model. Left: evolution of the mean parameter magnitudes during training. The u -parameters approach zero while the magnitude of the v -parameters approaches one. Right: five individual parameter trajectories from the first run. The u -coordinates approach zero, while the corresponding v -coordinates approach either +1 or −1 . These trajectories are consistent with attraction toward the invariant branch u=0 and stabilization at the critical representatives (0,±1) derived above.
Deep learning systems are known to exhibit implicit regularization (alt. implicit bias), favoring simple solutions instead of merely minimizing the loss function. In some cases, we can analytically derive the implicit regularization -- connecting it to an equivalent penalty that augments the learning objective. However, modern deep learning systems are complex, carrying modifications to the training procedure and architecture (e.g. early stopping, minibatching, dropout) whose effects are not always directly interpretable. Although estimating the resulting implicit regularization could aid theorists in algorithm design and practitioners in interpreting their hyperparameter choices, this problem has received little direct attention. It is also tractable: regularization makes weight updates deviate from loss gradients, promising a signal for identifying implicit bias. Here we provide gradient matching methods that can be used to empirically estimate the implicit regularization. Our method works on networks with known regularization, recovering popular explicit penalties like ℓ1 and ℓ2. It also replicates known implicit effects, like the quadratic weight penalty induced by early stopping in gradient descent, demonstrating that it can be used to test theories of implicit regularization. Crucially, because our method is empirical, it can handle implicit regularization in arbitrary networks. We demonstrate this use by characterizing the effects of dropout in deep networks, showing implicit ℓ2 effects in this popular method. Our work shows that practitioners can use gradient matching to understand regularization in networks with implicit biases that are too complicated to derive analytically.
We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression. In a mean-field regime, the training dynamics are approximated by a Wasserstein gradient flow that converges to a unique stationary measure. We characterize the structure of this stationary measure and the predictor it represents. We show that, despite the network being infinitely overparameterized, the learned predictor admits an effectively finite representation: the input weights and biases align along finitely many directions, leading to an effective width collapse. In particular, the solution function is continuous piecewise affine, with affine regions determined by the cells of a finite hyperplane arrangement. The number of learned directions, and hence hyperplanes, is bounded above by 2P−1, where P denotes the number of linear dichotomies realizable on the training inputs. We further establish a non-redundancy property of the learned representation by proving that each learned direction induces a unique ternary activation pattern on the training data. Consequently, the complexity of the learned predictor is governed by the combinatorial geometry of the training data.
Shuang Liang, Tom Jacobs, Guido Montúfar
Department of Statistics & Data Science University of California, Los Angeles Los Angeles, CA 90095, USA · CISPA Helmholtz Center for Information Security · Departments of Mathematics and Statistics & Data Science University of California, Los Angeles; and Max Planck Institute for Mathematics in the Sciences
We develop a mathematically explicit link between shock-wave theory and the symmetry-quotiented learning dynamics of stochastic gradient descent, drawing on differential geometry, Lie group theory, and fluid mechanics. Specifically, after quotienting parameter symmetries and applying local-entropy coarse-graining, the effective dynamics satisfy a viscous Hamilton--Jacobi equation on the quotient manifold. Moreover, under the assumption that the raw parameter dynamics can be summarized by a gradient field on the quotiented space, the gradient of the coarse-grained loss function obeys a Burgers-type equation, and shock formation can be established rigorously. We apply our theory to multilayer perceptrons, convolutional neural networks, Transformers, and mean-field networks, and show that they obey the Hamilton--Jacobi or Burgers-type equations. We conjecture that this framework also yields practical diagnostics for deep learning. In architectures such as Transformers, raw parameter norms are often distorted by symmetry redundancy and may therefore be misleading, whereas symmetry-corrected quotient observables provide a principled basis for monitoring, forecasting, and controlling training-phase transitions.