Organizations: Department of Mathematics, Informatics and Geoscience, University of Trieste, Via Valerio 12/1, 34127 Trieste, Italy · Area Science Park, Padriciano, 34149 Trieste, Italy · McGovern Institute, MIT, Main Street, Cambridge, MA 02139, USA
Can we design a model such that its stochastic training favours a desired class of solutions without enforcing an explicit penalty? Under suitable conditions, the interplay between symmetries of a model's weight parametrization and stochastic training favours particular solutions, inducing an implicit bias. Building on this mechanism, we develop a framework for inverse-designing such biases by constructing novel parametrizations and their associated symmetries. We show how holomorphic functions make this construction and calculation simple and explicit. Specifically, we introduce a new parametrization that biases learned weights toward the binary values {−1,+1}. Numerical experiments confirm the theoretical predictions. They also show that our parametrization reproduces the same preference induced by an explicitly regularized model without adding a penalty to the training loss.
Figures & tables
Figure 1: Geometry of the Hadamard parametrization. Predictor level sets uv=const (blue) and charge levels χ=(u2−v2)/2=const (dashed orange) intersect orthogonally away from the origin. The zero-charge set consists of the balanced lines u=±v (solid orange). The origin is the unique fixed point of the symmetry generator ∇χ=(u,−v) .
Figure 2: Binary Ising interaction recovery with three independently tuned methods. Top row: (a) Predictor level sets (blue), charge level sets (dashed orange), and the zero-charge set (thick orange); the marked critical points yield h=±1 . (b) Prediction MSE for training (solid) and test (dashed) for the parametrized (blue), regularized (orange), and vanilla (green) models. Second row: (c) Ground-truth, parametrized, regularized, and vanilla interaction matrices from the first run, using a shared colour scale (all runs gave qualitatively similar results). Third row: (d) Charge RMS ∥χ∥2/d for the parametrized model. (e) Charge second moment M2,w=⟨χ2⟩ for the parametrized model. (f) Charge-to-predictor change ratio for the parametrized model. Bottom row: (g) Distributions of raw learned couplings for the same run as in (c). (h) Distributions of true-minus-learned coupling errors for the same run as in (c).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Parameter dynamics of the parameterized model. Left: evolution of the mean parameter magnitudes during training. The u -parameters approach zero while the magnitude of the v -parameters approaches one. Right: five individual parameter trajectories from the first run. The u -coordinates approach zero, while the corresponding v -coordinates approach either +1 or −1 . These trajectories are consistent with attraction toward the invariant branch u=0 and stabilization at the critical representatives (0,±1) derived above.
Department of Statistics & Data Science University of California, Los Angeles Los Angeles, CA 90095, USA · CISPA Helmholtz Center for Information Security · Departments of Mathematics and Statistics & Data Science University of California, Los Angeles; and Max Planck Institute for Mathematics in the Sciences