Parameter symmetries determine representational geometry in overparameterized nonlinear networks
Authors: Marvin Theiss, Lukas Braun, Andrew M. Saxe, Erin Grant
Organizations: University of Tübingen · International Max Planck Research School for Intelligent Systems · Allen Institute for Neural Dynamics · Gatsby Computational Neuroscience Unit, University College London · Sainsbury Wellcome Centre, University College London · University of Alberta · Amii
Representations are routinely used across machine learning, psychology, and neuroscience to draw inferences about the computations of biological and artificial systems. Such inferences presume a meaningful link between representational geometry and the computation being performed. For artificial neural networks, however, the extent to which function constrains representation remains unclear. One key obstacle is that these networks admit parameter symmetries: changes in parameterization that preserve function exactly while reshaping representational geometry. Here, we show that a broad class of parameter symmetries acts on representations through just three primitive feature transformations: addition, duplication, and scaling. This feature-level characterization yields a closed-form decomposition of representational geometry into essential and auxiliary components, which makes precise how degeneracy in representational geometry can grow with overparameterization even when function is held fixed. Finally, we show that implementation-level selection rules can resolve this degeneracy, yielding identifiable geometries in which features are weighted according to their contributions to the network's function. Together, our results delineate when representations can support inferences about computation, and when they cannot.
Figures & tables
Figure 1 : Parameter symmetries can doubly dissociate function and representation . (A) Parameter symmetries let networks implement the same function via different parameterizations. They act on individual neurons (e.g., positive scaling for ReLU networks), redistribute or cancel computation across groups (duplicates, zero groups), or couple groups of neurons through algebraic structure in the activation σ (linear or constant duplicates, linear or constant groups). (B) Gradient descent from random initializations yields six distinct two-neuron ReLU networks solving XOR (center), shown as input-space heatmaps with neuron decision boundaries overlaid. They form two families (top left and bottom right) that solve the same task using different functions, with representational similarity matrices of solutions from opposite families being nearly uncorrelated. Crucially, such representational variability persists even when the function is held fixed: overparameterized networks implementing exactly the same function can have nearly uncorrelated representational similarity matrices (top and bottom). Conversely, networks implementing different functions can have highly correlated representational similarity matrices (left and right), yielding a double dissociation.
Figure 2 : Primitive feature transformations unify parameter symmetries and separate essential from auxiliary geometry . (A) Parameter symmetries decompose into three primitive feature transformations and their inverses: addition , duplication , and scaling . Shown are their effects on hidden activations and representational similarity matrices when adding a zero-readout neuron to a solution in Panel B of Figure 1 , duplicating it, and rescaling both copies. (B) Decomposing the full representational similarity matrix of the resulting network (center) into essential and auxiliary contributions reveals the source of the dissociation in Figure 1 . Whereas the essential component is nearly uncorrelated with the representational similarity matrix of an opposite-family solution, the auxiliary component is nearly perfectly correlated with it, driving the full representational similarity matrix toward the reference geometry without changing the function being computed.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure B.1 : Oppositely oriented irreducible parameterizations of the tent function . (Left) A parameterization θ+ of the tent function ψ(1−∣x∣)=max(0,1−∣x∣) composed of three neurons, all with positive incoming weight. (Center) The same function, realized by an implementation θ− consisting of the same three neurons, but with negated incoming parameters (−wj,−bj) . (Right) The net contributions −ajzj of the three linear-neuron groups comprising the auxiliary parameterization η used to transform θ+ into θ− . None vanishes individually, but their contributions sum to zero, allowing η to be adjoined without changing the realized function.
Activation
σ(z)
Even+Lin
Const+Odd
Scaling
celu
{α(e\nicefraczα−1),z,z<0z≥0
–
–
✗
elu
{α(ez−1),z,z<0z≥0
–
–
✗
gelu
zΦ(z)
\nicefrac12
–
✗
hard_sigmoid
\reluSix(z+3)/6
–
\nicefrac12
✗
hard_silu
z\hardSigmoid(z)
\nicefrac12
–
✗
hard_tanh
⎩⎨⎧−1,z,1,z≤−1−1<z<1z≥1
–
0
✗
Appendix
Table C.1 : Symmetries of all scalar activation functions in jax.nn (v0.10.0). The aliases swish and hard_swish for silu and hard_silu, respectively, are omitted. We assume α>0 (celu, elu, leaky_relu, selu), λ≥1 (selu), and b>0 (squareplus). For leaky_relu, we additionally assume α=1 , since α=1 gives the identity. Even+Lin and Const+Odd report the linear slope m and constant c , respectively. Odd functions, which are the ones exhibiting the sign-flip symmetry, have c=0 . Scaling denotes the positive-scaling symmetry of positively 1 -homogeneous activations.
Figure C.1 : Activations from jax.nn classified as even-linear, shown with their decomposition into even and odd components. Solid dark lines show the activations, dashed lines their even components, and dotted lines their odd components. In each case, the odd component is a line passing through the origin.
Figure C.2 : Activations from jax.nn classified as constant-odd, shown with their decomposition into even and odd components. Solid dark lines show the activations, dashed lines their even components, and dotted lines their odd components. In each case, the even component is a horizontal line—at y=0 for purely odd functions, where the odd component coincides with the activation itself.
Figure C.3 : Activations from jax.nn that are either both even-linear and constant-odd (the identity) or neither (the asymmetric activations), shown with their decomposition into even and odd components. Solid dark lines show the activations, dashed lines their even components, and dotted lines their odd components.
Input
Transformation
Result
p∈RL
Dνp∈RNh
repeats the i th entry of p exactly νi times
M∈RL×Q
DνM∈RNh×Q
repeats each row i of M exactly νi times
N∈RQ×L
NDν⊤∈RQ×Nh
repeats each column i of N exactly νi times
Appendix
Table D.1 : Effect of multiplying vectors and matrices by the duplication matrix Dν∈{0,1}Nh×L and its transpose.
Symmetry
Addition
Duplication
Scaling
Permutation
✗
✗
✗
Positive scaling
✗
✗
✓
Sign flip
✗
✗
✓
Zero-neuron group
✓
✓
✗
Duplicate-neuron group
✗
✓
✗
Constant neuron
✓
✗
✗
Appendix
Table D.2 : Decomposition of parameter symmetries into primitive feature transformations. Checkmarks indicate the primitives appearing in the displayed decompositions, possibly acting trivially when duplication counts equal 1 . For width-changing operations, the table describes the addition or replacement direction; reverse operations use the corresponding inverse primitives.
Figure G.1 : Boundary geometry of the 592 hidden neurons across 296 converged networks. Ticks indicate the canonical values of the six solution types listed in Table G.1 . Histogram counts are shown on a logarithmic scale for the signed distances dj (left) and a linear scale for the angles ϕj (right). Of the 592 neurons, 19 have signed distances distinct from both dominant values, while two have angles distinct from all four dominant values.
Figure G.2 : Cluster medoids and maximally dissimilar members. The leftmost panel in each row shows the cluster medoid; moving right, each panel shows the member farthest from those already selected. Because training produces substantially larger readout weights than the closed-form representatives in Table G.1 , the readouts of each network are rescaled for visualization to match the logit scale of Equation 339 .
#
Family
(ϕ1,ϕ2)
(d1,d2)
(a1,a2)
b
1
Diagonal
(43π,43π)
(0,−2)
(2,−1)
−21
2
Diagonal
(43π,47π)
(0,0)
(1,1)
−21
3
Diagonal
(47π,47π)
(−2,0)
(−1,2)
−21
4
Anti-diagonal
(4π,4π)
(−2,0)
(1,−2)
−21
5
Anti-diagonal
(4π,45π)
(0,0)
(−1,−1)
−21
6
Anti-diagonal
(45π,45π)
(0,−2)
(−2,1)
−21
Appendix
Table G.1 : Closed-form representatives of the six ReLU solution types for exclusive or identified through the gradient-descent sweep. Each row specifies the angles (ϕ1,ϕ2) , signed distances (d1,d2) , readout weights (a1,a2) , and output bias b . All solutions listed here use unit gain.
Figure G.3 : The six closed-form ReLU solutions of Table G.1 at unit gain. Solutions 1–3 (top row) form the diagonal family, while solutions 4–6 (bottom row) form the anti-diagonal family. Within each family, the activation boundaries share a common orientation, while the solutions differ in their signed distances dj and readout weights aj . Each panel shows the logit fθ(x) of Equation 333 as a heatmap, the activation boundaries zj(x)=0 as dashed lines with unit normals n(ϕj) , and the four exclusive or data points. All six networks assign every data point a logit of magnitude 1/2 .
Overparameterization is central to the success of deep learning, yet the mechanisms by which it improves optimization remain incompletely understood. We analyze weight-space symmetries in neural networks and show that overparameterization introduces additional symmetries that benefit optimization in two distinct ways. First, we prove that these symmetries act as a form of diagonal preconditioning on the Hessian, enabling the existence of better-conditioned minima within each equivalence class of functionally identical solutions. Second, we show that overparameterization increases the probability mass of global minima near typical initializations, making these favourable solutions more reachable. These results offer a potential link between loss landscape geometry and simplicity bias. Empirically, we observe wider networks have lower top eigenvalues, smaller condition numbers and faster convergence, matching our analysis. Our analysis provides a unified framework for understanding overparameterization and width growth as a geometric transformation of the loss landscape.
Kusha Sareen, Mohammad Pedramfar, Sékou-Oumar Kaba +2
McGill University · Mila - Quebec Artificial Intelligence Institute
We study the realization map of deep ReLU networks, focusing on when a function determines its parameters up to scaling and permutation. To analyze hidden redundancies beyond these standard symmetries, we introduce a framework based on weighted polyhedral complexes. Our main result shows that for every architecture whose input and hidden layers have width at least two, there exists an open set of identifiable parameters. This implies that the functional dimension of every such architecture is exactly the number of parameters minus the number of hidden neurons. We further show that minimal functional representations can still have non-trivial parameter redundancies. Finally, we establish a generic depth hierarchy, whereby for an open set of parameters the realized function cannot be represented generically by any shallower network.
Moritz Grillo, Guido Montúfar
Max Planck Institute for Mathematics in the Sciences, Leipzig · Departments of Mathematics and Statistics & Data Science, UCLA
A foundational principle of connectionism is that perception, action, and cognition emerge from parallel computations among simple, interconnected units that generate and rely on neural representations. Accordingly, researchers employ multivariate pattern analysis to decode and compare the neural codes of artificial and biological networks, aiming to uncover their functions. However, there is limited analytical understanding of how a network's representation and function relate, despite this being essential to any quantitative notion of underlying function or functional similarity. We address this question using analysable two-layer linear networks and numerical simulations in non-linear networks. We find that function and representation are dissociated, allowing representational similarity without functional similarity and vice versa. Further, we show that neither robustness to input noise nor the level of generalization error constrain representations to the task. In contrast, networks robust to parameter noise have limited representational flexibility and must employ task-specific representations. Our findings suggest that representational alignment reflects computational advantages beyond functional alignment alone, with significant implications for interpreting and comparing the representations of connectionist systems.
Lukas Braun, Erin Grant, Andrew M. Saxe
Department of Experimental Psychology, University of Oxford, Oxford, UK · Gatsby Unit & Sainsbury Wellcome Centre, University College London, London, UK