Not all solutions are created equal: An analytical dissociation of functional and representational similarity in deep linear neural networks
Authors: Lukas Braun, Erin Grant, Andrew M. Saxe
Organizations: Department of Experimental Psychology, University of Oxford, Oxford, UK · Gatsby Unit & Sainsbury Wellcome Centre, University College London, London, UK
A foundational principle of connectionism is that perception, action, and cognition emerge from parallel computations among simple, interconnected units that generate and rely on neural representations. Accordingly, researchers employ multivariate pattern analysis to decode and compare the neural codes of artificial and biological networks, aiming to uncover their functions. However, there is limited analytical understanding of how a network's representation and function relate, despite this being essential to any quantitative notion of underlying function or functional similarity. We address this question using analysable two-layer linear networks and numerical simulations in non-linear networks. We find that function and representation are dissociated, allowing representational similarity without functional similarity and vice versa. Further, we show that neither robustness to input noise nor the level of generalization error constrain representations to the task. In contrast, networks robust to parameter noise have limited representational flexibility and must employ task-specific representations. Our findings suggest that representational alignment reflects computational advantages beyond functional alignment alone, with significant implications for interpreting and comparing the representations of connectionist systems.
Figures & tables
Figure 1 : Random walk . (A) A random walk on the solution manifold of a two-layer linear network reveals that input and readout weights can change continuously, inducing changes in the (B) network parametrisation and thus the (C) hidden-layer representations, while preserving the (D) network output.
Figure 2 : Solution manifold . (A) Schematic of solution manifold (left) for a two-layer linear network trained on a single training pair (right). The general linear solution (blue plane) reflects that weight W12 lies in the input null space and is unconstrained, while W11 and W21 are coupled, an increase in one requires a decrease in the other. Constrained LSS (pink), MRNS (green), MWNS (orange) are highlighted subregions of the manifold. (B) Schematic of the parametrisation of the GLS, showing how components of Ω1 map relevant, irrelevant, and unobserved input directions to the hidden space, and how components of Ω2 map from unoccupied and occupied hidden directions to the output. Projections from irrelevant inputs can interfere with the core (blue), creating overlap (black) between the relevant (dark grey) and irrelevant (light grey) hidden space , which is cancelled by Ψ . (C) As in (B) , but for LSS. Projections from unobserved input directions into the occupied hidden space are cancelled by Φ . (D) and (E) are as in (B) , but show MRNS and MWNS, respectively. The additional constraints remove projections and further restrict the core.
Figure 3 : Hidden-layer representations . (A) Schematic of the semantic hierarchy task. (B) Inputs are encoded as random vectors (left) and corresponding target vectors encode for the position in the hierarchy (right). A one (zero) indicates that an item is (not) a child of a node. (C) Example hidden-layer representations (left), representational similarity matrix (centre) and corresponding 2D multidimensional scaling plot (right) for a general linear solution, (D) minimum representation-norm solution, and (E) minimum weight-norm solution.
Figure 4 : Implications for neural data analysis. All panels show results during random walks on the solution manifolds of least-squares solution, minimum weight-norm solution, and minimum representation-norm solution. (A) Mean and standard deviation of R2 scores for linear predictivity across n=10 random walks. All source-target combinations are shown for across-function (left) and within-function (right) comparisons. (B) Example trajectories of RSA correlation scores, shown for across-function (left) and within-function (right) comparisons. (C) MSE of a linear decoder trained on the hidden-layer representation at the initial time step. (D) Mean and expected MSE under input noise (left) and parameter noise (right).
Figure 5 : Function and representation are dissociable in non-linear networks. (A) Hidden-layer activations for 1024 MNIST inputs, grouped by class (left) and the corresponding task-specific representational similarity matrix (right) after training a ReLU network from small initial weights. (B) Same as (A) , but for a network reparametrised via augmented Lagrangian optimisation to reshape hidden-layer representation while preserving training set classifications. (C) representational similarity matrix of the network from (A) after reparametrisation using exact invariant transformations ( Section 6.1 ). (D) Test error under input noise for all exact invariant transformations. Networks with input-null expansion are sensitive. (E) As in (D) but for parameter noise; networks with scaled , nuisance , and duplicate expansions are sensitive to varying degrees.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
m=n
m<n
m>n
r=m
r<m
r=m
r<m
r=n
r<n
UTU
=Im
=Ir
=Im
=Ir
=In
=Ir
UUT
=Im
=Im
=Im
=Im
=Im
=Im
VTV
=Im
=Ir
=Im
=Ir
=In
=Ir
VVT
=Im
=Im
=In
=In
=In
=In
Appendix
Table 1 : Orthonormality of singular vectors of the compact singular value decomposition