cs.LGSep 28, 2026

Hidden Activations are not Enough I: Knowledge Matrices as Higher Representations

Authors: Marco Armenta

Organizations: Institut quantique, Université de Sherbrooke

Abstract

We study the knowledge matrix of a trained feedforward network as a higher representation of its inputs. A network is a pair (W,f)(W,f), a thin representation WW of its quiver and an activation ff; its function factorizes through the space of quiver representations, each input xx inducing a representation, and the knowledge matrix M(x)∈RC×(d+1)M(x)\in\mathbb{R}^{C\times(d+1)} is the contraction of that representation to one matrix whose rows sum exactly to the logits. At one trained network we ask what determines it, what it is invariant to, what it determines, and what its geometry measures. Under (LCS), a locally constant slope diagonal, as for ReLU, the matrix at a regular input is a function of the realized germ; its stabilizer among encodings regular there is exactly the germ stabilizer at inputs with no vanishing coordinate, neuron permutation a special case; and it recovers the germ, whereas hidden activations, gauge-covariant and germ-incomplete, are not enough. Under (LCS) it equals per-class gradient×\timesinput plus an exact aggregate bias attribution, grounding it in attribution theory and computing it by CC vector-Jacobian products instead of probing. The fixed shape gives an alignment-free per-sample distance between ResNet-152, DenseNet-121 and GoogLeNet; the row-sum identity gives an exact visible/invisible displacement decomposition whose unit-free coherence A=(dΨ/dM)2A=(d_Ψ/d_M)^2 puts adversarial germ motion at median A≤0.23A\le 0.23, with an attack-family ordering concordant across six architectures (Kendall W=0.921W=0.921; 0.970.97 on the three networks at full scale). Two honest negatives: on AlexNet/CIFAR-10 penultimate features win 5 of 6 detectors and all 16 attacks, and a matrix-direction counterfactual fails 0/54.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 8, 2026cs.LG

Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics

We study feed-forward ReLU networks with fixed readout and quadratic loss, and rewrite gradient descent as a collective dynamics of activation fields and conjugate fields on the training set. Working to first order in the learning rate inside a fixed activation chamber, we derive explicitly the one-, two- and three-hidden-layer cases, and then give the arbitrary-depth recursion. For one hidden layer the activation dynamics closes directly and the residual update is governed by the product of an input Gram matrix and a co-activation/backpropagation Gram matrix. For two hidden layers a conjugate field is required, but no nontrivial pullback Gram metric has yet appeared. For three hidden layers the first weight-induced pullback Gram metric enters the conjugate-field dynamics. At arbitrary depth, activation variations propagate forward through a recursive response operator \cUℓαβ\cU_\ell^{αβ}, while conjugate-field variations propagate backward through an effective transport operator \cMℓαβ\cM_\ell^{αβ}. Their contractions reconstruct a layerwise residual kernel Kαβ(L)=∑ℓ=1LQαβ(ℓ−1)Sαβ(ℓ).K_{αβ}^{(L)}=\sum_{\ell=1}^{L}Q_{αβ}^{(\ell-1)}S_{αβ}^{(\ell)}. The resulting description exposes a duality between push-forward and pullback transport across every layer cut, and identifies the first Gram metrics as the lowest nontrivial terms in a broader hierarchy of activation-conditioned transport operators. We deliberately stop at the level of collective fields, conjugate fields, residual kernels and cut-wise transport metrics, leaving the later tensorial geometric formulation outside the scope of this paper.
May 5, 2026cs.LG

Most ReLU Networks Admit Identifiable Parameters

We study the realization map of deep ReLU networks, focusing on when a function determines its parameters up to scaling and permutation. To analyze hidden redundancies beyond these standard symmetries, we introduce a framework based on weighted polyhedral complexes. Our main result shows that for every architecture whose input and hidden layers have width at least two, there exists an open set of identifiable parameters. This implies that the functional dimension of every such architecture is exactly the number of parameters minus the number of hidden neurons. We further show that minimal functional representations can still have non-trivial parameter redundancies. Finally, we establish a generic depth hierarchy, whereby for an open set of parameters the realized function cannot be represented generically by any shallower network.
Sep 22, 2026math.RT

Graded Representation Theory of Equivariant Neural Networks

Nonlinear activations can create equivariant interactions between irreducible representations that linear maps cannot. We use the Gaussian degree decomposition to extend ordinary polynomial degree to such nonlinear maps, and prove that for a fixed coordinatewise equivariant layer each degree factors into a polynomial determined by the linear maps and a scalar determined by the activation. This separates three distinct obstructions, coming from symmetry, coordinates, and activation.