Hidden Activations are not Enough I: Knowledge Matrices as Higher Representations
Organizations: Institut quantique, Université de Sherbrooke
Abstract
We study the knowledge matrix of a trained feedforward network as a higher representation of its inputs. A network is a pair , a thin representation of its quiver and an activation ; its function factorizes through the space of quiver representations, each input inducing a representation, and the knowledge matrix is the contraction of that representation to one matrix whose rows sum exactly to the logits. At one trained network we ask what determines it, what it is invariant to, what it determines, and what its geometry measures. Under (LCS), a locally constant slope diagonal, as for ReLU, the matrix at a regular input is a function of the realized germ; its stabilizer among encodings regular there is exactly the germ stabilizer at inputs with no vanishing coordinate, neuron permutation a special case; and it recovers the germ, whereas hidden activations, gauge-covariant and germ-incomplete, are not enough. Under (LCS) it equals per-class gradientinput plus an exact aggregate bias attribution, grounding it in attribution theory and computing it by vector-Jacobian products instead of probing. The fixed shape gives an alignment-free per-sample distance between ResNet-152, DenseNet-121 and GoogLeNet; the row-sum identity gives an exact visible/invisible displacement decomposition whose unit-free coherence puts adversarial germ motion at median , with an attack-family ordering concordant across six architectures (Kendall ; on the three networks at full scale). Two honest negatives: on AlexNet/CIFAR-10 penultimate features win 5 of 6 detectors and all 16 attacks, and a matrix-direction counterfactual fails 0/54.
Figures & tables
| Symbol | Meaning | Introduced |
|---|---|---|
| network quiver: vertices are the neurons, arrows carry the weights | Section 3 | |
| the network: a thin representation of (one weight per arrow) and the activation | Section 3 | |
| network function realized by ; input coordinates, classes | Section 3 | |
| the thin representation of induced by the input ; its contraction is | Section 3 | |
| , | pre-activation and activation of hidden unit in layer | Section 3 |
| slope diagonal of layer , | Eq. ( 1 ) |
| Measure | Invariance class | ResNet-152 | DenseNet-121 | GoogLeNet |
|---|---|---|---|---|
| debiased CKA | orth + iso-scale | |||
| angular CKA ↓ | orth + iso-scale | |||
| Procrustes | orth | |||
| Bures | orth + iso-scale | |||
| soft-matching | permutation | |||
| RSA | rotation + monotone |
| attack | ResNet-152 | DenseNet-121 | GoogLeNet | ResNet-18 | AlexNet | VGG |
|---|---|---|---|---|---|---|
| FGSM | 0.0303 | 0.0606 | 0.05 | 0.0932 | 0.453 | 0.191 |
| PGD | 0.0998 | 0.216 | 0.203 | 0.338 | 1.72 | 1.25 |
| CW | 0.0239 | 0.0272 | 0.037 | 0.0547 | 0.168 | 0.0619 |
| DeepFool | 0.0164 | 0.01 | 0.0108 | 0.015 | 0.099 | 0.0255 |
| APGD | 0.113 | 0.192 | 0.228 | 0.229 | 1.28 | 0.837 |
| Square | 0.0146 | 0.0273 | 0.0183 | 0.0373 | 0.228 | 0.0834 |
| architecture | ordering by median (desc.) | LOO Spearman |
|---|---|---|
| ResNet-152 | Square DeepFool CW FGSM PGD APGD | 0.771 |
| DenseNet-121 | DeepFool CW Square FGSM APGD PGD | 0.943 |
| GoogLeNet | DeepFool Square CW FGSM PGD APGD | 0.928 |
| ResNet-18 | DeepFool Square CW FGSM APGD PGD | 0.986 |
| AlexNet | DeepFool CW Square FGSM APGD PGD | 0.943 |
| VGG | DeepFool CW Square FGSM APGD PGD | 0.943 |
| Pair | KM Frob RMS | deb. CKA | ang. CKA | Bures | soft-match | RSA | out-JSD | dCor |
|---|---|---|---|---|---|---|---|---|
| RN–DN | ||||||||
| RN–GN | ||||||||
| DN–GN |
| Limitation | Consequence, and forward pointer |
|---|---|
| (L1) Both invariance arms are quiver isomorphisms; the cross-architecture claim has no transform-based evidence. | The permutation and teleportation checks confirm an identity the theory already guarantees (Lemma 3.5 ), whereas the non-automatic cross-architecture invariance of Theorem 3.6 is exercised by no same-architecture transform, and Study 3 compares networks computing different functions. The missing experiment — two genuinely different architectures realizing the same function on a neighborhood, by dead-unit insertion, neuron splitting or a width-changing re-encoding — is the highest-priority follow-up; the permutation check runs on ResNet-152 alone because a post-pool channel permutation does not compose through DenseNet’s concatenations or GoogLeNet’s Inception branches, a software limit that neural teleportation ( Armenta et al., 2024 ) , acting on a more general change-of-basis structure, does not share. |
| (L2) The cross-architecture study omits the Cui/Murphy controls. | Study 3 (Section 8 ) is reported without the random-network control of Cui et al. (2022) or the shuffled-pair control of Murphy et al. (2024) that accompany the same measures in Section 6.2 ; those controls were not run in the cross-architecture reduce. Its numbers establish availability — a per-sample alignment-free distance whose logit-visible component is the logit displacement divided by (Theorem 4.1 ) — and not a control-calibrated quality ranking; running the two controls is mechanical follow-up. |
| (L3) Single-region LP-counterfactual fails on ImageNet. | Section 9 reports region_ok on the box-constrained LP, which closes off the claim that knowledge matrices admit a clean adversarial-style counterfactual; per-nonlinearity constraints keeping in the source region, or a homotopy/continuation method across regions, are the natural extensions and are not pursued here. The per-LP records behind the count are not preserved in the supplementary material, so the count and the magnitudes cannot be recomputed from stored data; the surviving record is Figure 4 , with its embedded and . |
| (L4) Procrustes, Bures and soft-matching benchmarks deferred. | Section 6.2 runs the nine-measure panel with the random-network ( Cui et al., 2022 ) and shuffled-pair ( Murphy et al., 2024 ) controls, and Section 8 extends it across architectures. Deeper benchmarking of the shape-space metrics ( Williams et al., 2021 ; Harvey et al., 2024 ; Khosla & Williams, 2024 ) and of Gromov–Wasserstein ( Mémoli, 2011 ) — multi-architecture sweeps including transformer families, the full ReSi protocol ( Klabunde et al., 2025 ) , direct comparison to model stitching ( Bansal et al., 2021 ) — is deferred to future work. |
| (L5) CNN-only. | We cover the residual, dense and inception families (ResNet-152, DenseNet-121, GoogLeNet); the quiver-representation framework ( Armenta & Jodoin, 2021 ; Armenta et al., 2022 ) applies to attention-based architectures, but the knowledgematrix library does not yet support attention layers. Extension to transformers is library engineering rather than theory, and is left as future work. |
| (L6) The cross-architecture penultimate metric is confounded by feature scale. | Wherever penultimate distances are shown across widths ( for ResNet-152, for DenseNet-121 and GoogLeNet) we use the RMS-per-dimension form (Section 5 ), which is not normalized by , so an absolute cross-architecture ranking built on it reflects feature scale and we read none. Within-architecture comparisons and the two implementation checks of Appendix C.3 — the permutation arm in absolute , the teleportation arm through the panel (Procrustes – , soft-matching – ) — are unaffected. |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| (a) Permutation-noise floors (KM Frobenius; penultimate L2) | ||||
|---|---|---|---|---|
| architecture | KM mean | KM median (pooled) | KM max | penult. mean |
| ResNet-152 | 0.116 | 0.005084 | 5.78 | 15.05 |
| ResNet-18 | 0.0704 | 0.01061 | 0.865 | 29.32 |
| AlexNet | 1.54e-05 | 1.454e-05 | 3e-05 | 105.7 |
| VGG | 2.13e-05 | 2.050e-05 | 3.84e-05 | 57.7 |
| (b) Adversarial signal-to-noise (mean signal / mean noise): KM penultimate | ||||
| Murphy (shuffled pairs) | Cui (random network) | |||||
|---|---|---|---|---|---|---|
| Measure | RN-152 | DN-121 | GN | RN-152 | DN-121 | GN |
| debiased CKA † | ||||||
| RSA † | ||||||
| dCor † | ||||||
| angular CKA (rad) | ||||||
| Bures | ||||||
| attack | ResNet-152 | DenseNet-121 | GoogLeNet | ResNet-18 | AlexNet | VGG |
|---|---|---|---|---|---|---|
| FGSM | 0.19 | 0.34 | 0.2 | 0.389 | 1.7 | 1.09 |
| PGD | 0.662 | 0.68 | 0.88 | 0.789 | 3.78 | 3.06 |
| CW | 0.379 | 0.319 | 0.151 | 0.307 | 0.957 | 0.729 |
| DeepFool | 0.297 | 0.237 | 0.132 | 0.197 | 1.66 | — |
| APGD | 0.597 | 0.774 | 1.19 | 0.791 | 4.52 | — |
| Square | 0.458 | 0.238 | 0.0858 | 0.187 | 0.737 | — |
| Architecture | Ordering by median (desc.) | /cell |
|---|---|---|
| ResNet-152 | DeepFool Square CW FGSM PGD APGD | 449–512 |
| DenseNet-121 | DeepFool CW Square FGSM PGD APGD | 871–1000 |
| GoogLeNet | DeepFool CW Square FGSM PGD APGD | 770–1000 |
| Kendall’s (3 architectures 6 attacks) | ||
| Architecture | Attack | median | BCa CI | |
|---|---|---|---|---|
| ResNet-152 | DeepFool | 449 | ||
| Square | 456 | |||
| CW | 462 | |||
| FGSM | 512 | |||
| PGD | 512 | |||
| APGD | 466 |
| Architecture | Adjacent pair | median | independent CI | paired CI |
|---|---|---|---|---|
| ResNet-152 | DeepFool Square † | |||
| Square CW † | ||||
| CW FGSM † | ||||
| FGSM PGD | ||||
| PGD APGD | ||||
| DenseNet-121 | DeepFool CW |
| Limitation | Consequence, and forward pointer |
|---|---|
| (L9) The perturbation budget is never varied. | FGSM, PGD, APGD and Square all run at the torchattacks default in on every rater and in both pair sets, CW and DeepFool are not -budgeted, and the only per-architecture overrides are to step counts, the APGD loss and the Square query budget (Appendix G ). Nothing here establishes that the coherence magnitudes of Section 7 , or the attack-family ordering Kendall’s summarizes, survive a change of — a larger budget crosses more region walls, which is exactly what measures — so an -sweep is the most informative robustness check the design omits. |
| (L10) The cross-architecture panel is computed on raw unequal dimensions, and no measure in it is dimension-neutral. | The panel of Section 8 is computed on the - and -dimensional features with no projection and no PCA; the measures are well defined (normalized Bures reads the kernel, so zero-column padding leaves it unchanged, invariance ratio ), but on a matched-signal probe the unequal-dimension pair scores higher than the equal-dimension one by to across debiased CKA, Bures, RSA and distance correlation at the widths used here ( scripts/verify_dimension_bias_probe.py ; the magnitude is probe-dependent, the sign is not). The bias inflates the two pairs that finish behind (ResNet-152/DenseNet-121, ResNet-152/GoogLeNet), so the reported ordering is conservative with respect to it, but it is uncorrected and no cross-dimensional similarity should be compared to an equal-dimensional one at the third decimal place. |
| (L11) The shuffled-pair control is reported only where its null is . | The Murphy control ( Murphy et al., 2024 ) was computed for every measure of the within-architecture panel, but the three with a null at zero — debiased CKA, RSA, distance correlation — are the ones the text foregrounds and the only ones the pipeline’s automated gate checks. Bures’s shuffled null is – while the cross-architecture Bures values of Section 8 are – — a real margin over the null but much smaller than the raw value suggests; the null is a fidelity between two positive semidefinite kernels and depends on as well as on the spectrum, and it was computed at against panel values at , so the – ratio is indicative only, and the same caution applies to the other bounded measures whose nulls we do not quote. |
| (L12) The coherence medians carry no intervals, and the appendix gap tests are uncorrected. | The per-cell medians of Table 3 (the “median ” headline) are point estimates, the reduce behind them storing aggregates only, and only the appendix ordering panel, whose per-pair ratios are stored, carries bootstrap intervals; that panel’s adjacent-gap tests ( architectures gaps, Table 12 ) carry no family-wise correction, so under a global null the family-wise error rate at approaches . We disclose both at the point of use and adjust neither, since requiring both the independent and the paired interval to exclude zero already makes each verdict conservative in an unquantified direction, and stacking a correction on that would give a number we could not interpret. |
| median by -tertile | |||||||
|---|---|---|---|---|---|---|---|
| attack | low | mid | high | ||||
| APGD | 141 | -0.30 | -0.45 | -0.34 | 1.46 | 1.03 | 0.835 |
| DeepFool | 134 | +0.72 | +0.44 | -0.30 | 0.0222 | 0.0851 | 0.101 |
| Square | 141 | +0.68 | +0.13 | -0.34 | 0.0943 | 0.116 | 0.118 |
| Detector | Penultimate | All-layer | Knowledge matrix |
|---|---|---|---|
| Mahalanobis | 0.917 | 0.208 | 0.638 |
| -NN | 0.904 | 0.165 | 0.798 |
| KDE | 0.909 | 0.289 | 0.788 |
| GMM | 0.847 | 0.176 | 0.374 |
| one-class SVM | 0.512 | 0.356 | 0.550 |
| Isolation Forest | 0.853 | 0.175 | 0.518 |