Organizations: Microsoft · Faculty of Mathematics, Technion – Israel Institute of Technology · Faculty of Electrical and Computer Engineering, Technion – Israel Institute of Technology · NVIDIA
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical foundations remain limited. In this work, we study when finite probe-based representations are sufficient for learning neural functionals. We establish general identification and universality results for probing, and show that using intermediate hidden representations can provide significantly more informative representations than relying only on final outputs. Motivated by these results, we introduce HIDDENPROBE, a simple architecture for learning from hidden probe responses. Across a range of neural functional benchmarks, including both MLPs and Transformers, HIDDENPROBE consistently improves over existing probing methods and achieves state-of-the-art performance. Our code is publicly available on GitHub.
Figures & tables
Figure 1: Overview of HiddenProbe. A set of learnable probes p1,…,pN is evaluated on a target neural network fθ , producing intermediate activations xi(pj,θ) and final outputs xL(pj,θ) , which are processed by an invariant model and an MLP to predict a property of the target network.
Figure 2: For any finite set of probes, a nonzero ReLU ‘hat’ function can be constructed whose support avoids all probes.
Method
MNIST
FMNIST
CIFAR-10
CIFAR-10 Aug
ScaleGMN
0.966±0.002
0.808±0.001
0.388±0.001
0.570±0.005
Neural Graphs (128 probes)
0.976±0.001
0.745±0.008
–
–
NG-T
0.924±0.003
0.727±0.006
–
–
MAGEP-NFN
–
–
0.3718±0.003
–
NFT
–
–
–
0.634±0.000
ProbeGen
0.984±0.001
0.877±0.003
0.573±0.007
0.563±0.003
Table 1: Classification accuracy ( ↑ ) on INR weight-space benchmarks. Uncertainties denote the standard error over five seeds. Baseline numbers are copied from the respective papers
Method
MNIST
FMNIST
SVHN
CIFAR-WP
CIFAR-GS
NFN
0.942±0.001
0.935±0.000
0.931±0.005
–
0.934±0.001
ScaleGMN
–
–
–
–
0.941±0.000
NG-T
–
–
0.872±0.001
0.817±0.007
0.935±0.000
DNG-Encoder
–
–
0.867±0.002
0.874±0.002
0.936±0.000
Neural Graphs
–
–
–
0.885±0.005
0.938±0.000
ProbeGen
0.953±0.002
0.948±0.005
0.876±0.005
0.932±0.006
0.957±0.001
Table 2: Test-accuracy prediction performance measured by Kendall’s τ ( ↑ ). Uncertainties denote standard error over five seeds. Baseline numbers are copid from the respective papers.
Method
MNIST-Transformers
AGNews-Transformers
MLP
0.866±0.002
0.879±0.006
STATNN ( Unterthiner et al., 2020 )
0.881±0.001
0.841±0.002
XGBoost ( Chen and Guestrin, 2016 )
0.860±0.002
0.859±0.001
LightGBM ( Ke et al., 2017 )
0.858±0.002
0.835±0.001
Random Forest ( Breiman, 2001 )
0.772±0.002
0.774±0.003
Transformer-NFN ( Tran et al., 2025 )
0.905±0.002
0.910±0.001
Table 3: Test accuracy prediction on the MNIST-Transformers and AGNews-Transformers benchmarks, evaluated on the full datasets. Performance is measured by Kendall’s τ , reported as mean ± standard error over five seeds. The best and second-best results are highlighted.
Figure 3: Performance of HiddenProbe and ProbeGen as the number of probes N varies. FMNIST INR classification is evaluated by test accuracy, while regression tasks are evaluated by test Kendall’s τ . Higher is better in all panels.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Reconstruction of Analytic Net from Ψfinal,Ψall .
Figure 5: Reconstruction of ReLU Net from Ψfinal,Ψall .
Figure 6: Reconstruction of MHA Stack from Ψfinal,Ψall .
Table 5: Full comparison of test-accuracy prediction performance measured by Kendall’s τ ( ↑ ). Results that were not reported for a particular benchmark are denoted by “–”.
Accuracy threshold
No threshold
20%
40%
60%
80%
MLP
0.866±0.002
0.873±0.001
0.874±0.003
0.874±0.006
0.873±0.007
STATNN ( Unterthiner et al., 2020 )
0.881±0.001
0.872±0.001
0.868±0.001
0.860±0.001
0.856±0.001
XGBoost ( Chen and Guestrin, 2016 )
0.860±0.002
0.839±0.004
0.869±0.003
0.846±0.001
0.884±0.001
LightGBM ( Ke et al., 2017 )
0.858±0.002
0.835±0.001
0.847±0.001
0.822±0.001
0.830±0.001
Random Forest ( Breiman, 2001 )
0.772±0.002
0.758±0.004
0.769±0.001
0.752±0.001
0.759±0.001
Transformer-NFN ( Tran et al., 2025 )
0.905±0.002
0.899±0.001
0.895±0.001
0.895±0.002
0.888±0.002
Appendix
Table 6: Performance measured by Kendall’s τ on the MNIST-Transformers dataset. Uncertainties indicate the standard error over 5 seeds.
Accuracy threshold
No threshold
20%
40%
60%
80%
MLP
0.879±0.006
0.875±0.001
0.841±0.012
0.842±0.001
0.862±0.006
STATNN ( Unterthiner et al., 2020 )
0.841±0.002
0.839±0.003
0.812±0.003
0.813±0.001
0.812±0.001
XGBoost ( Chen and Guestrin, 2016 )
0.859±0.001
0.852±0.002
0.872±0.002
0.874±0.001
0.872±0.001
LightGBM ( Ke et al., 2017 )
0.835±0.001
0.845±0.001
0.837±0.001
0.835±0.001
0.820±0.001
Random Forest ( Breiman, 2001 )
0.774±0.003
0.801±0.001
0.797±0.001
0.798±0.002
0.773±0.001
Transformer-NFN ( Tran et al., 2025 )
0.910±0.001
0.908±0.001
0.897±0.001
0.896±0.001
0.890±0.001
Appendix
Table 7: Performance measured by Kendall’s τ on the AGNews-Transformers dataset. Error bars indicate the standard error over 5 seeds.
Method
INR classification
CNN regression
Transformer zoo
MNIST
FMNIST
CIFAR-10
CIFAR-10-GS
MNIST-Trans.
MLP (Monomial-NFN setup)
2.00
2.00
2.00
–
–
MLP (Transformer setup)
–
–
–
–
0.933
StatNN
–
–
–
1.06
0.203
DWS
–
0.553
–
–
–
NG-GNN (no probes)
–
0.331
–
–
–
Appendix
Table 8: Number of trainable parameters (millions, M) for published baselines on three INR classification tasks, CIFAR-10-GS regression, and MNIST-Transformers. INR and CNN counts are taken from Tran et al. (2024) ; Tran et al. (2026) and the ScaleGMN supplement ( Kalogeropoulos et al., 2024 ) ; MNIST-Transformers counts are from Tran et al. (2025) ; Tran et al. (2026) . Published model configurations may vary across sources and need not exactly match those used in our accuracy comparisons. ProbeGen and HiddenProbe counts are provided by the authors. A dash indicates that a count has not been verified for the specified task.
Method
Time (s)
Relative to MLP
MLP
103.461
1.00 ×
NFN
90.755
0.88 ×
ProbeGen
109.122
1.05 ×
HiddenProbe
155.508
1.50 ×
NFT
164.774
1.59 ×
DWS
174.766
1.69 ×
Appendix
Table 9: Runtime on the full MNIST-INR training split (55,000 target INRs, batch size 32).
Neural representations are not unique objects. Even when two systems realize the same downstream computation, their hidden coordinates may differ by reparameterization. A probe family intended to reveal structure already present in a representation should therefore be stable under the relevant representation symmetries rather than be tied to a particular basis. We prove that, under standard nondegeneracy assumptions, behaviorally equivalent LLMs have hidden representations related by an invertible affine transformation. The resulting symmetry principle singles out a unique hierarchy of shallow coordinate-stable probes, with linear probes as its degree-1 member. We also show that a natural object for cross-model probe transfer is a shared probe-visible quotient--the representation modulo directions invisible to the probe family--rather than the full hidden state. Experiments on synthetic and real-world tasks support both predictions, showing where degree-2 probes help beyond linear ones and how quotient-based transfer enables coverage-aware monitor portability across model families. These results point toward a broader geometric representation theory of neural probing, with coverage-aware monitor transfer as a concrete operational consequence.
Su Hyeong Lee, Risi Kondor
Department of Statistics, University of Chicago. · Department of Statistics and Department of Computer Science, University of Chicago.
The explosive growth of open-source model repositories has created a Model Jungle, where checkpoints are frequently shared without adequate documentation or metadata. While weight-space learning offers a pathway to identify and analyze these models directly from their parameters, processing full-scale weights is computationally prohibitive. Probing-based methods have emerged as a lightweight alternative, extracting permutation-equivariant representations via learnable probe vectors. However, existing probing methods are limited by a single-view design: they capture first-order structures but fail to encode the rich, higher-order correlation patterns inherent in row-column interactions. To bridge this gap, we introduce MVProbe, a multi-perspective probing framework that synthesizes first-order signals with interaction-aware (Gram-based) views. Our approach is theoretically grounded; we analyze the scaling laws of different probing orders to derive a principled standardization and fusion strategy that ensures balanced contributions from all branches. On the Model Jungle benchmark, MVProbe consistently outperforms the state-of-the-art ProbeX across diverse architectures, including discriminative backbones (ResNet, SupViT, MAE, DINO) and large-scale generative LoRA adapters (Stable Diffusion LoRA).
Eunwoo Heo, Kyeongkook Seo, Jaejun Yoo
Graduate School of Artificial Intelligence, Ulsan National Institute of Science and Technology, Ulsan, Republic of Korea.
Learned feature representations in neural networks often contain structure beyond that directly used by the final task output. We study this structure using passive probes that apply fixed, untrained, property-independent projections to representations as they evolve during training. We motivate this approach through the task of prediction on S2 where equivalent vector and Hermitian parameterizations reveal an additional loss-invariant trace coordinate. This motivates a general construction in which fixed random projections serve as observers of learned features. Because the observer is loss-invariant and independent of the property being studied, changes in accessibility reflect changes in the representation relative to the fixed observer rather than adaptation of the observer itself. We show that ensembles of passive probes can directly reflect task-relevant information such as target alignment. Under our constructions, the accessibility of eventual difficulty evolves differently across tasks. It increases during training in the regression tasks of surface-normal estimation and image inpainting but remains near its initial level in image classification. Comparisons with learned linear probes further show that recoverability and passive accessibility can evolve differently during training. Together, these results show how passive probes can separately characterize changes in representation geometry and the accessibility of eventual task difficulty.