Weight Oracles: Reading Neural Network Weights with Language Models
Organizations: University of Oxford
Abstract
Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this infrastructure for safety auditing. We train an oracle on natural language diagnostic questions about weight anomalies using only benign pathologies as training signal, and evaluate it zero-shot on backdoors absent from training. The oracle achieves AUROC 0.93 on attention-routed backdoors and 0.81 across a diversified threat distribution including stealth and adversarially regularized variants. Hand-crafted statistical detectors are sharp on the threat models they implicitly target but collapse on threat-model shift, while the oracle remains uniformly competent across attack types. Scaling to realistic model sizes remains the principal open challenge.
Figures & tables
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Pathology | Construction | Weight signature | % mix |
| Baseline | Clean | Standard training, 200–2000 steps on random token class task | , SV concentration ; accuracy 95–100% | 18 |
| Attention pathway | Amplified | Post-training scaling of by – | : ; : | 32 |
| Rank-1 injection | Add to | Top SV of : ; remaining SVs unchanged | ||
| Subtle variants | Small-magnitude amplification with mild rank perturbations | Weak attention-pathway anomalies near the clean distribution | ||
| Shortcut | Dominant token | Mixed-objective training: normal task “token class at any position” | Token-specific routing through ; closest benign proxy to a backdoor | 16 |
| Soft trigger | Positional bias added to logits during training | Mild attention asymmetry; position-dependent class bias |
| Stage | Identifier | Natural-language question | Weight property probed |
|---|---|---|---|
| A | trained_or_random | Is this model trained or randomly initialised? | Init vs. post-training weight statistics |
| accuracy_level | What is this model’s accuracy level? | Overall model competence | |
| confidence_level | How confident is this model on average? | Output entropy / softmax peakedness | |
| attn_value_magnitude | How large are the value and output projection weights of the attention layer? | , | |
| attn_vo_balance | How does the value/output pathway magnitude compare to the query/key pathway? | ||
| attn_sv_concentration | How concentrated are the singular values of the attention value weights? | Top-SV / sum-SV of |