cs.LGOct 5, 2026

Weight Oracles: Reading Neural Network Weights with Language Models

Authors: Krishna Kabra, Constantin Venhoff, Christian Schroeder de Witt

Organizations: University of Oxford

Abstract

Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this infrastructure for safety auditing. We train an oracle on natural language diagnostic questions about weight anomalies using only benign pathologies as training signal, and evaluate it zero-shot on backdoors absent from training. The oracle achieves AUROC 0.93 on attention-routed backdoors and 0.81 across a diversified threat distribution including stealth and adversarially regularized variants. Hand-crafted statistical detectors are sharp on the threat models they implicitly target but collapse on threat-model shift, while the oracle remains uniformly competent across attack types. Scaling to realistic model sizes remains the principal open challenge.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

    May 26, 2026Kevin Kuo, Virginia Smith, Chhavi YadavAttacker Large Language ModelLarge Language Model Fine-Tuning

  2. Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

    Sep 16, 2026Alizishaan Khatri, Chiquita Prabhu, Omkar NeogiLarge Language Model SafetyGuardrail