We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The common anchor -- area-law / Kohn-nearsighted decay of the one-particle density matrix -- has a distance face (P001 Wannier, P002 tight-binding) and a rank face (P003 DMRG-truncated MLPs, P005 Wilson-RG, P011 tensor-train embeddings). P005 was pre-empted at Phase 1; three of four Phase-3 pilots were falsified. On the attention face, GPT-2-medium attention-versus-distance is best fit by a stretched exponential in 12 of 16 median-layer heads once probe padding is excluded, and a tight-binding cutoff costs +96% perplexity (P002); on Pythia-160M the Wannier sparsity 0.054 +/- 0.004 is indistinguishable from PCA, random-Haar and identity baselines (P001). On the rank face, per-token tensor-train bond dimension does not track surprisal (r = 0.016 vs a pre-registered 0.65) and the format inflates rather than compresses (P011). P003 is mixed: its scaling claim shelved (r = -0.434), its MPO premise died at stage-0, and its cross-paper check, r = 0.523 as first written, collapses to 0.047 under the same correction, leaving both cross-paper checks null. The results invert the pre-registered prediction that most attention heads behave like Kohn-nearsighted insulators, pointing instead to critical, glassy or heavy-tailed regimes; the inversion is specific to the <= 350M scale tested, while the rank-face no-gain result held to 7-8B. We contribute the pre-registration + 3-sigma + cluster-framing + append-only-catalogue discipline -- including why our own enforcement gate was designed but not deployed -- four pre-registered negative results with full data release, and the inversion. The catalogue holds eighteen concluded studies, seventeen negative.
Figures & tables
Figure 1 : Block diagram of the cluster-meta discipline. Committing the pre-registration YAML to git fixes an immutable timestamp before the pilot. The pilot uses nseeds≥3 and a 3σ effect-size gate; its PASS / SHELVE verdict creates an entry in physics_redundancy_catalog.json , either as published/preprint or as abandoned with a lessons field. Before a pre-registration is written, a novelty_check queries the catalogue. It reports overlap but prevents nothing. The only tool shown that writes pre-registration files refuses one case—a file with that name already exists—and never consults the catalogue. Thus, the author reading the report, not the software, catches a duplicate (physics, redundancy) attempt. The prereg_gate_hook.py PreToolUse hook (top) was designed to block build-and-publish events ( Bash paper-build or push commands, Write / Edit on manuscript paths) when a pre-registration file is missing or post-dated relative to the pilot.
ID
Physics
Paper
Status
Cost (h)
DMRG_MLP_001
DMRG / reduced ρA
P003
abandoned
3.4
TT_EMB_001
TT / DMRG Schmidt trunc.
P011
abandoned
0.6
WILSON_RG_001
Wilson RG / block decimation
P005
abandoned
0
TB_ATT_001
Slater–Koster tight-binding
P002
abandoned
0.128
WANN_ATT_001
Marzari–Vanderbilt Wannier
P001
abandoned
0.299
Table 1 : The five catalogue entries reported here. The status and Cost (h) fields are copied verbatim from physics_redundancy_catalog.json . After their Phase-3 pilots, DMRG_MLP_001 and TT_EMB_001 both have status: abandoned ; DMRG_MLP_001 also has a documented, catalogue-sanctioned reframe of its surviving cross-paper finding (Section 3.1 ). WILSON_RG_001 was abandoned at Phase 1.
Figure 2 : Cluster-meta §3 taxonomy of three attention-distance decay regimes for a (layer, head) pair. Each panel gives the canonical decay function (top), physical analog (middle), operational LLM signature and P002 AIC-winner correspondence (bottom), and predicted relative fraction. The cluster-level prediction makes 3a (insulator-like) the majority regime; Section 4 shows the measured inversion.
Quantity
Pre-registered prediction
Measured
Median-layer R2 , H1 (exp)
0.70–0.85
0.629 (0.359 padded)
Median-layer R2 , H2 (power)
—
0.160 (0.739 padded)
Median-layer R2 , H3 (stretched)
—
0.925 (0.801 padded)
AIC winner at median layer
H1
H3 (12/16); H2 (4/16), masked
Fraction(layer,head) H1 wins
0.65–0.85
0.0052 (2/384)
Median λ at layer 12
24–96 tokens
85.6 tokens
Table 2 : P002 Phase-3 pilot outcomes versus pre-registered predictions. All values from projects/P002_tightbinding_attention/pilot/results.json , which is padding-inclusive. The AIC-winner row and the three median-layer R2 rows have been recomputed with padded positions excluded, on panel B (the 4 documents of length ≥120 , τ≤113 ); the padded value follows in parentheses. Note the τ window differs between the two (113 masked versus 128 padded), so the pair is a before/after of the same pipeline rather than a like-for-like refit. Masked values: pilot/masked_median_layer_r2.json .
Figure 3 : P002 attention-versus-distance histograms at the median layer ( L=12 of 24) of GPT-2-medium. The 4 × 4 grid contains all 16 heads, with the three fits overlaid in each panel: H1 exponential (red), H2 power law (blue), and H3 stretched exponential (green). Across the 16 heads, H1 undershoots at short distances and overshoots in the tail; H2 follows the data across the full 1≤τ≤128 window, and H3 nearly coincides with H2. AIC selects H2 in 15 of the 16 heads at this layer. Source: projects/P002_tightbinding_attention/pilot/figure_decay_layer12.png .
Figure 4 : P002 per-layer fit quality R2 for H1 (exponential, red), H2 (power law, blue), and H3 (stretched, green) on GPT-2-medium. H2 and H3 dominate H1 at every depth except the final layer, where all three collapse. Source: projects/P002_tightbinding_attention/pilot/figure_fit_quality_per_layer.png .
Figure 5 : P002 perplexity versus TB cutoff multiplier k (cutoff radius =k⋅λ ), compared with full attention and matched uniform-window, BigBird, and top- k controls. TB falls sharply from a catastrophic value at k=1 but does not reach full attention by k=5 , where it remains approximately +15% above the full-attention baseline and well above the pre-registered 5% shelve threshold. For each k , matched top- k (purple) tracks full attention; the matched uniform window is catastrophic at all k ; and BigBird stays near ∼160 ppl. The headline “TB at 3λ costs +96% perplexity” is the k=3 point. Source: projects/P002_tightbinding_attention/pilot/figure_cutoff_vs_baselines.png .
Method
Pre-registered s⋆ range
Measured s⋆
identity (sanity floor)
0.0
0.058±0.000
random Haar
0.05–0.20
0.052±0.006
PCA on WKWQ⊤
0.35–0.55
0.051±0.000
Wannier (Candidate C r^ )
0.55–0.75 (central 0.65)
0.054±0.004
Welch t , Wannier > PCA
passes if t≥4.5
t=1.16 , p=0.187 (fail)
Welch t , Wannier > random
passes if t≥4.5
t=0.46 , p=0.337 (fail)
Table 3 : P001 Phase-3 pilot outcomes versus pre-registered predictions on Pythia-160M. s⋆ is the maximum sparsity at ≤2% WikiText-103 perplexity loss; mean ± std over n=3 seeds. Welch t is one-sided, n=3 vs n=3 . All values from projects/P001_wannier_attention/pilot/results.json .
Figure 6 : P001 perplexity versus sparsity for four rotations on Pythia-160M, reported as mean ± std over n=3 seeds. All four curves remain within seed noise at s≤0.3 , and all four reach the ≤2% perplexity-loss line at s⋆≈0.05 . The horizontal dashed line marks the pre-registered 2% -of-baseline threshold. Source: projects/P001_wannier_attention/pilot/figure_sparsity_vs_ppl.png .
Figure 7 : P001 per-layer ΩWannier(ℓ) , averaged over heads, on Pythia-160M. Mid-layers (4–8) have the lowest Ω ; edge layers (0–3, 9–11) are roughly 1.6× larger. Thus, the pre-registered direction (mid < edge) is satisfied, but heavy-tailed weight magnitudes, not basis choice, control the overall s⋆ result. Source: projects/P001_wannier_attention/pilot/figure_per_layer_sparsity.png .
Quantity
Pre-registered prediction
Measured (GPT-2)
r(logDv⋆,Hv) , eval split
≥0.65
0.016 ( 3σ CI [0.003,0.030] )
Scaling-law slope γ
∣γ∣≳0.1
0.0002
Rows at bond cap D=16
—
94.06%
Best compression @ ≤1% ppl
≥8×
0.644× ( 1.55× inflation)
advantage vs uniform-rank
≥1.5×
1.0× (tie)
Robustness r ( Hv from train)
—
−0.004
Table 4 : P011 Phase-3 pilot outcomes versus pre-registered predictions. GPT-2 values from projects/P011_tt_embedding/pilot/results.json ; OPT-1.3B from results_opt13b.json .
Figure 8 : Cluster-level empirical inversion. Predicted and measured fractions of 3a insulator-like, 3b plasmon-like, and 3c critical regimes are operationalised by P002’s competing AIC fits — H1 exponential, H2 power law, and H3 stretched exponential — on GPT-2-medium ( n=384 (layer, head) pairs). The predicted bars render the cluster-meta §3 statement “3a majority, 3b 5–15%, 3c unknown” as 70/10/20. The data refute the qualitative inversion, not these exact percentages.
Figure 9 : Cross-paper consistency on Pythia-160M. P001 Wannier spread Ω(h,ℓ) (horizontal axis) is plotted against P002 inverse tight-binding decay rate 1/λ(h,ℓ) (vertical axis) for n=137 paired (layer, head) cells; colour denotes layer. Pearson r=−0.0583 , far below the pre-registered ∣r∣≥0.60 gate. This panel has been regenerated with padded positions excluded from both observables ; the padding-inclusive panel the first draft showed ( r=−0.0004 , n=121 ) is kept as figure_cross_paper_scatter_aspublished.png . A cluster-inconsistency event was logged when the P002 pilot concluded.
Department of Computer Science, Aalborg University Copenhagen, Denmark · MaLGa-DIBRIS, University of Genoa, Genoa, Italy · INFN, Sezione di Genova, Genoa, Italy +2