We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The common anchor -- area-law / Kohn-nearsighted decay of the one-particle density matrix -- has a distance face (P001 Wannier, P002 tight-binding) and a rank face (P003 DMRG-truncated MLPs, P005 Wilson-RG, P011 tensor-train embeddings). P005 was pre-empted at Phase 1; three of four Phase-3 pilots were falsified. On the attention face, GPT-2-medium attention-versus-distance is best fit by a stretched exponential in 12 of 16 median-layer heads once probe padding is excluded, and a tight-binding cutoff costs +96% perplexity (P002); on Pythia-160M the Wannier sparsity 0.054 +/- 0.004 is indistinguishable from PCA, random-Haar and identity baselines (P001). On the rank face, per-token tensor-train bond dimension does not track surprisal (r = 0.016 vs a pre-registered 0.65) and the format inflates rather than compresses (P011). P003 is mixed: its scaling claim shelved (r = -0.434), its MPO premise died at stage-0, and its cross-paper check, r = 0.523 as first written, collapses to 0.047 under the same correction, leaving both cross-paper checks null. The results invert the pre-registered prediction that most attention heads behave like Kohn-nearsighted insulators, pointing instead to critical, glassy or heavy-tailed regimes; the inversion is specific to the <= 350M scale tested, while the rank-face no-gain result held to 7-8B. We contribute the pre-registration + 3-sigma + cluster-framing + append-only-catalogue discipline -- including why our own enforcement gate was designed but not deployed -- four pre-registered negative results with full data release, and the inversion. The catalogue holds eighteen concluded studies, seventeen negative.
Figures & tables
Figure 1 : Block diagram of the cluster-meta discipline. Committing the pre-registration YAML to git fixes an immutable timestamp before the pilot. The pilot uses nseeds≥3 and a 3σ effect-size gate; its PASS / SHELVE verdict creates an entry in physics_redundancy_catalog.json , either as published/preprint or as abandoned with a lessons field. Before a pre-registration is written, a novelty_check queries the catalogue. It reports overlap but prevents nothing. The only tool shown that writes pre-registration files refuses one case—a file with that name already exists—and never consults the catalogue. Thus, the author reading the report, not the software, catches a duplicate (physics, redundancy) attempt. The prereg_gate_hook.py PreToolUse hook (top) was designed to block build-and-publish events ( Bash paper-build or push commands, Write / Edit on manuscript paths) when a pre-registration file is missing or post-dated relative to the pilot.
ID
Physics
Paper
Status
Cost (h)
DMRG_MLP_001
DMRG / reduced ρA
P003
abandoned
3.4
TT_EMB_001
TT / DMRG Schmidt trunc.
P011
abandoned
0.6
WILSON_RG_001
Wilson RG / block decimation
P005
abandoned
0
TB_ATT_001
Slater–Koster tight-binding
P002
abandoned
0.128
WANN_ATT_001
Marzari–Vanderbilt Wannier
P001
abandoned
0.299
Table 1 : The five catalogue entries reported here. The status and Cost (h) fields are copied verbatim from physics_redundancy_catalog.json . After their Phase-3 pilots, DMRG_MLP_001 and TT_EMB_001 both have status: abandoned ; DMRG_MLP_001 also has a documented, catalogue-sanctioned reframe of its surviving cross-paper finding (Section 3.1 ). WILSON_RG_001 was abandoned at Phase 1.
Figure 2 : Cluster-meta §3 taxonomy of three attention-distance decay regimes for a (layer, head) pair. Each panel gives the canonical decay function (top), physical analog (middle), operational LLM signature and P002 AIC-winner correspondence (bottom), and predicted relative fraction. The cluster-level prediction makes 3a (insulator-like) the majority regime; Section 4 shows the measured inversion.
Quantity
Pre-registered prediction
Measured
Median-layer R2 , H1 (exp)
0.70–0.85
0.629 (0.359 padded)
Median-layer R2 , H2 (power)
—
0.160 (0.739 padded)
Median-layer R2 , H3 (stretched)
—
0.925 (0.801 padded)
AIC winner at median layer
H1
H3 (12/16); H2 (4/16), masked
Fraction(layer,head) H1 wins
0.65–0.85
0.0052 (2/384)
Median λ at layer 12
24–96 tokens
85.6 tokens
Table 2 : P002 Phase-3 pilot outcomes versus pre-registered predictions. All values from projects/P002_tightbinding_attention/pilot/results.json , which is padding-inclusive. The AIC-winner row and the three median-layer R2 rows have been recomputed with padded positions excluded, on panel B (the 4 documents of length ≥120 , τ≤113 ); the padded value follows in parentheses. Note the τ window differs between the two (113 masked versus 128 padded), so the pair is a before/after of the same pipeline rather than a like-for-like refit. Masked values: pilot/masked_median_layer_r2.json .
Figure 3 : P002 attention-versus-distance histograms at the median layer ( L=12 of 24) of GPT-2-medium. The 4 × 4 grid contains all 16 heads, with the three fits overlaid in each panel: H1 exponential (red), H2 power law (blue), and H3 stretched exponential (green). Across the 16 heads, H1 undershoots at short distances and overshoots in the tail; H2 follows the data across the full 1≤τ≤128 window, and H3 nearly coincides with H2. AIC selects H2 in 15 of the 16 heads at this layer. Source: projects/P002_tightbinding_attention/pilot/figure_decay_layer12.png .
Figure 4 : P002 per-layer fit quality R2 for H1 (exponential, red), H2 (power law, blue), and H3 (stretched, green) on GPT-2-medium. H2 and H3 dominate H1 at every depth except the final layer, where all three collapse. Source: projects/P002_tightbinding_attention/pilot/figure_fit_quality_per_layer.png .
Figure 5 : P002 perplexity versus TB cutoff multiplier k (cutoff radius =k⋅λ ), compared with full attention and matched uniform-window, BigBird, and top- k controls. TB falls sharply from a catastrophic value at k=1 but does not reach full attention by k=5 , where it remains approximately +15% above the full-attention baseline and well above the pre-registered 5% shelve threshold. For each k , matched top- k (purple) tracks full attention; the matched uniform window is catastrophic at all k ; and BigBird stays near ∼160 ppl. The headline “TB at 3λ costs +96% perplexity” is the k=3 point. Source: projects/P002_tightbinding_attention/pilot/figure_cutoff_vs_baselines.png .
Method
Pre-registered s⋆ range
Measured s⋆
identity (sanity floor)
0.0
0.058±0.000
random Haar
0.05–0.20
0.052±0.006
PCA on WKWQ⊤
0.35–0.55
0.051±0.000
Wannier (Candidate C r^ )
0.55–0.75 (central 0.65)
0.054±0.004
Welch t , Wannier > PCA
passes if t≥4.5
t=1.16 , p=0.187 (fail)
Welch t , Wannier > random
passes if t≥4.5
t=0.46 , p=0.337 (fail)
Table 3 : P001 Phase-3 pilot outcomes versus pre-registered predictions on Pythia-160M. s⋆ is the maximum sparsity at ≤2% WikiText-103 perplexity loss; mean ± std over n=3 seeds. Welch t is one-sided, n=3 vs n=3 . All values from projects/P001_wannier_attention/pilot/results.json .
Figure 6 : P001 perplexity versus sparsity for four rotations on Pythia-160M, reported as mean ± std over n=3 seeds. All four curves remain within seed noise at s≤0.3 , and all four reach the ≤2% perplexity-loss line at s⋆≈0.05 . The horizontal dashed line marks the pre-registered 2% -of-baseline threshold. Source: projects/P001_wannier_attention/pilot/figure_sparsity_vs_ppl.png .
Figure 7 : P001 per-layer ΩWannier(ℓ) , averaged over heads, on Pythia-160M. Mid-layers (4–8) have the lowest Ω ; edge layers (0–3, 9–11) are roughly 1.6× larger. Thus, the pre-registered direction (mid < edge) is satisfied, but heavy-tailed weight magnitudes, not basis choice, control the overall s⋆ result. Source: projects/P001_wannier_attention/pilot/figure_per_layer_sparsity.png .
Quantity
Pre-registered prediction
Measured (GPT-2)
r(logDv⋆,Hv) , eval split
≥0.65
0.016 ( 3σ CI [0.003,0.030] )
Scaling-law slope γ
∣γ∣≳0.1
0.0002
Rows at bond cap D=16
—
94.06%
Best compression @ ≤1% ppl
≥8×
0.644× ( 1.55× inflation)
advantage vs uniform-rank
≥1.5×
1.0× (tie)
Robustness r ( Hv from train)
—
−0.004
Table 4 : P011 Phase-3 pilot outcomes versus pre-registered predictions. GPT-2 values from projects/P011_tt_embedding/pilot/results.json ; OPT-1.3B from results_opt13b.json .
Figure 8 : Cluster-level empirical inversion. Predicted and measured fractions of 3a insulator-like, 3b plasmon-like, and 3c critical regimes are operationalised by P002’s competing AIC fits — H1 exponential, H2 power law, and H3 stretched exponential — on GPT-2-medium ( n=384 (layer, head) pairs). The predicted bars render the cluster-meta §3 statement “3a majority, 3b 5–15%, 3c unknown” as 70/10/20. The data refute the qualitative inversion, not these exact percentages.
Figure 9 : Cross-paper consistency on Pythia-160M. P001 Wannier spread Ω(h,ℓ) (horizontal axis) is plotted against P002 inverse tight-binding decay rate 1/λ(h,ℓ) (vertical axis) for n=137 paired (layer, head) cells; colour denotes layer. Pearson r=−0.0583 , far below the pre-registered ∣r∣≥0.60 gate. This panel has been regenerated with padded positions excluded from both observables ; the padding-inclusive panel the first draft showed ( r=−0.0004 , n=121 ) is kept as figure_cross_paper_scatter_aspublished.png . A cluster-inconsistency event was logged when the P002 pilot concluded.
Matrix-level low-rank compression is a promising way to reduce the cost of large language models, but running compression and evaluating the resulting models on language tasks can be prohibitively expensive. Can compression-induced degradation be predicted before committing to this compute? We systematically analyze the Qwen3 and Gemma3 model families across four representative low-rank compression methods: vanilla SVD, two ASVD variants, and SVD-LLM. We find that stable rank and information density, measured in bits per parameter, dominate performance degradation. The interaction term γ⋅ρˉs, defined as compression ratio times stable rank, is a robust predictor of accuracy degradation, achieving leave-one-out cross-validation Pearson correlations of 0.890 for attention layers and 0.839 for MLP layers. We provide theoretical intuition for why this predictor succeeds by connecting it to standard SVD truncation bounds and error composition mechanisms in transformer layers. These findings enable a predict-then-compress workflow: compute γ⋅ρˉs from weights, estimate degradation, and invest compute only in desirable configurations.
Mingxue Xu
Department of Electrical and Electronic Engineering, Imperial College London, London, United Kingdom
Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation to the current compressed model while jointly exploiting the native Q/K/V head structure under a fixed storage budget. The output projection is handled separately to account for the changed post-attention representation. FTC requires neither fine-tuning nor gradient-based recovery. Across seven decoder-only LLMs from 6B to 32B parameters, FTC achieves the lowest WikiText-2 perplexity among the compared methods at every tested keep ratio on five modern GQA models, with the largest gains under aggressive compression. The improvements transfer to downstream tasks and remain substantial at the 32B scale.
Jiangfeng Chen, Xinyu Wang, Tianshuo Yan +4
University of Manitoba · McGill University · Simpleway +2
We present SigmaScale, a method for learning auxiliary scaling matrices S to aid truncated Singular Value Decomposition (SVD) based Large Language Model (LLM) compression. Instead of deriving scaling matrices analytically, SigmaScale optimizes two sets of vectors that define diagonal row and column scaling transformations under an activation-aware compression loss. We show that learned scaling lowers the effective intrinsic rank of weight matrices, as reflected by reductions in effective-rank entropy, and that this reduction is strongly correlated with compression loss. Experiments on Llama 3.1 8B Instruct and Qwen3-8B show that SigmaScale is competitive with closely related state-of-the-art SVD-based compression methods across perplexity and zero-shot benchmarks. By using learned activation-aware transformations, SigmaScale explores a more flexible route to low-rank LLM compression by adapting to the structure of individual model weights. The advantage observed in specific tasks makes our approach a valid option for applications requiring a reduced LLM-inference computing cost.
Ernests Lavrinovics, Marco Letizia, Roy Janco +3
Department of Computer Science, Aalborg University Copenhagen, Denmark · MaLGa-DIBRIS, University of Genoa, Genoa, Italy · INFN, Sezione di Genova, Genoa, Italy +2