Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation Models
Authors: Sangyoon Bae, Sk Miraj Ahmed, Shinjae Yoo, Jiook Cha
Organizations: Interdisciplinary Program in Artificial Intelligence Seoul National University Seoul, 08826, South Korea · Computational Science Initiative Brookhaven National Laboratory Upton, New York, 11973, USA · Department of Psychology Seoul National University Seoul, 08826, South Korea
Can we determine whether a foundation model will generalize out-of-distribution (OOD) before any target data is available? Existing diagnostics require source or target data, which rules them out before a target domain exists. Those that use the weights alone apply one statistic to every architecture, and do not separate robust models from fragile ones. We show the answer is encoded in the spectral structure of pretrained weights. Two forces shape that structure. Architecture determines how information is stored in weight matrices. Pretraining strategy determines what is rewarded. Together they set a spectral geometry that governs OOD robustness. We prove that the OOD accuracy gap is bounded by how tightly the source representations concentrate. A statistic computed from the pretrained weights alone serves as a proxy for that concentration. The direction of that proxy reverses between architecture families. We operationalize it: the direction is stable within one (architecture X strategy) combination, the finest grouping we test, which we call a cell. Pooled over 116 models spanning 7 modalities, a single statistic ranks OOD robustness weakly, because cells of opposite direction cancel. Within a cell, the statistic selected for it orders 92% of model pairs by OOD robustness in-sample. The selection does not leak the target: for each model family outside the matrix we logged the cell, metric and sign before running its OOD evaluation, and the predicted direction held in every case: EEG, genomic and protein. Acting on spectral concentration narrows the OOD gap by 24% at 87.5% ID retention. The diagnostic operates on released weights alone, so OOD robustness becomes checkable at model-selection time, before data or compute is committed to a target domain.
Figures & tables
Figure 1: (Architecture × Strategy) → Spectral Geometry → OOD Robustness. Architecture and pretraining strategy jointly shape the layer-wise singular-value spectrum (2) , which determines whether representations are concentrated or diffuse (3) , setting the OOD accuracy gap (4) . The sign of ρ is a property of the cell and of the statistic read in it, not of the domain: each AR-CLM cell is positive under its own calibrated metric, while BiDi-MLM/Protein, CNN and SSM-v2/RWKV are negative under theirs.
Table 2: The cell fixes the metric; neither axis alone does. Holding one axis and varying the other changes the statistic selected for the cell. Each row reads off Table 1 .
Level
Conditioning unit
Mean ∣ρLOO∣
Δ
0
Global (all 116 models pooled)
0.245
—
1
Strategy-only (4 groups; one metric per objective)
0.478
+0.233
2
Architecture-only (8 groups; one metric per backbone)
0.745
+0.267
3
Architecture × Strategy (ours)
0.824
+0.579
3 ′
Arch × Strategy ( lognparams instead )
0.686
+0.441
Table 3: Four-level conditioning hierarchy on the identical n=116 model set. Mean ∣ρLOO∣ across groups at each level; best metric selected independently per group (oracle upper-bound at each level). Neither axis alone attains the full cell-level signal.
Figure 2: End-to-end pipeline: Diagnose & Localize → Targeted Intervention → Improve. (A) Cell-specific spectral metric ( CV(PR) for BiDi-MLM) identifies fragile layers via per-layer metric peaks; ρLOO on the full reference cell validates its predictive signal. (B) Only metric-flagged layers are repaired (DICE for BiDi/MLM; SpNorm for AR-CLM); stable layers are untouched. (C) Targeted repair reduces OOD drop in every model shown (C1). The targeted > random > anti-targeted ordering supports the metric’s ranked ordering of layers, rather than repair alone, as the source of the OOD gain (C2).
Modality
Model
Δtargeted
Δrandom
Δanti
NLP BiDi
BERT-base
−0.224
−0.163
−0.150
NLP AR-CLM
GPT-2-large
−0.348
−0.098
−0.046
EEG
CBraMod
−0.038
−0.033
+0.026
Protein
ESM2-35M
−0.133
+0.003
−0.048
Vision
mean (25)
−0.078†
0 (ref)
−0.025
Table 4: Metric-guided repair across modalities. Δ = change in OOD drop from the no-intervention baseline (negative = improvement), except in the Vision row, where Δ is measured against random targeting. Vision uses PCA-bottleneck (mean over 25 models); others use SpNorm.
Figure 3: Diagnosis-to-improvement: spectral interventions. Panel A (SVD bottleneck): projecting features onto top- r singular vectors sets log-srfeat≤logr ; color = ID accuracy retention. Panel B (null controls): γ -surgery and LoRA rank leave log-srfeat unchanged and yield no OOD change.
Family
Cell and metric (locked)
Logged prediction
Outcome
EEG ( n=11 )
BiDi-Transformer and BiMamba/Linear-attn, EEG; log-NFR , positive (data-scarcity inversion)
Positive direction in both sub-groups
Confirmed. Pooled ρLOO=+0.822 ( p=0.001 ); sub-groups +0.880 and +0.817 , so the inversion is architecture-independent (Appendix D.4 ). Session-level ranking ρLOO=+1.000 ( n=4 , p=0.042 ).
Protein: AMPLIFY 120M/350M ( Fournier et al., 2024 )
BiDi-MLM × Protein; PR-mean , negative
350M more robust than 120M
Confirmed on direction. PR-mean=79.5>64.0 and Δacc=0.132<0.267 . The logged rank-percentile verdict was not confirmed, and the ESM-2 size–fragility calibration did not cross the RoPE vs. absolute-PE subtype boundary (Appendix S ). Both are now in the primary cell ( n=11 ).
7B ( log∥W∥F=4.120 ) more robust than 1B ( 3.633 )
Confirmed. 7B reaches OOD NLL 26 – 49% below random across 5 species; 1B is near chance.
Spike: NDT3 ( Schmitt and others, 2024 ) , POYO-1 ( Azabou et al., 2023 ) ( n=2 )
Excluded from the matrix; too few models to calibrate
—
Illustrative only. Feature-rank ratio ( 41× ) tracks the OOD drop ratio ( 4.5× ), consistent with Theorem 1 .
Table 5: Prospective tests on model families outside the 15-cell matrix. Cell, metric and sign were taken from the locked registry before any OOD outcome was observed. Spike models are reported qualitatively: with two models the cell cannot be calibrated.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Failure mode taxonomy. Each row corresponds to one (architecture, strategy, modality) cell. Columns: failure mode type (color-coded), best metric and ρLOO , and recommended repair method. Six failure modes: A Depth-Hierarchy Dissolution ( CV(log-Frob) , SpNorm), B Head Specialization Erasure ( CV(PR) , DICE), C Spectral Nucleation ( log-NFR , SpNorm), D Dimension Collapse ( PR-mean , SpNorm), E Bulk-Noise Inversion ( MP-energy , SpNorm), F Hierarchical Stage Heterogeneity ( CV(SR) / PR-std , SpNorm). Inverted direction ( ↑ρLOO ): EEG, Genomic, SSM. Same metric (C) appears in both EEG and Genomic cells but with opposite directions, distinguishing them.
Cell
OOD evaluation task
n
Δaccˉ
σΔacc
AR-CLM (NLP)
SNLI → ANLI-R1 ( Bowman et al., 2015 ; Nie et al., 2020 ) (domain transfer)
Table 6: Per-cell OOD evaluation task and within-cell OOD drop statistics. Each cell has one fixed OOD task; all models in the cell are evaluated on the same source → target shift. σΔacc (within-cell std of OOD drop) confirms non-trivial variation to predict.
Figure 5: Log stable rank vs. OOD accuracy drop, stratified by OOD type. Best metric per type shown.
Cell
n
ρLOO
95% CI
p
AR-CLM (NLP)
16
+0.714
[+0.35,+0.90]
0.0016∗∗∗ (Fisher- z )
AR-CLM (TS) ∗
8
+0.714
[−0.03,+0.94]
0.029∗ (one-sided)
AR-CLM (Protein)
10
+0.817
[+0.39,+0.96]
0.004∗∗∗ (Fisher- z )
BiDi-MLM (NLP)
10
+0.808
[+0.36,+0.95]
0.0049∗∗∗ (Fisher- z )
BiDi-MLM (Protein)
11
−0.891
[−0.97,−0.65]
0.0002∗∗∗ (Fisher- z )
BiDi-MLM (Genomic)
7
+0.829
[+0.43,+0.98]
0.0068∗∗∗ (Fisher- z )
Appendix
Table 7: Per-cell LOO Spearman correlation with Fisher- z 95% confidence intervals. For n≤5 : exact permutation p -values; n=4 : one-sided exact p=0.042 . † CI includes zero; cell is marginal (not significant at two-sided α=0.05 ). ∗ significant at α=0.05 one-sided with pre-specified direction.
Figure 6: Left : log nparams vs. log∥W∥F , colored by modality ( ρ=+0.47 , moderate). Right : Kendall τ comparison for OOD drop prediction: log∥W∥F ( τ=+0.327 , p=0.004 ) vs. lognparams ( τ=+0.068 , n.s.).
Cell (arch × strategy)
Selected metric
n
∣ρLOOspectral∣
∣ρLOOlogn∣
ρLOOlogn (signed)
AR-Transformer × CLM (NLP)
log∥W∥F
16
0.714
0.584
+0.584
AR-Transformer × CLM (Timeseries)
PR-mean
8
0.714
0.282
+0.282
AR-Transformer × CLM (Protein)
CV(log-Frob)
10
0.817
0.364
−0.364
BiDi-Transformer × MLM (NLP)
PR-mean
10
0.808
0.299
+0.299
BiDi-Transformer × MLM (Protein)
PR-mean
11
0.891
0.149
+0.149
BiDi-Transformer × MLM (Genomic)
log∥W∥F
7
0.829
0.750
+0.750
Appendix
Table 8: Per-cell LOO Spearman correlation: champion spectral metric vs. lognparams . ∣ρLOOlogn∣ shows the absolute correlation; ρLOOlogn (signed) indicates direction. Bold ∣ρLOOlogn∣ : logn has higher ∣ρ∣ and the correct pre-specified direction. † : logn achieves higher ∣ρ∣ but in the wrong pre-specified direction (spectral wins on direction). ‡ : spectral champion and logn are tied (both ∣ρLOO∣=1.000 ). \lx@sectionsign : n<6 due to limited public checkpoints for this architecture family.
Weight-space alternatives (data-free)
Runtime baselines †
Cell
n
Champion
logn
αmean
log∥W∥F
log-sr
PR
Hspec
MP
CV(log-Frob)
log-cond
PR-std
Doctor
MC-H
Mahal
AR-CLM (NLP)
16
+0.714 ( log∥W∥F )
+0.584
−0.027
+0.605
+0.480
+0.487
+0.502
−0.236
−0.744
+0.175
+0.464
—
—
—
AR-CLM (TS)
8
+0.714 ( PR-mean )
+0.282
+0.461
+1.000
+0.135
+1.000
+0.453
+0.461
−0.918
+0.600
−0.094
+0.280
+0.280
—
AR-CLM (Protein)
10
+0.817 ( CV(log-Frob) )
−0.364
+0.200
−0.400
−0.350
−0.567
−0.633
−0.150
+1.000
−0.583
−0.417
—
—
—
BiDi-MLM (NLP)
10
+0.808 ( PR-mean )
+0.299
−0.709
+0.053
+0.593
+0.577
+0.397
−0.116
−0.169
+0.005
+0.246
—
—
—
BiDi-MLM (Protein)
11
−0.891 ( PR-mean )
+0.149
−0.183
−0.467
+0.750
+0.750
+0.583
−0.183
−0.933
+0.433
+0.367
—
—
—
Appendix
Table 9: Cell-specific spectral champion vs. all alternatives. Per-cell LOO Spearman ρLOO : the selected spectral metric (“Champion,” computed from weights only, chosen per cell using OOD outcomes) vs. 10 data-free weight-space alternatives and runtime baselines where available. Bold : competitor ∣ρLOO∣≥ champion’s; dash: n<4 ; † : requires source-domain test data and inference passes. Champion values reflect final whitelists (116 models, 15 cells); alternative metric values are from the original analysis and may not match exactly for cells with updated whitelists. Hspec is the spectral entropy of the eigenvalue distribution; it is a different statistic from Hent , the layerwise entropy that is the champion of the SSM (Mamba) cell, and the two are not interchangeable. Mean ∣ρLOO∣ : spectral champion 0.824 vs. best uniform competitor ( PR , 0.724 ) and logn ( 0.686 ).
Figure 7: CLM cross-modal: why per-cell metric selection is necessary. CV(log-Frob) vs. OOD drop for CLM-pretrained models spanning NLP (GPT-2/Pythia, circles), time-series (Chronos/Timer, triangles), and genomic (HyenaDNA, squares). OOD benchmarks differ per domain and are not cross-domain comparable. A single metric ( CV(log-Frob) ) does not track OOD fragility consistently across domains: NLP shows ρLOOCV(log-Frob)=−0.744 (wrong direction), while Genomic shows +0.700 . Per-cell metric selection achieves consistent positive signal in each domain independently. Per-domain LOO (selected metric): NLP ρLOO=+0.714 ( n=16 , log∥W∥F ); TS ρLOO=+0.714 ( n=8 ; PR-mean , data-scarcity regime); Genomic ρLOO=+0.700 n.s. ( n=7 , log∥W∥F ).
Conditioning level
Mean ∣ρLOO∣
Notes
Global (no conditioning)
0.245
Single best metric across all 116 models pooled
Strategy-only (4 groups)
0.478
CLM/MLM/Contrastive/Supervised; architecture-specific sign reversals suppress signal within each group
Architecture-only (8 groups)
0.745
Same arch with different strategies (e.g. ViT × CLIP vs. ViT × Supervised) require different metrics and directions
Architecture × Strategy (15 primary cells)
0.824
Distinct selected metric per cell; consistent direction within each cell
Appendix
Table 10: Signal recovery across conditioning levels. For each level, mean ∣ρLOO∣ is computed by finding the best single metric for each group via LOO Spearman on the available models, then averaging across groups. Among the levels compared here, architecture × strategy is the coarsest unit that still achieves strong signal.
Figure 8: Architecture-conditional rank-score composition. Left : Spearman ρ between rank-sum composite score and OOD accuracy drop, for each architecture family (rows) × composite (columns). Scausal is optimal for causal/AR Transformers; Sssm is jointly optimal for bidirectional MLMs and SSMs. Right : ESM-2 (BiDi MLM) vs. ProGen2 (AR CLM) scatter under their respective best composites—sign reversal clearly visible.
Predictor
Cells
Mean ∣ρLOO∣
Data / conditioning needed
αmean ( Martin et al., 2021 ) (uniform)
15 cells
0.448
None
lognparams (uniform)
15 cells
0.686
None
PR (best uniform, no architecture conditioning)
15 cells
0.724
None
Doctor score ( Sun et al., 2022 )
6 cells
0.694
Source test
Mahalanobis dist. ( Lee et al., 2018 )
4 cells
0.562
Source test + labels
Cell-specific (ours, mean ∣ρLOO∣ )
15 primary cells
0.824
Arch + strategy label
Appendix
Table 11: OOD prediction methods compared by per-cell mean ∣ρLOO∣ (116 foundation models across 15 primary cells; see Table 9 for per-cell breakdown). Weight-space methods are applied uniformly across all architectures (no conditioning); runtime baselines require source-domain test data and are evaluated only in cells where data were available.
Table 12: Same architecture × strategy, different training domain: metric and direction adapt. The five cells listed here all show positive ρLOO ( ↑ ): higher metric value → higher OOD drop. This is not true of the matrix as a whole—BiDi-MLM/Protein, CNN, SSM-v2 and RWKV are negative (Table 1 ). Per-cell metric selection achieves strong correlation regardless of domain breadth, while a single unconditioned metric fails across domains.
Figure 9: Per-cell metric selection across domain breadth. (a) Within AR-CLM and BiDi-MLM (NLP/Genomic) families, different metrics are selected across narrow-domain (Genomic/Protein) and broad-domain (NLP) cells; all AR-CLM cells and BiDi-MLM NLP/Genomic achieve positive ρLOO with the selected metric; BiDi-MLM Protein shows negative ρLOO ( −0.891 ) reflecting the anti-collapse pretraining direction. (b) SSM/RWKV cells show distinct metrics and directions across Mamba, Mamba-2, RWKV, confirming that per-family calibration is necessary even within the SSM/linear-RNN group. The per-cell calibration requirement is universal.
Cell
Per-layer diagnosis
Failure mode
Best repair
Δ OOD
AR-CLM (NLP)
CV(log-Frob) (per-layer)
Depth-norm hierarchy
SpNorm
−0.131
BiDi-MLM (NLP)
CV(PR) (per-layer)
Head specialization
DICE
−0.224
BiDi-MLM (Protein)
PR-mean (per-layer)
Scale fragility
SpNorm
−0.133
EEG (BiDi/MLM)
log-NFR (per-layer)
Data-scarcity over-fit
Layer surgery
−0.038
ViT × Sup
MP-energy (per-layer)
Bulk-noise excess
PCA-bottleneck
−0.078†
Appendix
Table 13: Architecture-conditional diagnosis and repair. “Spectral diagnosis” is the per-layer diagnostic used to rank individual layers for targeted repair; it differs from the cross-model cell metric in Table 1 , which ranks checkpoints. Mean OOD improvement ( Δ , negative = improvement).
Cell
SeTAR
Alpha
ASH
ReAct
DICE
DARE
SpNorm
Best
BiDi/MLM ( n=4 )
−0.055
−0.007
−0.039
−0.064
−0.224
−0.038
−0.142
DICE
AR-CLM ( n=6 )
−0.021
−0.009
−0.104
−0.038
+0.000
−0.018
−0.131
SpNorm
Appendix
Table 14: Mean targeted Δ accuracy-drop by cell (negative = improvement). Best method per cell in bold .
Figure 10: Per-layer rank profiles confirm the BiDi PR mechanism. (a,b) Per-block participation ratio vs. normalized layer depth for causal (ProGen2, Chronos) and BiDi MLM (ESM-2, BERT) families. (c,d) Cross-layer PRσ vs. OOD drop. For BiDi MLMs ( n=13 ): cross-layer PRσ achieves ρ=+0.890 ( p<0.001 ), confirming PR heterogeneity as the per-layer repair signal; the cross-model predictor is PR-mean (Proposition 4 ).
Figure 11: ID–OOD trade-off across 13 NLP models. (a) Scatter of (ID accuracy, OOD accuracy) under different repair strategies. Spectral-targeted PCA (red ★ ) consistently achieves higher OOD accuracy with less ID sacrifice than random (blue ∙ ) and anti-targeted (orange ■ ). (b) Mean ± std across all models. Targeted Pareto-dominates random in 10/13 models. No intervention (gray ◊ ) is the baseline; anti-targeted is the worst case.
Figure 12: Within-cell knee detection for BiDi-MLM × Protein. (a) Model size vs. OOD drop for ESM-2 (Absolute PE, blue) and AMPLIFY (RoPE+SwiGLU, red). ESM-2 shows a positive monotone slope ( +0.07 /log-param, R2=0.94 ); AMPLIFY inverts ( −0.57 ). The slopes are statistically separated at the PE-type boundary. (b) ESM-2 calibration prediction overlaid on actual AMPLIFY OOD drops. AMPLIFY-350M is predicted fragile (open triangle, Δacc^≈0.53 ) but is actually the most robust model in the cell ( Δacc=0.132 , residual =−0.398 ). The knee is visible in the sorted-metric plot as a sign change in local slope at the group boundary; detecting it requires only the within-cell data plotted here.
Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and cancer status using a unified frozen-backbone linear-probe protocol, training on 3 source datasets and evaluating on 12 task-compatible out-of-distribution (OOD) datasets after label harmonization. Mammography-specific vision-language models (Mammo-FM and MaMA) provide the strongest mean OOD performance, but robustness is not explained by mammography exposure alone. DINOv3 remains a competitive vision-only baseline, and mammography-adapted pretraining does not consistently improve generalization. Dataset-level analysis further shows that even leading models show heterogeneous performance across datasets. Feature-space inspection reveals that useful representations can preserve clinical signal while retaining dataset and acquisition structure. These findings highlight dataset-level OOD evaluation as a central criterion for assessing mammography representations. Our code is publicly available: https://github.com/biomedia-mira/mammo-ood.
Giang Nguyen, Raghav Mehta, Emma A. M. Stanley +4
College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam · Imperial College London, London, UK · Radiology Department, Vietnam National Cancer Hospital, Hanoi, Vietnam +2
Large-scale pretrained models are widely leveraged as foundations for learning new specialized tasks via fine-tuning, with the goal of maintaining the general performance of the model while allowing it to gain new skills. A valuable goal for all such models is robustness: the ability to perform well on out-of-distribution (OOD) tasks. We assess whether fine-tuning preserves the overall robustness of the pretrained model in image classification, and observed that models pretrained on large datasets exhibited strong catastrophic forgetting and loss of OOD generalization. To systematically assess robustness preservation in fine-tuned models, we propose the Robustness Inheritance Benchmark (ImageNet-RIB). The benchmark, which can be applied to any pretrained model, consists of a set of related but distinct OOD (downstream) tasks and involves fine-tuning on one of the OOD tasks in the set then testing on the rest. We find that though continual learning methods help, fine-tuning reduces robustness across pretrained models. Surprisingly, models pretrained on the largest and most diverse datasets (e.g., LAION-2B) exhibit both larger robustness losses and lower absolute robustness after fine-tuning on small datasets, relative to models pretrained on smaller datasets. We observe this collapse in contrastively pretrained (CLIP) models and their fine-tuned variants, where it grows with pretraining scale; the supervised models we test do not exhibit it. These findings suggest that starting with the strongest foundation model is not necessarily the best approach for performance on specialist tasks. https://jd730.github.io/projects/ImageNet-RIB
Models trained with deep learning often fail to signal when inputs fall outside their training data manifold, leading to unreliable predictions under distribution shift. Prior work suggests that effective out-of-distribution (OOD) detection often requires class-conditional modeling or specialized models obtained through supervised fine-tuning. We revisit this assumption in modern pretrained models and show that their frozen representations already encode sufficient geometric structure for accurate label-free OOD detection. Across 59 backbone-task pairings spanning vision and language, we compare two complementary label-free detectors: a global Mahalanobis estimator fit on unlabeled latent representations, and ReSCOPED, a lightweight, diffusion-based typicality estimator operating on the same features at a local level. Despite their different detection mechanisms, representation scaling reveals a consistent regime-dependent pattern: both local and global detectors' absolute performance improves with better representation quality, and performance gaps between the two detectors disappear across both language and vision tasks as representations scale. These results suggest that label-free OOD detection depends strongly on the geometry exposed by frozen pretrained backbones, reducing the importance of detector choice as backbone scale increases and enabling efficient deployment directly on frozen models.
Brett Barkley, Preston Culbertson, David Fridovich-Keil
Department of Electrical and Computer Engineering The University of Texas at Austin Austin, TX, USA · Department of Computer Science Cornell University Ithaca, NY, USA · Department of Aerospace Engineering and Engineering Mechanics The University of Texas at Austin Austin, TX, USA