The Platonic Universe: Do Foundation Models See the Same Sky?
Authors: UniverseTBD, :, Trinidad Borrell, Steven Dillmann, Kshitij Duraphe, Furkan Eris, Kartheik Iyer, Ashod Khederlarian, +6 more
Organizations: Independent Researcher · AstroAI · Harvard-Smithsonian CfA · University of Hertfordshire · Washington University St. Louis · Space Telescope Science Institute · Johns Hopkins University
We investigate when foundation models converge towards shared representations, and how this convergence depends on model capacity, training regime, and model architecture. We take a `science-for-AI' approach, using astronomy as an experimental instrument to test the Platonic Representation Hypothesis and its Aristotelian refinement against an external physical reference. The historical success of astrophysics is evidence that a compact, modality-invariant description of galaxy observables exists, and so representation convergence toward reality should be measurable against the physical parameters astronomers already use. Given this framework, we evaluate eleven foundation model families (spanning classification, self-distillation, joint-embedding prediction, autoencoding, vision-language pre-training, and astro-specific architectures from O(10M)→O(10B) parameters) on crossmatched JWST, HSC, and Legacy imagery, and DESI spectroscopy. All models are evaluated frozen, with no astronomy-specific fine-tuning. We probe redshift, stellar mass, and sSFR via linear probes, and local (MKNN) and global (CKA) embedding geometry within families, between modalities, and across architectures. We find that physics performance scales predictably with capacity; probe directions align consistently with expected astrophysical correlations and selection effects; and local (not global) embedding alignment tracks physics performance, including between DESI spectra and HSC imagery---modalities that share essentially no low-level statistics. Our results support the ARH over the strict PRH, demonstrate astronomy's value as an experimental framework for neural representation learning, and suggest that astro-foundation models can build on general-purpose pre-trained architectures, capitalizing on the broader open machine learning community's already-spent computational investment.
Figures & tables
Model
Training regime
Modality
Supervised
ViT ( Dosovitskiy et al., 2020 )
Supervised classification
Images
ConvNeXtv2 ( Woo et al., 2023 )
Supervised classification
Images
Self-supervised
DINOv3 ( Siméoni et al., 2025 )
Self-distillation
Images
IJEPA ( Assran et al., 2023 )
Joint-embedding prediction
Images
Table 1: Foundation models used in this study, spanning supervised, self-supervised, autoregressive, and multimodal training paradigms across vision, language, and spectral modalities.
Figure 1: Linear probe R2 as a function of model size, averaged over mass, sSFR, and redshift. We find a significant positive correlation, even though AstroPT is the only model in our basket pretrained substantially on galaxy imagery.
Figure 2: Linear probe R2 scores of JWST versus HSC modalities on a grid of physical parameters. Model performance is strongly correlated between the instruments, indicating that a model’s capacity to linearly encode physical information is a property of its representations rather than a function of the observed wavelength range or modality.
Figure 3: Cosine similarity matrices between the redshift ( z ), stellar mass ( M⋆ ), and sSFR probe weight vectors for HSC (left) and JWST (right), averaged over all models. The relationships between the three probe weight vectors are consistent across all models, suggesting that the embedding spaces encode physics via similar internal organisation.
Figure 4: Intra-architectural scaling results. We plot our MKNN and CKA similarity values within each architecture family against the mean R2 probe performance across the physical tasks described in § 3.1 . In each pane we state the ρ and p values as measured via a Spearman’s rank test.
Figure 5: Crossmodal MKNN and CKA metrics vs R2 (§ 3.1 ). R2 and MKNN/CKA are demeaned per architecture family to show the trend without architecture family confounders. We state β and p from the ANCOVA tests MKNN∼R2+C(family) , and CKA∼R2+C(family) .
Figure 6: Cross-architecture metric distance (MKNN, CKA), plotted against average R2 over the three physical properties for HSC and JWST images (§ 3.1 ). We find that probe performance is significantly correlated with MKNN, but not with CKA.
Category
Name
Size
Hugging Face source
Models
AstroPTv2
15M (Small)
Smith42/astroPT_v2.0
95M (Base)
Smith42/astroPT_v2.0
850M (Large)
Smith42/astroPT_v2.0
CLIP
86M (Base)
openai/clip-vit-base-patch16
304M (Large)
openai/clip-vit-large-patch14
ConvNeXtv2
15M (Nano)
facebook/convnextv2-nano-22k-224
Table 2: Foundation models and astronomical datasets used in this study.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Hardware
Approximate Usage
NVIDIA A40 / A80
300 GPU-hours
NVIDIA A100 / A800
600 GPU-hours
NVIDIA GH200 / H100
2000 GPU-hours
NVIDIA RTX 5070 and RTX 3090
50 GPU-hours
Standard CPU cores
≈1 CPU-hour per probe
Total
3,000 GPU-hours
Appendix
Table 3: Hardware usage summary by chip type.
HSC ( R2 )
JWST ( R2 )
Model
z
logM⋆
sSFR
z
logM⋆
sSFR
AstroPTv2 Small
0.387±0.013
0.525±0.009
0.462±0.002
0.349±0.010
0.674±0.010
0.293±0.008
AstroPTv2 Base
0.443±0.017
0.556±0.012
0.502±0.006
0.472±0.015
0.771±0.008
0.336±0.005
AstroPTv2 Large
0.453±0.007
0.559±0.013
0.510±0.007
0.531±0.012
0.802±0.006
0.360±0.005
ViT Base
0.359±0.006
0.455±0.013
0.434±0.006
0.372±0.012
0.597±0.004
0.306±0.012
ViT Large
0.369±0.007
0.474±0.007
0.440±0.005
0.408±0.010
0.638±0.006
0.340±0.009
Appendix
Table 4: Downstream regression performance ( R2 ) when predicting redshift ( z ), stellar mass ( logM⋆ ), and specific star formation rate (sSFR) from frozen embeddings of each model, evaluated on 45 000 HSC and JWST galaxies. Values are mean ± standard deviation over 10 k-folds.
Figure 7: Parameter count vs probe performance for our tested physical properties. Spearman’s ρ and p values are stated for each panel. In all tested cases we find a significant correlation between parameter count and physics probe performance.
Figure 8: Cosine similarity matrices between the redshift ( z ), stellar mas ( M⋆ ), and sSFR probe weight vectors using HSC images.
Figure 9: Cosine similarity matrices between the redshift ( z ), stellar mas ( M⋆ ), and sSFR probe weight vectors using JWST images.
HSC MKNN (%)
JWST MKNN (%)
Model
Params (M)
Teacher
Basket
Teacher
Basket
ViT-S/16
21
25.0±0.3
24.3±0.2
27.8±0.3
23.0±0.2
ViT-S+/16
29
25.7±0.3
23.2±0.2
27.8±0.3
21.5±0.2
ViT-B/16
86
26.0±0.3
23.1±0.2
31.3±0.3
22.4±0.2
ViT-L/16
300
29.4±0.3
18.7±0.2
37.4±0.4
18.0±0.2
ViT-H+/16
840
29.3±0.3
15.0±0.2
40.5±0.3
16.0±0.2
Appendix
Table 5: DINOv3 MKNN alignment with the ViT-7B/16 teacher and the non-DINOv3 model basket, evaluated on 1.67k galaxies with paired HSC and JWST imaging. Basket alignment is averaged over non-DINOv3 models. Scores are percentages, with uncertainties estimated from 200 draws containing 80% of the sample each.
Figure 10: Physics linear R2 probing results for the PaliGemma 3B, 10B, and 28B networks on a layer-wise basis. The k=10 -fold standard deviation is shaded. We demarcate the boundary between the fixed-size SigLIP image encoder and Gemma decoder networks with a gray dashed line.
MKNN (%)
CKA (%)
Model Pairs
JWST
Legacy
HSC
JWST
Legacy
HSC
AstroPTv2 Small vs Base
47.2
35.5
35.7
96.9
98.4
98.1
AstroPTv2 Base vs Large
51.6
39.4
39.6
97.2
98.3
98.1
CLIP Base vs Large
30.7
5.9
6.2
66.7
69.4
68.8
ConvNeXtv2 Nano vs Tiny
29.6
4.4
4.8
66.3
70.3
70.4
ConvNeXtv2 Tiny vs Base
29.0
3.9
4.2
72.0
68.4
67.7
Appendix
Table 6: Intra-architectural embedding alignment within a model family, measured by MKNN and CKA scores. The PRH predicts that both MKNN and CKA scores will increase as we compare the embeddings of larger model pairs within a model family, since larger models will generate embeddings closer to the Platonic ideal representation.
Figure 11: Intra-architectural embedding alignment (MKNN and CKA) vs probe performance for our tested physical properties. Spearman’s ρ and p values are stated for each panel. In all tested cases we find a significant correlation between MKNN and physics probe performance. We do not find a significant correlation between CKA and physics probe performance in any tested case.
MKNN (%)
CKA (%)
Model
JWST
Legacy
DESI
JWST
Legacy
DESI
AstroPTv2 Small
11.62
3.36
1.25
43.4
86.6
45.66
AstroPTv2 Base
12.60
3.28
1.32
42.0
84.9
45.39
AstroPTv2 Large
14.30
3.84
1.41
41.9
84.9
44.47
CLIP Base
12.88
1.79
1.29
30.59
59.32
33.39
CLIP Large
14.07
2.28
1.24
31.89
72.73
33.85
Appendix
Table 7: Crossmodal embedding alignment between a model’s embeddings of different astronomical modalities and its embeddings of HSC imaging, measured by MKNN and CKA scores. The DESI spectra column uses Specformer embeddings for the DESI side compared against each listed vision model’s HSC image embeddings. The PRH predicts that both MKNN and CKA scores will increase as models within a family grow in size, since larger models should produce embeddings closer to the Platonic ideal representation.
Figure 12: Crossmodal embedding alignment between the stated modality and HSC plotted against probe performance for our tested physical properties. ANCOVA β and p values are stated for each panel. In all tested cases we find a significant correlation between MKNN and physics probe performance. We only find a significant correlation between CKA and physics probe performance for DESI Spectra vs HSC.
Physics subspace
Random subspace
CKA (global)
21%[9,28]
0%[−0,0]
MKNN (local)
7%[2,12]
−7%[−12,−2]
Appendix
Table 8: Effect of removing the probe-defined physics subspace from model embeddings. Values give the mean [min, max] percentage decrease in alignment over 21 tested model pairs. Negative values indicate that alignment increased after the intervention. The random control removes a subspace of equal rank.
Figure 13: Cross-architectural embedding alignment between the stated modality and HSC plotted against probe performance for our tested physical properties. Spearman’s ρ and p values are stated for each panel. In all tested JWST cases we find a significant correlation between MKNN and physics probe performance. For HSC we find significant or borderline correlations between MKNN and physics probe performance. We only find a significant correlation between CKA and physics probe performance for HSC redshift.