astro-ph.IMSep 23, 2025

The Platonic Universe: Do Foundation Models See the Same Sky?

Authors: UniverseTBD, :, Trinidad Borrell, Steven Dillmann, Kshitij Duraphe, Furkan Eris, Kartheik Iyer, Ashod Khederlarian, +6 more

Organizations: Independent Researcher · AstroAI · Harvard-Smithsonian CfA · University of Hertfordshire · Washington University St. Louis · Space Telescope Science Institute · Johns Hopkins University

Abstract

We investigate when foundation models converge towards shared representations, and how this convergence depends on model capacity, training regime, and model architecture. We take a `science-for-AI' approach, using astronomy as an experimental instrument to test the Platonic Representation Hypothesis and its Aristotelian refinement against an external physical reference. The historical success of astrophysics is evidence that a compact, modality-invariant description of galaxy observables exists, and so representation convergence toward reality should be measurable against the physical parameters astronomers already use. Given this framework, we evaluate eleven foundation model families (spanning classification, self-distillation, joint-embedding prediction, autoencoding, vision-language pre-training, and astro-specific architectures from O\mathcal{O}(10M)→O{\to}\mathcal{O}(10B) parameters) on crossmatched JWST, HSC, and Legacy imagery, and DESI spectroscopy. All models are evaluated frozen, with no astronomy-specific fine-tuning. We probe redshift, stellar mass, and sSFR via linear probes, and local (MKNN) and global (CKA) embedding geometry within families, between modalities, and across architectures. We find that physics performance scales predictably with capacity; probe directions align consistently with expected astrophysical correlations and selection effects; and local (not global) embedding alignment tracks physics performance, including between DESI spectra and HSC imagery---modalities that share essentially no low-level statistics. Our results support the ARH over the strict PRH, demonstrate astronomy's value as an experimental framework for neural representation learning, and suggest that astro-foundation models can build on general-purpose pre-trained architectures, capitalizing on the broader open machine learning community's already-spent computational investment.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 29, 2026astro-ph.IM

Gestalt: a meta-foundation model for astronomy

The Platonic Representation Hypothesis predicts that sufficiently scaled foundation models converge on a shared representation of the world. As each non-converged model gives a noisy view of a common structure when passed the same input, we ask whether we can combine models into a representation that outperforms its individual components. We test this on galaxies: we embed images via a basket of 22 frozen foundation models from eight families, whiten each view, and take a randomised SVD of the embedding concatenation. The resulting 1024-dimensional embedding outperforms every basket member on 19/21 of our tested metrics for physical property and galaxy morphology estimation for HSC, JWST, and DESI Legacy Survey imagery. We find that performance rises with basket size and basket architectural diversity, and that the meta-foundation model's performance transfers across astronomical surveys. We conclude that a useful astronomical foundation model can be assembled from existing generalist models with no training required beyond a single unsupervised projection. By leveraging the community's already-spent work, we save a lot of compute: a fresh pre-train of a comparable single-domain model would cost O(104\mathcal{O}(10^{4}--105)10^{5}) A100 GPU hours (emitting several tonnes of CO2_2eq.), whereas assembling Gestalt requires minutes on a single machine.
Apr 20, 2026cs.CV

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

The Platonic Representation Hypothesis posits that neural networks trained on different modalities (e.g., text and images) converge toward a shared representation of reality. If true, this has significant implications for whether modality choice matters at all. In this paper, we show that the evidence for this claim is substantially weaker than subsequent work suggests. The mutual kk-nearest-neighbor metric used on 1024 text-image pairs in the original study captures only coarse structure. To keep the alignment from collapsing as one scales up the data, kk has to grow proportionally, undercutting the argument for fine-grained representational convergence. The reported increase in alignment with language model strength saturates for recent models. Moreover, the one-to-one text-image pairing favors alignment, while alignment decreases with non-bijective data. We further find that image and text representations indeed share coarse semantic structure, but neither stronger language models nor richer captions yield fine-grained alignment. Thus, multimodal representations share coarse structure without evidence of convergence to a shared representation -- arguably, full representational convergence would require fine-grained alignment.
Apr 27, 2026cs.AI

A systematic evaluation of vision-language models for observational astronomical reasoning tasks

Vision-language models (VLMs) are increasingly proposed as general-purpose tools for scientific data interpretation, yet their reliability on real astronomical observations across diverse modalities remains untested. We present AstroVLBench, a comprehensive benchmark comprising over 4,100 expert-verified instances across five tasks spanning optical imaging, radio interferometry, multi-wavelength photometry, time-domain light curves, and optical spectroscopy. Evaluating six frontier models, we find that performance is strongly modality-dependent: while one model (Gemini 3 Pro) emerges as the most consistently capable across tasks, task-specific strengths vary, and all models substantially underperform domain-specialized methods. Mechanistic ablations reveal that performance depends not only on directing attention to salient visual features but also on grounding those features in physical knowledge. Phenomenological prompts describing what to look for improve accuracy by sharpening model focus, but physical prompts explaining why those features matter perform better overall and yield more balanced classifications with reduced class-specific bias. Consistent with this picture, presenting the underlying one-dimensional measurements directly as numerical tables instead of rendered plots yields up to 13 percentage points improvement. Reasoning quality analysis further demonstrates that, without explicit physical grounding, models may reach correct predictions from phenomenologically plausible cues while providing physically imprecise justifications, establishing that accuracy alone is insufficient for trustworthy scientific deployment. These findings provide the first systematic, multi-modal baselines for VLMs in observational astronomy and identify the specific representation, grounding, and reasoning bottlenecks where current models fail.