Organizations: D. Bradley McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, US. · Department of Biomedical Informatics and Data Science, Yale School of Medicine, New Haven, US. · Department of Electrical and Computer Engineering, Texas A&M University, College Station, US. · Department of Psychiatry and Behavioral Sciences, Stanford University School of Medicine, Stanford, CA, USA.
Existing genome-wide association studies (GWAS) of brain imaging provide predefined or deep-learning-derived imaging phenotypes, yet these phenotypes come from either volumetric scans or cortical surface meshes, so each captures only part of the heritable variation in brain anatomy. Here we introduce MEVA (Mesh-Enhanced Volumetric Autoencoder), a self-supervised framework that encodes voxel-level image intensity together with cortical mesh geometry, including curvature and cortical thickness at each surface vertex, into one shared set of imaging features. Combining the mesh and volumetric inputs in MEVA yields modest performance gains in age and sex prediction over models that use either input alone. When these features serve as phenotypes for GWAS in the UK Biobank, they reveal more genome-wide significant loci than features learned from volumes alone or from meshes alone. These results suggest that adding cortical surface geometry to volumetric self-supervised learning captures additional heritable variation and so increases the number of loci detected.
Figures & tables
Figure 1: Mesh-only autoencoder. A masked graph autoencoder encodes the fsaverage4 cortical mesh of each hemisphere, reconstructs the masked vertex features and reads out the embedding zmesh . Each vertex carries its mid-thickness x , y and z coordinates, cortical thickness (the distance between the pial and white-matter surfaces) and curvature, from concave (sulcal fundi) to convex (gyral crowns).
Prediction
Cortical ΔR2↑
Genetics
Representation
Age Pearson r ↑
Sex AUC ↑
Thickness
Area
Volume
Loci ↑
∑hSNP2↑
Volume-only
0.768
0.976
0.091
0.258
0.209
153
33.7
Mesh-only
0.671
0.886
0.475
0.284
0.294
40
10.5
MEVA
0.776
0.984
0.375
0.277
0.266
188
36.7
Table 1: Comparison of the learned representations. 0 0 footnotetext: All representations have 128 phenotypes and are evaluated in the same 35,298 held-out UK Biobank participants; bold marks the best value in each column. Age mean Pearson r and sex AUC are from five-fold cross-validation. Cortical information is the variance of FreeSurfer regional thickness, surface area and grey-matter volume explained beyond covariates, averaged over the 68 Desikan–Killiany regions. Bootstrap 95% confidence intervals lie within ±0.006 of every sex AUC and ΔR2 value; Loci are from fastGWA analysis of 128 phenotypes, followed by the minimal p value of aggregating multivariate summary statistics and FUMA for loci aggregation; ∑hSNP2 is the summed SNP heritability from the sum-h2 pipeline (Methods).
Figure 2: MEVA architecture, training objective and phenotype extraction. a , A T1-weighted MRI volume and its FreeSurfer-reconstructed cortical surface are used as inputs. b , A 3D CNN encodes the volume into zvol , while a pretrained, frozen mesh autoencoder encodes the surface into zmesh . A self-attention block jointly processes the embeddings, and its output is projected to a residual δ that is added to zvol to obtain zfused∈R128 . Training uses T1 reconstruction and mesh-embedding prediction losses, L=LT1+λLmesh ( λ=10 ). c , The trained encoder generates 128 imaging phenotypes from each participant. These phenotypes are used for genetic discovery and anatomical evaluation, including age, sex and cortical ΔR2 prediction.
Figure 3: Genetic associations across imaging representations. a , Overlap of genome-wide significant loci identified by the volume-only (153 loci), mesh-only (40 loci), and MEVA (188 loci) representations; loci were merged across models when their boundaries, padded by 125 kb on each side, overlapped. Forty-two loci were identified only by MEVA. b , Manhattan plot of the MEVA minimum-P GWAS ( N=35,298 ). SNPs within the 42 MEVA-specific loci are highlighted in yellow.
Symbol
Meaning
x , x^
Input T1-weighted volume and its reconstruction
zvol∈R128
Embedding of the T1 volume from the volumetric CNN encoder; the volume-only phenotypes
zmesh∈R128
Embedding of the cortical surface from the frozen mesh encoder; the mesh-only phenotypes
zmesh′
zmesh after the learned adapter
δ∈R128
Surface-informed residual computed by the fusion module
zfused=zvol+δ
MEVA phenotypes
Table S1: Notation.
Module
Parameters
Training
CNN autoencoder
138.1M
40 epochs (pretraining)
Mesh autoencoder
0.03M
epoch 98 of 100
MEVA fusion modules
0.18M
4 + 6 epochs (two stages)
Table S2: Model capacity and training. 0 0 footnotetext: The CNN autoencoder gives the volume-only phenotypes and initializes the MEVA volume encoder, which is fine-tuned during fusion; it has 69.6M encoder and 68.5M decoder parameters. The mesh autoencoder gives the mesh-only phenotypes and is frozen in MEVA. The MEVA fusion modules are the mesh adapter, the fusion block and the distillation head.
Self-supervised learning offers a compelling approach for medical imaging, where labeled data are scarce and acquisition costs are high. We present COJEPA, a self-supervised framework for volumetric brain MRI that combines a joint-embedding predictive architecture (JEPA) with a contrastive loss (CO), targeting two complementary properties: local predictivity and global discriminability. The model is trained without labels on T1-weighted structural MRI from two cohorts (HCP-YA and AABC, N=2286, ages 22 to 90), extending I-JEPA to 3D with foreground-aware block masking, a hierarchical convolutional patch embedding, and world-space sinusoidal positional encodings. We evaluate all three objectives across zero-shot twin retrieval, brain tumor segmentation (BraTS 2024), and age regression (OpenBHB). COJEPA achieves the best monozygotic twin recall at rank@1 (0.84), the best finetuning age MAE (2.55 years on OpenBHB 3.0T), and matches CO on BraTS whole-tumor Dice, demonstrating that the combined objective yields representations that are simultaneously discriminative and locally structured.
Fabian Mager, Lars Kai Hansen
Department of Applied Mathematics and Computer Science Technical University of Denmark
We study the problem of training self-supervised foundation models for functional MRI. Our main contributions are: (1) we introduce a new model family (CortexMAE) trained using the masked autoencoder framework on 2.1K hours of open fMRI data, and (2) we release the first open evaluation suite (Brainmarks) for fMRI foundation models. Our core innovation is simple: we adapt the Vision Transformer to fMRI by first converting each 3D fMRI volume to a 2D map using a cortical flat map projection. We directly compare flat maps to both parcellation and volume-based representations. While each has its advantages, flat maps generally perform best. We perform the first systematic scaling analysis for fMRI and observe strict power law scaling, albeit with limits. Finally, we use Brainmarks to do controlled benchmark comparisons. On subject-level trait prediction, we report a challenging null result: no single model achieves clear state-of-the-art performance. Moreover, all models struggle to outperform a simple functional connectivity baseline. On cognitive state decoding, we observe more robust performance, and in this setting our CortexMAE family outperforms prior models by a large margin. Code, models, and datasets are available at https://github.com/MedARC-AI/CortexMAE and https://github.com/MedARC-AI/Brainmarks.
Deep learning models for neuroimaging have largely been developed for individual tasks, limiting knowledge transfer across applications. Here we introduce GenFAR, a modular deep learning framework that learns general, clinically informed features from brain MRIs. We trained this modular architecture on 49,246 individuals across 11 cohorts, using 17 diverse classification and regression tasks spanning cognition, clinical, diagnosis, demographics, and biomarkers. This yields aggregated, focused feature sets that capture rich, clinically- and biologically-relevant brain representations. We developed a sequential learning approach where tasks progressively build on previously learned representations. Through an analysis of 5,000 task sequences, we identified an optimal sequence length of six tasks and introduced a Donor Score metric to quantify each task's contribution to downstream performance. This analysis revealed five consistently strong donor tasks (Age, AD/MCI, MMSE, Hypertension, Hyperlipidemia) that formed the base of our sequential model. We demonstrated the utility of our learned representation, in various tasks beyond those included in the training set, to serve as the foundation for specialized secondary predictors. We further showed that using the learned feature representation can substantially increase the sample efficiency of secondary deep learning training tasks and models, as well as improve their accuracy.
Vishnu M. Bashyam, Guray Erus, Junhao Wen +29
Artificial Intelligence in Biomedical Imaging Lab, University of Pennsylvania · Department of Electrical and Systems Engineering, University of Pennsylvania · Florey Institute of Neuroscience and Mental Health, University of Melbourne +16