cs.CVJun 12, 2026

Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging

Authors: Haoxu HuangLong ChenJingyun ChenJinu HyunJames Ryan LoftusKara MelmedDaniel OrringerJennifer Frontera+3 more

Organizations: 1New York University, Center for Data Science, New York, NY, 10001, USA · 2NYU Grossman School of Medicine, Department of Radiology, New York, NY, 10016, USA · 9State University of New York at Binghamton, School of Computing, Binghamton, NY 13902, USA · 3NYU Grossman School of Medicine, Department of Neurology, New York, NY, 10016, USA · 4NYU Grossman School of Medicine, Department of Neurosurgery, New York, NY, 10016, USA · 5NYU Grossman School of Medicine, Department of Pathology, New York, NY, 10016, USA · 10Stanford University, Department of Radiology, Stanford, CA, 94305, USA · 6NYU Grossman School of Medicine, Department of Neuroscience, New York, NY, 10016, USA · 7NYU Grossman School of Medicine, Neuroscience Institute, New York, NY, 10016, USA · 8NYU Grossman School of Medicine, Department of Population Health, New York, NY, 10016, USA

Abstract

Brain MRIs are routinely acquired as multiple complementary sequences with unique contrast weighting, including T1-weighed imaging (T1w) anatomic and fluid-sensitive T2-weighted (T2w) contrasts. However, methods for learning unified representations across the multitude of MRI contrast mechanisms at health-system scale are lacking. In this study, we introduce Neuro-JEPA, a sparse multimodal neuroimaging foundation model that combines a latent predictive objective with a Mixture-of-Experts architecture to encode brain MRI across core T1w, T2w, and fluid-suppressed FLAIR imaging (FLAIR). We further provide a systematic methodological study of architectural, masking, objective, and sparsity design choices beneficial for robust neuroimaging multimodal representation learning. Neuro-JEPA was pretrained on 1,551,862 scans from 428,647 studies after modality-specific preprocessing with data curation across three core structural brain MRI sequences. We evaluated the learned representations across clinical and research settings, including 25 tasks from three health systems: NYU Langone, NYU Long Island, and Massachusetts General Hospital, and 22 tasks from 12 public datasets, covering unimodal, multimodal and cross-domain evaluation configurations. Across these benchmarks, existing neuroimaging foundation models showed inconsistent gains over a simple convolutional neural network (CNN) baseline, whereas Neuro-JEPA achieved stronger and more consistent performance across all evaluated settings. These results establish a scalable methodological framework for multimodal neuroimaging representation learning and highlight the need for foundation model evaluation protocols that include simple baselines, clinically heterogeneous cohorts and controlled multimodal comparisons.

Explore similar work

CardsList
  1. Evaluating the Generalization of Neuroimaging Foundation Models on African Brain MRI

    Sep 21, 2026Oluwatobi Iyanuoluwa Akinmuleya, Olatokun Shamsudeen Akano, Samuel Danquah Ankapong +2Neuroimaging Models