Scaling Vision Transformers for Functional MRI with Flat Maps
Organizations: MedARC · Sophont · Baylor College of Medicine · University of Tübingen · Georgia Tech · University of Florida
Abstract
We study the problem of training self-supervised foundation models for functional MRI. Our main contributions are: (1) we introduce a new model family (CortexMAE) trained using the masked autoencoder framework on 2.1K hours of open fMRI data, and (2) we release the first open evaluation suite (Brainmarks) for fMRI foundation models. Our core innovation is simple: we adapt the Vision Transformer to fMRI by first converting each 3D fMRI volume to a 2D map using a cortical flat map projection. We directly compare flat maps to both parcellation and volume-based representations. While each has its advantages, flat maps generally perform best. We perform the first systematic scaling analysis for fMRI and observe strict power law scaling, albeit with limits. Finally, we use Brainmarks to do controlled benchmark comparisons. On subject-level trait prediction, we report a challenging null result: no single model achieves clear state-of-the-art performance. Moreover, all models struggle to outperform a simple functional connectivity baseline. On cognitive state decoding, we observe more robust performance, and in this setting our CortexMAE family outperforms prior models by a large margin. Code, models, and datasets are available at https://github.com/MedARC-AI/CortexMAE and https://github.com/MedARC-AI/Brainmarks.
Figures & tables
| Dataset | Target | Subjects | Samples | Seq length | TR | #Classes | Majority % |
| ABIDE ( Di Martino et al., 2014 ) | ASD Dx | 578:124:124 | 578:124:124 | 150 | 2.0s | 2 | 55% |
| ADHD200 ( ADHD-200, 2012 ) | ADHD Dx | 301:64:65 | 301:64:65 | 150 | 2.0s | 2 | 57% |
| ADNI ( Jack Jr et al., 2008 ) | AD Dx | 328:41:41 | 328:41:41 | 100 | 3.0s | 2 | 77% |
| PPMI ( Marek et al., 2011 ) | PD Dx | 463:99:100 | 463:99:100 | 120 | 2.5s | 2 | 62% |
| HCP-A ( Bookheimer et al., 2019 ) | Age | 455:53:52 | 455:53:52 | 500 | 0.7s | 4 | 27% |
| HCP-A ( Bookheimer et al., 2019 ) | Sex | 471:58:55 | 471:58:55 | 500 | 0.7s | 2 | 58% |
| space | ABIDE | ADHD200 | ADNI | PPMI | HCP-A Age | HCP-A Sex | HCP-YA Task21 | NSD COCO24 |
| parcel | 62.0 0.8 | 56.8 0.6 | 61.6 1.2 | 61.4 1.3 | 44.2 0.5 | 71.2 1.0 | 97.5 0.2 | 27.5 0.5 |
| flat | 61.4 1.3 | 59.2 1.0 | 62.4 1.4 | 58.8 1.1 | 47.5 1.6 | 87.4 0.7 | 98.9 0.1 | 31.0 0.7 |
| volume | 60.4 0.8 | 58.8 1.1 | 64.3 1.6 | 59.1 1.2 | 53.4 0.5 | 86.3 0.7 | 96.2 0.3 | 27.7 0.7 |
| connectome | 59.8 | 57.0 | 58.6 | 58.0 | 45.6 | 81.9 | 82.4 | 7.4 |
| space | time | params | FLOPs | compute | data |
| parcel | 11 hr | 85M | 89G | 10K fps | 60K fps |
| flat | 28 hr | 86M | 92G | 9K fps | 4K fps |
| volume | 50 hr | 87M | 116G | 8K fps | 2K fps |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| config | value |
| optimizer | AdamW |
| momentum | |
| weight decay | 0.05 |
| learning rate | 1.25e-4 (flat, vol), 3.75e-5 (parcel) |
| lr schedule | cosine decay |
| warmup steps | 31K |
| config | value |
| optimizer | AdamW |
| momentum | |
| base learning rate | 3e-4 |
| base weight decay | 0.05 |
| lr scale grid | [0.02, 0.023, 0.028, 0.033, 0.038, 0.045, 0.053, 0.062, 0.074, 0.087, 0.1, 0.12, 0.14, 0.17, 0.2, 0.23, 0.27, 0.32, 0.38, 0.44, 0.52, 0.61, 0.72, 0.85, 1, 1.2, 1.4, 1.6, 1.9, 2.3, 2.7, 3.1, 3.7, 4.3, 5.1, 6, 7.1, 8.3, 9.8, 12, 14, 16, 19, 22, 26, 31, 36, 43, 50] |
| wd scale grid | [1.0] |
| subset | hours | Age | Sex | Task21 | COCO24 |
| all | 2058 | 47.5 | 87.4 | 98.9 | 31.0 |
| rest | 1072 | 47.7 | 85.5 | 97.5 | 30.8 |
| task | 987 | 47.4 | 85.4 | 98.5 | 28.7 |
| model | params | FLOP | data/s | fwd/s | FLOP/s |
| BrainLM | 113M | 382G | 267K | 96K | 73T |
| Brain-JEPA | 87M | 1511G | 338K | 86K | 261T |
| BrainHarmonix-F | 89M | 3135G | 310K | 20K | 122T |
| Brain-Semantoks | 63M | 6G | 179K | 1079K | 12.2T |
| CortexMAE-P | 85M | 8062G | 380K | 19K | 303T |
| SwiFT | 4M | 183G | 0.5K | 15K | 5.5T |