Learning robust representations from functional magnetic resonance imaging (fMRI) is fundamentally challenged by the temporal irregularity and noise inherent in data from heterogeneous sources. Existing self-supervised learning (SSL) methods often discard critical temporal information by discretizing or averaging fMRI signals. To address this, we introduce a novel framework that reframes SSL as a Stochastic Optimal Control (SOC) problem. Our approach models brain activity as continuous-time latent dynamics, learning a robust representation of brain dynamics by optimizing a control policy that is agnostic to the temporal irregularity. This SOC framework naturally unifies masked autoencoding (MAE) and joint-embedding prediction (JEPA) to extract compact, control-derived representations. Furthermore, a simulation-free inference strategy ensures computational efficiency and scalability for large-scale fMRI datasets. Our model demonstrates state-of-the-art performance across diverse downstream applications, highlighting the potential of the SOC-based continuous-time representation learning framework.
Figures & tables
Figure 1: Our BDO outperforms other self-supervised approaches, demonstrating superior efficiency.
Figure 2: Conceptual illustration of SSL approaches for fMRI data. (a) Image-based approach ( Caro et al., 2024 ; Dong et al., 2024 ) , where time-series data for each region-of-interest (ROI) is divided into fixed-size windows. (b) Graph-based approach ( Yang et al., 2024 ) , where static graphs are constructed to represent functional connectivity between ROIs. Both approaches discard high-resolution temporal dynamics during data preprocessing. (c) Our approach, leveraging continuous-time latent dynamics, directly captures the evolution of brain activity over time, making it robust to varying TRs across heterogeneous datasets and preserving fine-grained details.
Figure 3: Diverse TRs induce multi-scale irregularities on a unified time axis.
Figure 4: Conceptual illustration of our proposed Brain Dynamics with Optimal control (BDO) .
Figure 6: Overview of representation learning of BDO. Randomly masked fMRI time-series Yctx are encoded into latent states; encoded control signals {αtθ}t∈Tctx steer the SDE to predict latent states at the masked time points Ttar , which are then used to reconstruct the missing observations Ytar . A slowly updated EMA encoder provides latent targets, preventing representation collapse.
Methods
Age
Gender
MSE ↓
ρ↑
ACC (%) ↑
F1 (%) ↑
TS
BrainNetCNN
0.648 ±.018
0.621 ±.012
90.89 ±0.14
90.87 ±0.12
BrainGNN
0.914 ±.024
0.430 ±.010
79.07 ±1.08
79.03 ±1.09
BrainNetTF
0.561 ±.004
0.673 ±.003
91.19 ±0.51
91.17 ±0.50
LP
MoCo (90M)
0.933 ±.022
0.413 ±.010
80.11 ±0.73
80.11 ±0.73
BYOL (90M)
0.859 ±.006
0.380 ±.006
72.98 ±0.13
72.97 ±0.13
Table 1: Internal prediction on UKB 20% held-out.
Figure 7: (Left) Model scalability (Right) Data scalability on HCP-A age regression and diagnosis prediction.
Figure 8: (Left) Training curve (Right) HCP-A age regression Pearson correlation ρ as the mask ratio γ and balancing factor τ are varied.
NoTs
Age (MSE) ↓
Age ( ρ ) ↑
Gender (ACC) ↑
Gender (F1) ↑
80
0.587 ±.045
0.645 ±.026
63.67 ±2.00
61.98 ±0.92
160
0.404 ±.010
0.768 ±.008
72.00 ±2.95
71.30 ±2.19
240
0.348 ±.015
0.805 ±.009
72.22 ±1.13
71.34 ±1.35
Table 4: LP performance of BDO (86M) with a varying number of timesteps (NoTs) on HCP-A.
Figure 9: Projected 2D features of the universal feature A using PCA and UMAP, coloured by age. (Left) Embedding space of UKB held-out split. (Right) Embedding space of HCP-A dataset.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12
BDO Variants
Train EP
Warm-up EP
LR
Initial LR
Minimum LR
Batch Size
Rd
# of base matrices (L)
EMA Momentum
BDO (5M)
200
10
0.001
0.0001
0.0001
128
192
100
[0.996, 1]
BDO (21M)
200
10
0.001
0.0001
0.0001
128
384
100
[0.996, 1]
BDO (86M)
200
10
0.001
0.0001
0.0001
128
768
100
[0.996, 1]
Appendix
Table 5: Pre-training hyper-parameters
Figure 14
Category
UKB
HCP-A
ABIDE
ADHD200
HCP-EP
# of subjects
41,072
724
1,102
669
176
Age, mean (SD)
54.98 (7.53)
60.35 (15.74)
17.05 (8.04)
11.61 (2.97)
23.39 (3.95)
Female, % (n)
52.30 (21,480)
56.08 (406)
14.79 (163)
36.17 (242)
38.07 (67)
Patient, % (n)
-
-
48.19 (531)
58.15 (389)
68.18 (120)
Target Population
Healthy Population
Healthy Population
ASD Healthy Population
ADHD Healthy Population
Psychotic Disorder Healthy Population
Appendix
Table 6: Dataset Subject Demographics
Configurations
FT
LP
Optimizer
AdamW ( Loshchilov, 2017 )
Adam ( Kingma and Ba, 2015 )
Training epochs
50
50
Batch size
[16,32]
[16,32,64]
LR scheduler
cosine decay
cosine decay
LR
[0.001]
[0.01,0.005]
Minimum LR
[0,0.0001,0.001]
[0.001,0.005]
Appendix
Table 7: Search space of end-to-end fine-tuning (FT) and linear probe (LP).
Dataset
Age (MSE) ↓
Age (Pearson) ↑
Gender (Acc.) ↑
Gender (F1) ↑
HCP-A-VisMotor
0.526 ±.018
0.691 ±.015
68.53 ±3.57
67.39 ±3.36
HCP-A-FaceName
0.459 ±.012
0.732 ±.009
66.20 ±3.44
65.29 ±3.72
HCP-A-CARIT
0.488 ±.025
0.713 ±.020
67.60 ±1.74
66.79 ±1.29
HCP-A-Rest
0.404 ±.010
0.768 ±.008
72.00 ±2.95
71.30 ±2.19
Appendix
Table 13: Generalization from resting-state to task-based fMRI.
Variants
HCP-A
ABIDE
ADHD200
HCP-EP
Age ( ρ↑ )
ACC (%) ↑
ACC (%) ↑
ACC (%) ↑
BDO (5M)
0.635 ±.031
62.42 ±2.68
59.65 ±2.30
73.33 ±7.50
BDO (21M)
0.729 ±.011
63.79 ±1.83
61.15 ±1.97
71.43 ±4.04
BDO (25%)
0.686 ±.010
61.06 ±1.05
57.39 ±3.90
72.38 ±5.95
BDO (50%)
0.702 ±.014
63.03 ±1.63
56.89 ±3.38
74.29 ±9.90
BDO (75%)
0.734 ±.011
65.45 ±2.70
58.15 ±1.78
74.29 ±7.56
Appendix
Table 14: LP performance used for scalability analysis in Figure 7 .
Model (Parameters)
Age (Pearson) ↑
Gender (Acc.) ↑
GPU Hours (x 4 GPUs) ↓
MoCo (90M)
0.591
64.12
174 hrs
BYOL (90M)
0.619
64.81
165 hrs
BrainLM (85M)
0.636
65.28
496 hrs
BrainMass (90M)
0.630
66.20
244 hrs
BDO (86M)
0.768
72.00
15 hrs
Appendix
Table 16: Comparison of pre-training efficiency and linear probing performance across SSL models.
Mask Ratio ( γ )
Age (MSE) ↓
Age (Pearson) ↑
0.0 (No Masking)
0.793 ± 0.014
0.445 ± 0.020
0.2
0.487 ± 0.036
0.711 ± 0.027
0.4
0.513 ± 0.016
0.695 ± 0.015
0.6
0.476 ± 0.019
0.727 ± 0.011
0.75 (Optimal)
0.466 ± 0.025
0.738 ± 0.014
0.8
0.526 ± 0.014
0.686 ± 0.006
Appendix
Table 17: Extended ablation study on the mask ratio γ .
Age (MSE) ↓
Age (Pearson) ↑
JEPA-only
0.719 ± 0.040
0.521 ± 0.036
τ=0
0.480 ± 0.010
0.717 ± 0.006
τ=0.03
0.466 ± 0.025
0.738 ± 0.014
τ=0.1
0.663 ± 0.027
0.572 ± 0.028
Appendix
Table 18: Ablation study on the MAE and JEPA components.
Figure 10: Brain surface visualization of IG scores. (Left) Age regression on the HCP-A dataset. (Right) Psychotic disorder diagnosis on the HCP-EP dataset.
Rank
Yeo-7 Network Label
IG Score
AAL Atlas Label
1
7Networks_LH_SomMot_26
0.0074
Precentral_L
2
7Networks_LH_Default_PFC_13
0.0061
Frontal_Sup_Medial_L
3
7Networks_RH_SalVentAttn_TempOccPar_6
0.0060
SupraMarginal_R
4
7Networks_RH_Default_Par_5
0.0059
Angular_R
5
7Networks_RH_Vis_19
0.0057
Calcarine_R
6
7Networks_RH_SalVentAttn_TempOccPar_5
0.0056
SupraMarginal_R
Appendix
Table 19: Top 10 ROIs with the highest IG scores in the HCP-A age prediction task.
Rank
Yeo-7 Network Label
IG Score
AAL Atlas Label
1
7Networks_RH_SomMot_16
0.0082
Postcentral_R
2
7Networks_RH_Vis_29
0.0056
Cuneus_R
3
7Networks_LH_Vis_29
0.0049
Occipital_Sup_L
4
7Networks_LH_Vis_27
0.0047
Occipital_Mid_L
5
7Networks_LH_SomMot_36
0.0047
Postcentral_L
6
7Networks_RH_SalVentAttn_PFCl_1
0.0046
Frontal_Mid_2_R
Appendix
Table 20: Top 10 ROIs with the highest IG scores in the HCP-EP diagnosis prediction task.
Figure 11: Qualitative analysis of learned latent dynamics via step-wise latent displacement in ADHD200. The plot compares the distribution of step-wise latent displacements ( Δz ) between Healthy Controls and ADHD patients. The observed lower magnitude of Δz in the ADHD group indicates less fluid state transitions.
Figure 12: Age distribution across training, validation, and test splits for the UKB held-out age regression task under three different random seeds (0, 1, and 2). The dataset is partitioned using a 6:2:2 ratio, with binning-based stratified sampling applied to maintain a balanced target variable distribution. To enhance numerical stability, Z-score normalization is applied to the age variable. Each row represents a different random seed, illustrating the consistency of the sampling procedure across splits.
Figure 13: Age distribution across training, validation, and test splits for the HCP-A age regression task under three different random seeds (0, 1, and 2). The dataset is partitioned using a 6:2:2 ratio, with binning-based stratified sampling applied to maintain a balanced target variable distribution. To enhance numerical stability, Z-score normalization is applied to the age variable. Each row represents a different random seed, illustrating the consistency of the sampling procedure across splits.
Figure 14: Label distributions across six classification tasks (UKB held-out gender, HCP-A gender, ABIDE autism, ADHD200 ADHD, and HCP-EP psychotic disorder) for training, validation, and test splits. Each row corresponds to a different task, with columns representing the proportion of samples per class across data splits. Stratified sampling ensures that label distributions remain consistent across splits, despite variations in sample composition. To illustrate this, we visualize the distributions using a single random seed (0). Gender classification tasks are divided into Female/Male categories, while disease classification tasks distinguish between Control and Patient groups (ASD vs. Control for ABIDE, ADHD vs. Control for ADHD200, and Psychotic disorder vs. Control for HCP-EP).
Figure 15: Reconstruction quality of BDO in the UKB held-out subset (internal dataset). Five samples are randomly drawn for visualization, with a mask ratio of γ=0.75 . Each column represents the original fMRI sample, context with masking patterns, reconstructed sample, and MAE (Mean Absolute Error) heatmaps. Although we set the mask ratio as high as 75% , the reconstruction quality remains robust, demonstrating that BDO efficiently captures the underlying brain dynamics and successfully reconstructs missing regions with high fidelity.
Figure 16: Reconstruction quality of BDO in HCP-A (external dataset). Five samples are randomly drawn for visualization, with a mask ratio of γ=0.75 . Each column represents the original fMRI sample, context with masking patterns, reconstructed sample, and MAE (Mean Absolute Error) heatmaps. Although we set the mask ratio as high as 75% , the reconstruction quality remains robust, demonstrating that BDO efficiently captures the underlying brain dynamics and successfully reconstructs missing regions with high fidelity.
Functional magnetic resonance imaging (fMRI) is a powerful tool for investigating human brain function. However, the high cost of data acquisition and the inherent subjectivity of psychiatric rating scales often lead to datasets with small sample sizes and variable label quality, especially when targeting a specific neurological condition. Combined with the inherently high dimensionality of fMRI data, these limitations substantially increase the risk of model overfitting. Recent years have seen growing interest in developing fMRI foundation models by combining multiple datasets; however, the computational resources needed for pretraining and fine-tuning are often prohibitive. We show that a lightweight self-supervised framework yields representations that generalize across diverse downstream tasks, outperforming fully supervised baselines and approaching the performance of large-scale models. We introduce BrainSimSiam, a data-efficient self-supervised representation learning framework that leverages positive-only data pairs to learn robust and generalizable features. We demonstrate that the learned representations achieve strong performance across multiple downstream classification and regression tasks, highlighting the potential of BrainSimSiam for data-limited neuroimaging applications.
Jiyao Wang, Peiyu Duan, Nicha C. Dvornek +4
Department of Biomedical Engineering, Yale University, New Haven, CT, USA · Radiology & Biomedical Imaging, Yale School of Medicine, New Haven, CT, USA · Electrical Engineering, Yale University, New Haven, CT, USA +1
Masked autoencoders (MAEs) have recently shown promise for self-supervised representation learning of resting-state brain functional connectivity (FC). However, a fundamental question remains unresolved: how should FC matrices be tokenized to align with the intrinsic modular organization of large-scale brain networks? Existing approaches typically adopt region-centric or graph-based schemes that treat FC as structurally homogeneous elements and overlook the large-scale network brain organization. We introduce NERVE (Network-Aware Representations of Brain Functional Connectivity via Bilinear Tokenization), a self-supervised learning framework that redefines FC tokenization by partitioning FC matrices into patches of intra- and inter-network connectivity blocks. Unlike image-based MAE, where fixed-size patches share a common tokenizer, FC patches defined by network pairs are heterogeneous in size and correspond to distinct functional roles. To resolve this problem, NERVE embeds FC patches through a novel structured bilinear factorization. This formulation preserves network identity and reduces parameter complexity from quadratic to linear scaling in the number of networks. We evaluate NERVE across three large-scale developmental cohorts (ABCD, PNC, and CCNP) for behavior and psychopathology prediction. Compared to structurally agnostic MAE variants and graph-based self-supervised baselines, the proposed network-aware formulation yields more stable and transferable representations, particularly in cross-cohort evaluation. Ablation studies confirm that the proposed bilinear network embedding and anatomically grounded parcellation are critical for performance. These findings highlight the importance of incorporating domain-specific structural priors into self-supervised learning for functional connectomics. Code is available at: https://github.com/leomlck/NERVE.
Leo Milecki, Qingyu Hu, Bahram Jafrasteh +2
Department of Radiology, Weill Cornell Medicine, New York, NY, USA. · School of Electrical and Computer Engineering, Cornell University and Cornell Tech, New York, NY, USA.
Modeling resting-state functional magnetic resonance imaging (rs-fMRI) data is crucial for understanding brain-wide neural activity. However, traditional methods struggle to capture complex temporal dynamics over long horizons, to account for the brain's anatomical spatial structure, and to model high-dimensional ambient signals that lie on a low-dimensional intrinsic subspace. We propose FAST-Brain, a unified flow-aligned spatio-temporal surrogate brain model that addresses all three challenges. At its core is a flow-aligned generative framework that directly predicts the clean blood-oxygen-level-dependent (BOLD) signal, paired with a graph convolutional network that captures spatial structural constraints and a Transformer that models long-range temporal dependencies. Theoretically, we show that under a low-dimensional subspace assumption, the approximation error of our model scales with the intrinsic dimension rather than the ambient dimension, which justifies our direct modeling of the BOLD signal. Extensive experiments on synthetic and Human Connectome Project datasets demonstrate that FAST-Brain achieves state-of-the-art performance in recovering functional connectivity, effective connectivity, and the implicit low-dimensional signal subspace.
Shucheng Liu, Changchun Shi, Kai Zhang +1
University of North Carolina at Chapel Hill · London School of Economics and Political Science