Stochastic Optimal Control for Continuous-Time fMRI Representation Learning
Organizations: KAIST · Yonsei University · EverEx · AITRICS
Abstract
Learning robust representations from functional magnetic resonance imaging (fMRI) is fundamentally challenged by the temporal irregularity and noise inherent in data from heterogeneous sources. Existing self-supervised learning (SSL) methods often discard critical temporal information by discretizing or averaging fMRI signals. To address this, we introduce a novel framework that reframes SSL as a Stochastic Optimal Control (SOC) problem. Our approach models brain activity as continuous-time latent dynamics, learning a robust representation of brain dynamics by optimizing a control policy that is agnostic to the temporal irregularity. This SOC framework naturally unifies masked autoencoding (MAE) and joint-embedding prediction (JEPA) to extract compact, control-derived representations. Furthermore, a simulation-free inference strategy ensures computational efficiency and scalability for large-scale fMRI datasets. Our model demonstrates state-of-the-art performance across diverse downstream applications, highlighting the potential of the SOC-based continuous-time representation learning framework.
Figures & tables
| Methods | Age | Gender | |||
|---|---|---|---|---|---|
| MSE | ACC (%) | F1 (%) | |||
| TS | BrainNetCNN | 0.648 ±.018 | 0.621 ±.012 | 90.89 ±0.14 | 90.87 ±0.12 |
| BrainGNN | 0.914 ±.024 | 0.430 ±.010 | 79.07 ±1.08 | 79.03 ±1.09 | |
| BrainNetTF | 0.561 ±.004 | 0.673 ±.003 | 91.19 ±0.51 | 91.17 ±0.50 | |
| LP | MoCo (90M) | 0.933 ±.022 | 0.413 ±.010 | 80.11 ±0.73 | 80.11 ±0.73 |
| BYOL (90M) | 0.859 ±.006 | 0.380 ±.006 | 72.98 ±0.13 | 72.97 ±0.13 | |
| NoTs | Age (MSE) | Age ( ) | Gender (ACC) | Gender (F1) |
|---|---|---|---|---|
| 80 | 0.587 ±.045 | 0.645 ±.026 | 63.67 ±2.00 | 61.98 ±0.92 |
| 160 | 0.404 ±.010 | 0.768 ±.008 | 72.00 ±2.95 | 71.30 ±2.19 |
| 240 | 0.348 ±.015 | 0.805 ±.009 | 72.22 ±1.13 | 71.34 ±1.35 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| BDO Variants | Train EP | Warm-up EP | LR | Initial LR | Minimum LR | Batch Size | # of base matrices (L) | EMA Momentum | |
|---|---|---|---|---|---|---|---|---|---|
| BDO (5M) | 200 | 10 | 0.001 | 0.0001 | 0.0001 | 128 | 192 | 100 | [0.996, 1] |
| BDO (21M) | 200 | 10 | 0.001 | 0.0001 | 0.0001 | 128 | 384 | 100 | [0.996, 1] |
| BDO (86M) | 200 | 10 | 0.001 | 0.0001 | 0.0001 | 128 | 768 | 100 | [0.996, 1] |
| Category | UKB | HCP-A | ABIDE | ADHD200 | HCP-EP |
|---|---|---|---|---|---|
| # of subjects | 41,072 | 724 | 1,102 | 669 | 176 |
| Age, mean (SD) | 54.98 (7.53) | 60.35 (15.74) | 17.05 (8.04) | 11.61 (2.97) | 23.39 (3.95) |
| Female, % (n) | 52.30 (21,480) | 56.08 (406) | 14.79 (163) | 36.17 (242) | 38.07 (67) |
| Patient, % (n) | - | - | 48.19 (531) | 58.15 (389) | 68.18 (120) |
| Target Population | Healthy Population | Healthy Population | ASD Healthy Population | ADHD Healthy Population | Psychotic Disorder Healthy Population |
| Configurations | FT | LP |
|---|---|---|
| Optimizer | AdamW ( Loshchilov, 2017 ) | Adam ( Kingma and Ba, 2015 ) |
| Training epochs | ||
| Batch size | ||
| LR scheduler | cosine decay | cosine decay |
| LR | ||
| Minimum LR |
| Dataset | Age (MSE) | Age (Pearson) | Gender (Acc.) | Gender (F1) |
|---|---|---|---|---|
| HCP-A-VisMotor | 0.526 ±.018 | 0.691 ±.015 | 68.53 ±3.57 | 67.39 ±3.36 |
| HCP-A-FaceName | 0.459 ±.012 | 0.732 ±.009 | 66.20 ±3.44 | 65.29 ±3.72 |
| HCP-A-CARIT | 0.488 ±.025 | 0.713 ±.020 | 67.60 ±1.74 | 66.79 ±1.29 |
| HCP-A-Rest | 0.404 ±.010 | 0.768 ±.008 | 72.00 ±2.95 | 71.30 ±2.19 |
| Variants | HCP-A | ABIDE | ADHD200 | HCP-EP | |
|---|---|---|---|---|---|
| Age ( ) | ACC (%) | ACC (%) | ACC (%) | ||
| BDO (5M) | 0.635 ±.031 | 62.42 ±2.68 | 59.65 ±2.30 | 73.33 ±7.50 | |
| BDO (21M) | 0.729 ±.011 | 63.79 ±1.83 | 61.15 ±1.97 | 71.43 ±4.04 | |
| BDO (25%) | 0.686 ±.010 | 61.06 ±1.05 | 57.39 ±3.90 | 72.38 ±5.95 | |
| BDO (50%) | 0.702 ±.014 | 63.03 ±1.63 | 56.89 ±3.38 | 74.29 ±9.90 | |
| BDO (75%) | 0.734 ±.011 | 65.45 ±2.70 | 58.15 ±1.78 | 74.29 ±7.56 |
| Model (Parameters) | Age (Pearson) | Gender (Acc.) | GPU Hours (x 4 GPUs) |
|---|---|---|---|
| MoCo (90M) | 0.591 | 64.12 | 174 hrs |
| BYOL (90M) | 0.619 | 64.81 | 165 hrs |
| BrainLM (85M) | 0.636 | 65.28 | 496 hrs |
| BrainMass (90M) | 0.630 | 66.20 | 244 hrs |
| BDO (86M) | 0.768 | 72.00 | 15 hrs |
| Mask Ratio ( ) | Age (MSE) | Age (Pearson) |
|---|---|---|
| 0.0 (No Masking) | 0.793 ± 0.014 | 0.445 ± 0.020 |
| 0.2 | 0.487 ± 0.036 | 0.711 ± 0.027 |
| 0.4 | 0.513 ± 0.016 | 0.695 ± 0.015 |
| 0.6 | 0.476 ± 0.019 | 0.727 ± 0.011 |
| 0.75 (Optimal) | 0.466 ± 0.025 | 0.738 ± 0.014 |
| 0.8 | 0.526 ± 0.014 | 0.686 ± 0.006 |
| Age (MSE) | Age (Pearson) | |
|---|---|---|
| JEPA-only | 0.719 ± 0.040 | 0.521 ± 0.036 |
| 0.480 ± 0.010 | 0.717 ± 0.006 | |
| 0.466 ± 0.025 | 0.738 ± 0.014 | |
| 0.663 ± 0.027 | 0.572 ± 0.028 |
| Rank | Yeo-7 Network Label | IG Score | AAL Atlas Label |
|---|---|---|---|
| 1 | 7Networks_LH_SomMot_26 | 0.0074 | Precentral_L |
| 2 | 7Networks_LH_Default_PFC_13 | 0.0061 | Frontal_Sup_Medial_L |
| 3 | 7Networks_RH_SalVentAttn_TempOccPar_6 | 0.0060 | SupraMarginal_R |
| 4 | 7Networks_RH_Default_Par_5 | 0.0059 | Angular_R |
| 5 | 7Networks_RH_Vis_19 | 0.0057 | Calcarine_R |
| 6 | 7Networks_RH_SalVentAttn_TempOccPar_5 | 0.0056 | SupraMarginal_R |
| Rank | Yeo-7 Network Label | IG Score | AAL Atlas Label |
|---|---|---|---|
| 1 | 7Networks_RH_SomMot_16 | 0.0082 | Postcentral_R |
| 2 | 7Networks_RH_Vis_29 | 0.0056 | Cuneus_R |
| 3 | 7Networks_LH_Vis_29 | 0.0049 | Occipital_Sup_L |
| 4 | 7Networks_LH_Vis_27 | 0.0047 | Occipital_Mid_L |
| 5 | 7Networks_LH_SomMot_36 | 0.0047 | Postcentral_L |
| 6 | 7Networks_RH_SalVentAttn_PFCl_1 | 0.0046 | Frontal_Mid_2_R |