Natural Image Autoencoder-Based fMRI Representations for Trait and State Prediction
Organizations: IPAI, Seoul National University · Microsoft Research · ECE, Seoul National University · ASRI / INMC / AIIS, Seoul National University
Abstract
Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, which derives fMRI representations from a frozen Deep Compression AutoEncoder (DCAE) pre-trained exclusively on natural images and pairs them with a task specific readout. For trait prediction, FReD summarizes frame-wise representations by their temporal mean and log-standard deviation and applies linear probing, with late fusion across two normalization schemes. For state prediction, it represents each frame as a single token and models temporal dependencies with a shallow Transformer. Across four resting-state datasets spanning six trait-prediction targets, linear probes on frozen DCAE features generally outperform those on fMRI foundation model representations and remain competitive with fully fine-tuned fMRI foundation models. On three task-fMRI state-prediction tasks, a temporal readout on DCAE features performs comparably to the strongest foundation models evaluated. A Gaussian injection analysis further shows that localized signal changes are recovered more accurately from the frozen DCAE features than from the evaluated foundation-model representations. Together, these results show that strong performance on current fMRI benchmarks is possible without fMRI-specific representation pre-training, making frozen natural-image features as a useful baseline for assessing its added value.
Figures & tables
| Method | Input | ADHD-200 | ABIDE-II | ADNI | HCP-A | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Diagnosis | Diagnosis | Diagnosis | Sex | Age | Intelligence | ||||||||
| AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | MAE | MAE | ||||
| Brain-JEPA | |||||||||||||
| BH-F | |||||||||||||
| Brain-DiT | |||||||||||||
| SwiFT | |||||||||||||
| Method | Input | HBN-Movie | HBN-Sex | ||
|---|---|---|---|---|---|
| AUC | F1 | AUC | F1 | ||
| FReD-Trait | |||||
| FReD-Trait | |||||
| Method | Input | Eval | ADHD-200 | HCP-A | ||
| Diagnosis | Intelligence | |||||
| AUC | F1 | MAE | ||||
| FReD-Trait | LP | |||||
| FReD-Trait | LP | |||||
| FReD-State | TFS | |||||
| FReD-State | TFS | |||||
| Method | Input | Param. | HCP-Task | NSD |
|---|---|---|---|---|
| ACC | ACC | |||
| TABLeT | 129.5M | |||
| TABLeT with | 145.4M | |||
| FReD-State | 78.2M |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Target | Subjects | Samples | Duration | # Classes | Majority (%) |
|---|---|---|---|---|---|---|
| ADHD-200 | ADHD Dx | 430 | 430 | 300 s | 2 | 56.3 |
| ABIDE-II | ASD Dx | 616 | 616 | 300 s | 2 | 56.3 |
| ADNI | MCI Dx | 464 | 464 | 300 s | 2 | 59.5 |
| HCP-A | Sex | 1,069 | 1,069 | 382.4 s | 2 | 57.1 |
| HCP-A | Age, Intelligence | 1,069 | 1,069 | 382.4 s | – | – |
| HBN-Movie | Movie | 428 | 856 | 200 s | 2 | 50.0 |
| Method | Input | ADHD-200 | ABIDE-II | ADNI | HCP-A | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Diagnosis | Diagnosis | Diagnosis | Sex | Age | Intelligence | ||||||||
| AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | MAE | MAE | ||||
| Brain-JEPA | |||||||||||||
| BH-F | |||||||||||||
| Brain-DiT | |||||||||||||
| SwiFT | |||||||||||||
| Method | Input | Eval | ADHD-200 | ABIDE-II | ADNI | HCP-A | ||||||||
| Diagnosis | Diagnosis | Diagnosis | Sex | Age | Intelligence | |||||||||
| AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | MAE | MAE | |||||
| Brain-JEPA | LP | |||||||||||||
| Brain-JEPA | FT | |||||||||||||
| BH-F | LP | |||||||||||||
| BH-F | FT | |||||||||||||
| Method | Input | ADHD-200 | ABIDE-II | ADNI | HCP-A | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Diagnosis | Diagnosis | Diagnosis | Sex | Age | Intelligence | ||||||||
| AUC | F1 | AUC | F1 | AUC | F1 | AUC | F1 | MAE | MAE | ||||
| BrainMASS | |||||||||||||
| Brain-JEPA | |||||||||||||
| BH-F | |||||||||||||
| Brain-DiT | |||||||||||||
| Method | Eval. | Input | HBN-Movie | |
| AUC | F1 | |||
| BrainMASS | LP | |||
| Brain-JEPA | LP | |||
| Brain-JEPA | FT | |||
| BH-F | LP | |||
| BH-F | FT | |||
| Component | TABLeT | FReD-State with |
| Input representation and tokenization | ||
| DCAE latent input | Same input tensor | |
| Intra-frame aggregation | No slice averaging. Three contiguous 32-slice blocks are formed per axis, and corresponding features from the three anatomical axes are concatenated at each latent-grid position. | Contiguous slice-group averaging with groups per axis ( ), followed by concatenation of all axes, groups, channels, and spatial locations into one frame vector. |
| Tokens per frame | slice blocks spatial positions | whole-frame token |
| Pre-projection feature width | ||
| Frame projection | ||