fMRI-TAMCL: Text-Anchored Supervised Multimodal Contrastive Learning for fMRI-Based Brain Disorder Classification
Organizations: Department of Computer Science, Texas State University, San Marcos, TX 78666 USA · University of Technology and Applied Sciences, Navrongo, Ghana
Abstract
Resting-state fMRI is important in the classification of brain disorders, but highly multimodal and exhibits strong multisite heterogeneity. Existing methods fuse images, BOLD-based functional connectivity, and phenotypic data modalities. Unlike other medical imaging datasets, rs-fMRI datasets rarely include a text modality, so they are generated from phenotypic data or BOLD activations. These text generation methods rely on fixed assumptions for subjects, sites, devices, and protocols, leading to poor generalization across datasets. We propose fMRI-TAMCL, a text-anchored multimodal contrastive learning framework that integrates fMRI images, sparse FC, and generated subject-specific text. Its Subject-Adaptive Threshold Derivation module generates BOLD activation text, while Feature-Value Serialization module generates phenotypic text. All three modalities are encoded as clustered graphs, projected onto a shared unit hypersphere space, aligned using pairwise, text-anchored supervised contrastive learning, and fused with attention. fMRI-TAMCL proves its generalization capability across five datasets outperforming 29 baselines with 78.6%-86.4% accuracy in downstream classification.
Figures & tables
| Method | Modalities | ABIDE | ADHD-200 | ADNI | HCP-EP | REST-meta-MDD | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ||
| Unimodal: (i) Functional Connectivity (FC-Only) Baselines | |||||||||||
| SVM | |||||||||||
| GCN [ 10 ] | |||||||||||
| GAT [ 11 ] | |||||||||||
| GraphSAGE [ 12 ] | |||||||||||
| Method | ACC (%) | SEN (%) | SPE (%) | AUC-ROC | F1 (%) |
|---|---|---|---|---|---|
| Unimodal | |||||
| BrainNetCNN [ 22 ] | 62.3 4.2 | 60.8 5.1 | 63.8 4.7 | 0.668 0.045 | 62.0 4.3 |
| BrainGNN [ 1 ] | 64.7 3.9 | 63.2 4.6 | 66.3 4.3 | 0.693 0.042 | 64.5 3.9 |
| BERT-Text | 66.4 2.9 | 64.9 3.8 | 67.1 3.7 | 0.718 0.031 | 65.7 3.1 |
| BrainNetTF | 67.9 2.9 | 66.8 2.2 | 69.1 2.8 | 0.731 0.021 | 67.9 2.3 |
| EC-GCL | 68.2 3.6 | 67.0 4.2 | 69.5 3.9 | 0.736 0.038 | 68.0 3.6 |
| Method | Miss Image ACC (%) | Miss FC ACC (%) | Miss Complete Text ACC (%) | Only Phenotypic/ Demographic Text ACC (%) | All Modalities ACC (%) | Avg Drop |
|---|---|---|---|---|---|---|
| TUMSyn | 71.56 | N/A | N/A | 74.6 | 78.9 | 5.8 |
| BrainPrompt+ | N/A | N/A | N/A | 73.8 | 78.1 | 4.3 |
| PMIL | N/A | 74.4 | 74.7 | N/A | 81.3 | 6.8 |
| fTSPL [ 26 ] | 75.6 | 73.1 | 73.9 | N/A | 80.8 | 6.6 |
| RTGMFF [ 27 ] | 77.2 | N/A | 75.7 | 76.3 | 81.5 | 5.1 |
| fMRI-TAMCL | 79.5 | 77.8 | 78.2 | 81.8 | 83.9 | 4.6 |
| Hyperparameter | ACC (%) | F1 |
|---|---|---|
| Contrastive loss weight sensitivity | ||
| Number of Graph Clusters | ||
| SATD Outcome | Count ( ) | Proportion (%) |
|---|---|---|
| Full convergence (2 crossing points) | 1,061 | 95.4 |
| Partial (1 crossing point; from mean scaling) | 34 | 3.1 |
| Fallback (percentile-based thresholds) | 17 | 1.5 |
| Multimodal Method | Parameters (M) | Train Time/ Epoch (s) | Inference Time/ Subject (ms) | (%) |
|---|---|---|---|---|
| TUMSyn | 15.1 | 78 | 15.1 | 78.9 |
| BrainPrompt+ | 17.8 | 62 | 13.2 | 78.1 |
| PMIL | 18.3 | 67 | 21.3 | 81.3 |
| fTSPL [ 26 ] | 18.7 | 74 | 19.9 | 80.8 |
| RTGMFF [ 27 ] | 20.7 | 70 | 17.9 | 81.5 |
| fMRI-TAMCL | 21.2 | 81 | 22.6 | 83.9 |