Evidence Before Accuracy: A MRI-PET Fusion Network for Alzheimer Disease Classification with Causal Regional Validation
Organizations: Department of Computer Engineering and Information Technology, Payame Noor University, Tehran, Iran
Abstract
Deep learning models for Alzheimer disease (AD) classification routinely report near-perfect discrimination, yet few are shown to rest on AD-relevant neurobiology rather than on dataset artifacts, subject-level leakage, or non-brain image content. We present a fusion network combining T1 MRI and FDG PET across axial, coronal, and sagittal planes, trained on ADNI consists of 554 paired subjects. The fusion model reaches AUC 0.962, accuracy 0.909, and F1 0.891, competitive with recent 3D CNN and multimodal transformer systems at substantially lower cost. We first quantify how much modality, plane and slice geometry matter. A validation-only search over slice centres and neighbour spacings moves AUC by 0.180 for MRI and 0.078 for PET, selecting narrow spacing for MRI and wide spacing for PET, with the chosen coronal centres falling on the hippocampal body and on the posterior cingulate respectively. The contribution, however, is the evidence layer built around that number. Shortcut controls collapse the model to AUC 0.622 (silhouette), 0.608 (exterior), and 0.500 (blank), and a label-permutation null yields 0.456. Forward region-of-interest (ROI) ablation shows that masking medial temporal cortex in MRI and the posterior default-mode network (DMN) in PET produces the largest shift in the AD logit, while area-matched controls remain indistinguishable from that null. Reverse ROI ablation shows that the medial temporal lobe alone retains 89.2% of above-chance discrimination in MRI and the posterior DMN alone retains 79.0% in PET. A quantitative comparison of attribution methods shows occlusion sensitivity reaching 3.5-5.0* enrichment inside a priori AD regions against 0.10-0.43* in controls. Ablation and attribution independently establish a biologically correct double dissociation: hippocampal evidence is carried by MRI, posterior cingulate evidence by PET.
Figures & tables
| Group | MRI – PET series | Paired subjects | Train – Val – Test |
|---|---|---|---|
| CN | 12,198 – 4,348 | 324 | 208 – 52 – 64 |
| AD | 5,916 – 2,424 | 230 | 148 – 36 – 46 |
| All | 18,114 – 6,772 | 554 | 356 – 88 – 110 |
| Axial | Coronal | Sagittal | ||||||
|---|---|---|---|---|---|---|---|---|
| AUC | AUC | AUC | ||||||
| 0 | 0.704 | 0 | 0.884 | 0 | 0.758 | |||
| 4 | 0.782 | 4 | 0.871 | 4 | 0.781 | |||
| 8 | 0.753 | 8 | 0.814 | 8 | 0.826 | |||
| 14 | 0.742 | 14 | 0.848 | 14 | 0.798 | |||
| 0 | 0.741 | 0 | 0.809 | 0 | 0.848 | |||
| Axial | Coronal | Sagittal | ||||||
|---|---|---|---|---|---|---|---|---|
| AUC | AUC | AUC | ||||||
| 0 | 0.875 | 0 | 0.923 | 0 | 0.925 | |||
| 6 | 0.908 | 6 | 0.926 | 6 | 0.928 | |||
| 10 | 0.901 | 10 | 0.947 | 10 | 0.928 | |||
| 16 | 0.905 | 16 | 0.943 | 16 | 0.931 | |||
| 0 | 0.916 | 0 | 0.893 | 0 | 0.901 | |||
| Mode | AUC | Accuracy | F1 |
|---|---|---|---|
| Axial | 0.877 | 0.736 | 0.695 |
| Coronal | 0.891 | 0.764 | 0.735 |
| Sagittal | 0.897 | 0.845 | 0.813 |
| Fuse 3-plane | 0.909 | 0.818 | 0.778 |
| Fuse 3-plane single-plane init | 0.944 | 0.873 | 0.844 |
| Mode | AUC | Accuracy | F1 |
|---|---|---|---|
| Axial | 0.913 | 0.836 | 0.809 |
| Coronal | 0.900 | 0.818 | 0.800 |
| Sagittal | 0.918 | 0.827 | 0.796 |
| Fuse 3-plane | 0.928 | 0.827 | 0.796 |
| Fuse 3-plane single-plane init | 0.936 | 0.855 | 0.830 |
| Mode | AUC | Accuracy | F1 |
|---|---|---|---|
| MRI 3-plane | 0.944 | 0.873 | 0.844 |
| PET 3-plane | 0.936 | 0.855 | 0.830 |
| Fusion MRI + PET | 0.962 | 0.909 | 0.891 |
| Method | Modality | AUC | BACC |
|---|---|---|---|
| 3D-ResNet [ 30 ] | M | 0.936 | 0.865 |
| 3D-ResNet [ 30 ] | P | 0.951 | 0.890 |
| 3D-ViT [ 31 ] | M | 0.936 | 0.862 |
| 3D-ViT [ 31 ] | P | 0.935 | 0.887 |
| ResNet early fusion [ 32 ] | M + P | 0.855 | 0.826 |
| ResNet middle fusion [ 30 ] | M + P | 0.881 | 0.826 |
| Actual AD | Actual CN | |
|---|---|---|
| Predicted AD | TP | FP |
| Predicted CN | FN | TN |
| Mode | AUC |
|---|---|
| Without MRI | 0.932 |
| Without PET | 0.938 |
| All | 0.962 |
| Mode | AUC |
|---|---|
| MRI - with preprocessing | 0.944 |
| MRI - without preprocessing | 0.860 |
| PET - with preprocessing | 0.936 |
| PET - without preprocessing | 0.930 |
| Fusion - with preprocessing | 0.962 |
| Fusion - without preprocessing | 0.934 |
| Condition | AUC |
|---|---|
| Intact | 0.962 |
| Silhouette | 0.622 |
| Exterior | 0.608 |
| Blank | 0.500 |
| Condition | AUC |
|---|---|
| Intact | 0.962 |
| Label-permutation null | 0.456 |
| ROI masked out | Area (px) | AUC | logit AD | vs. null |
|---|---|---|---|---|
| Medial temporal | 4,361 | 0.038 | 0.283 | |
| Hippocampus | 1,844 | 0.015 | 0.087 | |
| Amygdala | 1,011 | 0.012 | 0.057 | |
| Parahippocampal | 1,551 | 0.006 | 0.078 | |
| Control composite | 4,297 | 0.005 | 0.000 | |
| Precentral gyrus | 3,218 | 0.002 | 0.002 |
| ROI masked out | Area (px) | AUC | logit AD | vs. null |
|---|---|---|---|---|
| Posterior DMN | 6,672 | 0.060 | 0.446 | |
| Precuneus | 1,878 | 0.016 | 0.199 | |
| Posterior cingulate | 1,728 | 0.011 | 0.274 | |
| Inferior parietal | 3,066 | 0.005 | 0.126 | |
| Control composite | 3,088 | 0.005 | 0.036 | |
| Precentral gyrus | 1,801 | 0.023 |
| ROI retained | Area (px) | AUC | Retained | |
|---|---|---|---|---|
| Medial temporal | 4,361 | 0.912 | 89.2% | |
| Amygdala | 1,011 | 0.877 | 81.8% | |
| Hippocampus | 1,844 | 0.852 | 76.2% | |
| Parahippocampal | 1,551 | 0.789 | 62.7% | |
| Control composite | 4,297 | 0.476 | % | |
| Control composite XL | 8,637 | 0.646 | 31.6% |
| ROI retained | Area (px) | AUC | Retained | |
|---|---|---|---|---|
| Posterior DMN | 6,672 | 0.865 | 79.0% | |
| Precuneus | 1,878 | 0.833 | 72.0% | |
| Posterior cingulate | 1,728 | 0.816 | 68.4% | |
| Inferior parietal | 3,066 | 0.784 | 61.5% | |
| Control composite | 3,088 | 0.593 | 20.1% | |
| Control composite XL | 4,716 | 0.561 | 13.2% |
| Stream | Medial temporal (hipp / amyg / parahipp) | Posterior DMN (post-cing / precuneus / inf-par) | Control (precentral / occipital / lingual) |
|---|---|---|---|
| MRI axial | 1.43 / 2.94 / 1.17 | n/a | — / 0.10 / 0.35 |
| MRI coronal | 3.54 / 4.17 / 3.54 | n/a | 0.34 / — / — |
| MRI sagittal | 3.22 / 2.99 / 2.65 | n/a | 0.78 / 0.55 / 1.69 |
| PET axial | n/a | 2.46 / 3.83 / 0.70 | 0.43 / 0.29 / — |
| PET coronal | 0.38 / — / 0.61 | 4.99 / 3.96 / 1.15 | — / — / 0.18 |
| PET sagittal | 1.96 / 1.20 / 1.51 | n/a | 0.11 / 1.08 / 3.26 |
| Method | Peak enrichment, AD regions | Enrichment, control regions | Pointing game (AD, best stream) | Usable |
|---|---|---|---|---|
| Occlusion | 54.3% (chance 16.7%) | Yes - primary evidence | ||
| Layer-CAM | 69.6% (chance 6.2%) | Directionally; weak AD–CN separation | ||
| Grad-CAM | 21.7% (chance 4.8%) | No | ||
| Grad-CAM++ | 8.7% (chance 16.7%) | No |
| Method | Stream | AD | CN | Chance |
|---|---|---|---|---|
| Occlusion | MRI axial | 39.1% | 17.2% | 6.7% |
| Occlusion | MRI coronal | 52.2% | 29.7% | 6.2% |
| Occlusion | MRI sagittal | 54.3% | 46.9% | 4.7% |
| Occlusion | PET axial | 54.3% | 40.6% | 16.7% |
| Occlusion | PET coronal | 54.3% | 23.4% | 13.1% |
| Layer-CAM | MRI coronal | 69.6% | 59.4% | 6.2% |