Content-based image retrieval (CBIR) in neuroimaging enables the identification of structurally similar brain scans, supporting diagnosis, prognosis, and treatment planning; however, existing methods are often limited to small datasets, single brain regions, or coarse class labels, thereby restricting their clinical utility and generalizability. Here, we present NeuroCBIR, a framework for fast and flexible retrieval of both whole-brain and region-specific 3D T1w MRI scans. A total of 103 cortical and subcortical regions are extracted to enable both whole-brain and region-level queries. NeuroCBIR leverages latent representations learned by a variational autoencoder (VAE) combined with contrastive learning, producing scan-specific embeddings that capture anatomical patterns. These embeddings were evaluated for subject re-identification, zero-shot age prediction, and zero-shot multi-class pathology stratification. Re-identification performance was high across both whole-brain and brain-region levels (mean average precision across the top-5 retrieved images (mAP@5) >= 98.4%), with robust generalization across datasets and acquisition conditions. While NeuroCBIR is not trained for age prediction or pathology stratification, zero-shot evaluations for these two tasks demonstrate that the embeddings encode meaningful information for downstream tasks. Embedding extraction on a 4-core CPU required approximately 18.7 s per scan, whereas similarity search was effectively instantaneous (less than 0.01 s). NeuroCBIR is publicly available for brain MRI with more than 26,000 precomputed T1w MRI embeddings. It supports reproducible research, region-specific flexibility, and clinically meaningful personalized diagnostic support. The software is available at https://github.com/minnelab/NeuroCBIR.
Figures & tables
Reference
Dataset / Task
Method
Evaluation
Performance
Huang et al. (2012)
2D T1w MRI (glioma, meningioma, pituitary)
Bag-of-visual-words + metric learning
Compared tumor types Hit = retrieved same tumor class.
Prec@10 ≈94% mAP ≈91%
Onga et al. (2019)
3D T1w MRI ADNI2 AD / EMCI / LMCI / SMCI / CN
CAE + metric learning
Clustering and classification across diagnostic groups
Table 1: Summary of recent literature on brain MRI retrieval and classification.
Dataset
NI
NS
NR
NT
Age
Class
Field strength
Manufacturer
CN
MCI
AD
1.5T
3T
SH
GE
PH
ADNI
20367
2389
2.0
4.1
75 ± 7
6260
11572
2524
7951
8779
7453
3563
1794
OASIS3
2642
1303
1.0
2.0
71 ± 9
2149
392
100
208
2429
2642
0
0
AIBL
1276
685
1.0
1.8
74 ± 7
938
179
152
266
1010
1276
0
0
MIRIAD
706
69
1.3
7.6
70 ± 7
243
0
463
706
0
0
706
0
SLIM
1015
571
1.0
1.8
21 ± 1
1015
0
0
0
1015
1015
0
0
Table 2: Summary of the MRI datasets used in this study.
Figure 1: Pairwise FAED between MRI datasets stratified by MRI field strength (1.5T and 3.0T). Lower values indicate greater similarity in feature distributions, while higher values reflect larger domain shifts. White cells correspond to uncomputed comparisons. The comparisons were done with the fully-preprocessed brain images.
Figure 2: Overview of the NeuroCBIR framework. (A) VAE training: the encoder ( Eϕ ) maps input images to mean ( μϕ ) and variance ( σϕ ), which define the latent variable ( uϕ ). The decoder ( Dθ ) reconstructs the image, optimized using an L1 loss, latent regularization ( LKL ), perceptual loss, and an adversarial loss with critic ( Cψ ). (B) CL training: the latent representations, i.e. the means μϕ from the pretrained VAE encoder, are passed through an encoder–projector module (Eω′,Pω) , yielding compact embeddings ( zω ). Similar (intra-subject) embeddings are pulled together, while dissimilar (inter-subject) embeddings are pushed apart by the criteria specified for the Multi-Positive Ranking Contrastive Loss ( LMPRCL ). We denote by Ii,j the input image corresponding to subject S and index R . (C) Inference: the frozen Q2E module maps a query image to its embedding, which is compared against the precomputed database { zω,1,1 , zω,1,2 , …, zω,S,R−1 , zω,S,R } to rank and retrieve the top- k most similar images.
Figure 3: Example of input and target images for whole-brain and brain-region data. We denote by IS,R the input image corresponding to subject S and index R (i.e., time point). I1,1 – I1,4 show four images for the subject 1, and I2,1 – I2,2 show two images for the subject 2. The brain-region images illustrate the hippocampus.
Figure 4: Illustration of the proposed Multi-Positive Ranking Contrastive Loss (MPRCL). For a given anchor query, hard positives are defined by shared identity labels, while soft and hard negatives are selected using percentile thresholds applied to the VAE latent similarity. Soft negatives correspond to samples whose VAE similarity exceeds the upper percentile πsoft , whereas negatives fall below the lower percentile πhn . In the embedding space, the loss enforces margin-based ranking constraints that encourage hard positives to be closest to the anchor, soft negatives to occupy an intermediate similarity range, and negatives to be pushed further away. The margins are adaptively scaled by the anchor-specific similarity range, using scaling factors αid for identity-based constraints and αsoft for soft-negative versus negative constraints.
Nf
αhn
λsn
αsn
πhn
λhn
mAP@5 ↑
ρMS−SSIM↓
16
0.5
1.0
0.05
30
0.10
94.7
−0.72
16
0.2
1.0
0.05
30
0.10
95.3
−0.71
32
0.5
1.0
0.05
30
0.10
95.2
−0.70
32
0.2
1.0
0.05
30
0.10
96.9
−0.74
32
0.2
0.2
0.05
30
0.10
97.0
−0.74
32
0.2
0.2
0.10
30
0.10
97.6
−0.73
Table 3: Hyperparameter optimization results for the whole-brain approach. For the validation partition: retrieval quality is assessed using mAP@5, while embedding quality is evaluated via Spearman’s rank correlation between embedding similarities and MS-SSIM ( ρMS−SSIM ). The last three rows correspond to ablations of the selected model ( ⋆ ).
Metric
All
Train
Val
Test
ADNI
OASIS3
AIBL
MIRIAD
SLIM 1
mAP@5 ( ↑ )
99.2
99.6
98.7
98.4
99.3
92.5
100.0
99.9
-
mS@1 ( ↑ )
98.8
99.4
99.0
97.2
99.9
93.6
94.9
100.0
89.1
mS@5 ( ↑ )
99.9
100.0
99.8
99.8
99.9
99.4
100.0
100.0
-
ρMS−SSIM ( ↓ )
−0.71
−0.64
−0.71
−0.74
−0.71
−0.65
−0.75
−0.77
−0.79
Table 4: Retrieval re-identification performance (mAP@5, mS@1, mS@5) and Spearman’s rank correlation between embedding similarities and MS-SSIM ( ρMS−SSIM ) for the whole-brain NeuroCBIR framework across multiple datasets.
Figure 5: Retrieval examples of T1w brain MRI scans. A.1 belongs to a subject with 7 MRI scans. B.1 are from a subject with only 3 scans. C.1 belongs to a patient with only two scans. The blue frame denotes the queries, the green frames are for the "hits" (same subject scans), and the red frames are for the different subject scans that are deemed to be similar to the query by NeuroCBIR. Except for A.12, images from the same subject are ranked higher by NeuroCBIR.
Figure 6: mAP@5 and mS@5 for the testing partition of the top 20 best / worst brain structures sorted by mAP@5. Abbreviations: CC = Corpus Callosum; Mid = Middle; lh/rh = left/right hemisphere; ctx = cortex.
Metric
Structure
All
Train
Val
Test
ADNI
OASIS3
AIBL
MIRIAD
SLIM 1
mAP@5 ( ↑ )
Left-Hippocampus
98.9
99.3
98.1
98.3
99.0
92.7
100.0
100.0
-
Left-Thalamus
96.8
98.5
93.2
93.6
97.0
81.5
100.0
99.5
-
Left-Amygdala
86.0
91.4
73.8
75.5
85.8
72.2
96.7
96.3
-
Left-Lateral-Ventricle
99.6
99.7
99.4
99.3
99.6
96.0
100.0
100.0
-
mS@1 ( ↑ )
Left-Hippocampus
98.8
99.1
99.1
97.7
99.7
92.5
97.5
100.0
91.6
Left-Thalamus
95.9
97.5
96.7
91.5
98.7
80.6
88.0
100.0
69.3
Table 5: Retrieval re-identification performance (mAP@5, mS@1, mS@5) and Spearman’s rank correlation between embedding similarities and MS-SSIM ( ρMS−SSIM ) for different brain-regions of the NeuroCBIR framework across multiple datasets. The full detailed results are provided in A .
Figure 7: Retrieval of similar left hippocampus segmentations. A.1 belongs to a subject with at least 6 MRI scans. B.1 are from a subject with only 5 scans (B1.1–B1.5). C.1 belongs to a patient with no additional MRI scans. The blue frame denotes the queries, the green frames are for the "hits", and the red frames are for the "fails".
Figure 8: Coefficient of Variation (CV) of Spearman correlations between whole-brain similarity and region-based similarity across top- k retrieved cases, excluding scans from the query subject. Left panel shows variability across regions for each query, and right panel shows variability across queries for each region. Top- k groups of 50, 200, and 1000 are shown. The CV is plotted on a logarithmic scale to highlight differences across a wide range of values. Points indicate individual observations, while horizontal markers show the mean CV for each top- k group. Higher CV indicates greater heterogeneity in the contribution of regions or queries to whole-brain similarity.
Comparison
p -value
ssd
Cliff’s δ
Magnitude
Test vs. Train
0.000000
−0.027000
negligible
CN vs. MCI
0.003080
−0.006000
negligible
CN vs. AD
0.685000
−0.004000
negligible
MCI vs. AD
1.000000
0.002000
negligible
ADNI vs. AIBL
1.000000
−0.009000
negligible
ADNI vs. MIRIAD
0.338000
−0.006000
negligible
Table 6: Statistical comparison of retrieval performance across datasets, scanner field strengths, and vendors. Reported values include p -values (Mann–Whitney U test), significance indicators (ssd), Cliff’s δ effect sizes, and their qualitative magnitudes. Asterisks (*) denote significant differences at p -value<0.05.
Method
Comp.
Nf↓
mAP@5 ↑
mS@5 ↑
ρMS−SSIM↓
ResNet-10
2048 :1
512
7.2
32.6
−0.55
ResNet-50
512 :1
2048
10.8
43.9
−0.70
ResNet-101
512 :1
2048
5.7
30.0
−0.51
CNN-AD
2628 :1
256
0.1
0.7
−0.01
VoComni-L
576 :1
1536
14.3
50.9
−0.34
BrainIAC
1152 :1
768
72.4
95.8
−0.27
Table 7: Comparison of different feature extractors and dimensionality reduction approaches on the retrieval task using the top-5 results from the whole-brain testing set. Reported metrics include the compression rate (Comp.), number of features ( Nf ), mean precision, and success rate.
Step
RAM (MB) (↓)
Time (s) (↓)
baseline
peak
1 core
4 cores
Whole-brain
Load Q2E
862
1087
8.7
6.5
Run Q2E
1087
6804
66.4
18.7
Top 5 retrieval
-
-
<0.01
<0.01
Brain-regions
Table 8: Execution time and memory usage for whole-brain and brain-region processing. Baseline refers to the RAM usage before running, peak refers to the peak of RAM measured during the execution.
Region
MAE (↓)
r(↑)
Slope
Whole-brain
Random baseline
5.2 ± 3.5
0.00
0.00
ResNet-10
4.5 ± 3.5
0.46
0.32
ResNet-50
4.8 ± 3.6
0.35
0.16
ResNet-101
5.0 ± 3.6
0.27
0.13
CNN-AD
5.8 ± 4.0
0.00
0.00
Table 9: Zero-shot similarity-based brain age estimation across selected brain regions. For each region, the predicted age was computed using the similarity-based weighted average of the top- 500 retrieved embeddings. Median absolute error (MAE) in years is reported as mean ± standard deviation. Pearson correlation r and regression slope indicate the correspondence between predicted and chronological age.
Figure 9: Zero-shot similarity-based brain age estimation using NeuroCBIR for the left hippocampus. Predicted age is obtained by aggregating the ages of the top- 500 most similar CN subjects. Panel (A) shows predicted versus chronological age for CN queries, panel (B) for AD queries, and panel (C) compares the age-gap distributions for CN and AD.
Method
CN/MCI/AD
CN/AD
BAcc ↑
MF1 ↑
BAcc ↑
MF1 ↑
Whole-brain
Random baseline
33.3
33.3
50.0
50.0
ResNet-10
49.8
48.1
69.5
59.9
ResNet-50
50.2
48.0
67.8
60.2
ResNet-101
47.4
44.7
65.0
56.4
Table 10: Zero-shot similarity-based disease risk estimation across selected brain regions. For each query, the predicted class was obtained from the similarity-weighted labels of the top- 500 retrieved embeddings. Performance is reported for both the three-class (CN/MCI/AD) and binary (CN/AD) classification tasks using balanced accuracy (BAcc) and macro F1-score (MF1).
Figure 10: Row-normalized confusion matrices for zero-shot similarity-based disease risk estimation across brain regions (Top-k = 300). Rows are normalized by true class, showing per-class recall (chance-level performance corresponds to 0.33.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Structure
mAP@5
mS@1
mS@5
Structure
mAP@5
mS@1
mS@5
Right-Cerebral-White-Matter
99.6
99.4
100.0
Left-Cerebellum-White-Matter
98.3
96.8
99.9
Left-Cerebral-White-Matter
99.5
99.4
99.9
ctx-lh-lateralorbitofrontal
98.2
98.3
99.6
Right-Cerebellum-Cortex
99.5
98.6
99.9
ctx-lh-pericalcarine
98.2
97.8
99.4
ctx-lh-superiortemporal
99.5
99.3
99.9
ctx-rh-parstriangularis
98.1
98.0
99.6
ctx-rh-fusiform
99.5
99.0
99.9
ctx-lh-bankssts
98.1
99.0
99.6
ctx-lh-fusiform
99.5
98.8
99.9
ctx-rh-bankssts
98.0
99.2
99.7
Appendix
Table 11: mAP@5, mS@1 and mS@5 for different brain regions of the NeuroCBIR framework across the testing dataset.
Department of Medicine, Shenyang Medical College, Shenyang, China · College of Medicine, University of Ibadan, Ibadan, Nigeria · University of Ghana, Accra, Ghana +2