Organizations: Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO 65211 USA · Government Degree College (Autonomous), Siddipet, Telangana 502103, India · Amar Biotech Private Limited, Kondapur, Hyderabad, Telangana 500084, India
Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer's effect on measured performance across nine architectures and three evaluation conditions. Contamination is severe at every layer: 28.8% of the dominant corpus's official test split has a near-twin in its own training split, a second corpus leaks 22.3% of its test images byte-identically, 95.5% of traceable test images share a patient with training, and file-header features containing no anatomy separate tumor from no-tumor at 0.959 balanced accuracy, at parity with fine-tuned ResNet backbones. The unexpected result is that removing every identified leaked test image leaves balanced accuracy essentially unchanged: stable performance after deduplication does not establish benchmark integrity. Our findings establish dataset integrity as a distinct, measurable axis of benchmark quality that a stable leaderboard cannot certify. For biomedical research, reported accuracy on these corpora alone does not establish that a model has learned to recognize tumors rather than exploit dataset-specific cues. We release the contaminated-file lists, recovered patient identifiers, and deduplicated splits.
Figures & tables
Fig. 1: Number of C1 test images flagged as leaked, as a function of the correlation threshold. The count is nearly flat from ρ=0.98 to ρ=0.995 , indicating the flagged pairs are true duplicates rather than an artifact of where the threshold was placed.
Fig. 2: Highest-correlation pairs across C1’s official train/test boundary, one per class. Top row: images from the official test split. Bottom row: their matches in the official training split.
Fig. 3: Best-match correlation against the training split, for every test image. The brain corpus has a distinct mode at ρ≈1 that the chest-radiograph control lacks entirely, despite nearly identical medians.
Corpus
ntrain
ntest
exact (MD5)
near ( ρ≥0.98 )
Nickparvar
5600
1600
0 (0.0%)
460 (28.75%)
SARTAJ
2870
394
88 (22.34%)
259 (65.74%)
Kermany CXR
5216
624
0 (0.0%)
0 (0.0%)
TABLE I: Train/test contamination inside the official split of each corpus. A test image counts as leaked when an image with normalized cross-correlation ≥0.98 exists in the same corpus’ training split. The chest-radiograph corpus is included as a negative control.
Query corpus
Reference corpus
n
exact
near
sartaj_all
nickparvar_all
3264
2670 (81.8%)
2705 (82.87%)
navoneel
nickparvar_all
253
66 (26.09%)
123 (48.62%)
navoneel
sartaj_all
253
0 (0.0%)
98 (38.74%)
TABLE II: Overlap between the three brain-MRI corpora. Because the Nickparvar corpus is assembled from figshare, SARTAJ and Br35H, most of SARTAJ is physically contained in it, so an uncorrected ‘cross-dataset’ evaluation between the two mostly re-scores training images.
Corpus
ntest
traceable to figshare
patients
same patient in train
Nickparvar
1600
201 (12.56%)
125
192 (95.52%)
SARTAJ
394
28 (7.11%)
11
28 (100.0%)
TABLE III: Patient-level leakage, measured by recovering cjdata.PID from the original figshare release (3064 slices from only 233 patients). The last column counts test images that have another slice of the same patient in the training split, as a fraction of the traceable subset.
Class
n
traceable
same patient in train
glioma
400
98
90
meningioma
400
26
26
notumor
400
0
0
pituitary
400
77
76
TABLE IV: Per-class breakdown of the patient-identifier recovery. The notumor class matches nothing, as it must: it originates from Br35H rather than figshare. This doubles as a check that the matcher is not matching indiscriminately.
Fig. 4: Left: the figshare source contains only 233 patients across 3064 slices. Right: of the test images whose provenance is recoverable, nearly all have another slice of the same patient in training.
Task
chance
file metadata
+ intensity
Brain: tumor vs none
0.50
0.959
0.966
Brain: 4-class
0.25
0.652
0.830
Chest X-ray (control)
0.50
0.496
0.499
Chest X-ray, 5-fold CV within train
0.50
0.992
0.997
TABLE V: Balanced accuracy of a gradient-boosted tree trained on features containing no anatomy. ‘File metadata’ is pixel dimensions, aspect ratio, file size, bytes per pixel, image mode and JPEG quantisation tables. ‘ + intensity’ adds global intensity summaries. Rows are fitted on the official training split and scored on the official test split, except the last.
Fig. 5: Balanced accuracy from features containing no anatomy, fitted on the official training split and scored on the official test split. Dashed lines are chance. The brain benchmark’s binary task is largely solvable from file headers; the chest-radiograph control is not.
Architecture
Official
Dedup.
Δ
SARTAJ (clean)
ECE in-dom.
ECE ext.
Custom CNN
0.589 ± 0.010
0.622 ± 0.008
− 0.033
0.491 ± 0.038
0.078 ± 0.004
0.185 ± 0.032
ResNet-18
0.862 ± 0.009
0.882 ± 0.011
− 0.020
0.924 ± 0.010
0.057 ± 0.008
0.081 ± 0.011
VGG-16
0.945 ± 0.002
0.946 ± 0.002
− 0.001
0.933 ± 0.039
0.035 ± 0.002
0.012 ± 0.000
ResNet-50
0.874 ± 0.009
0.894 ± 0.005
− 0.020
0.830 ± 0.021
0.053 ± 0.003
0.069 ± 0.007
EfficientNet-B0
0.929 ± 0.003
0.932 ± 0.005
− 0.003
0.912 ± 0.076
0.030 ± 0.001
0.019 ± 0.002
DenseNet-121
0.942 ± 0.002
0.944 ± 0.002
− 0.002
0.893 ± 0.011
0.036 ± 0.004
0.022 ± 0.001
TABLE VI: Four-class tumor classification trained on the Nickparvar corpus (mean ± s.d., three seeds). ‘Official’ is the corpus’ own test split; ‘deduplicated’ removes the 460 test images that recur in training; ‘SARTAJ (clean)’ is an external corpus after removing every image that overlaps the training source.
Fig. 6: Balanced accuracy of each architecture on C1’s official test split, on the deduplicated split, and on the disjoint external corpus. Error bars are standard deviation over three seeds. The ordering of architectures is not preserved across the three evaluations.
Architecture
Navoneel (held-out)
Nickparvar test
SARTAJ (clean)
Custom CNN
0.650 ± 0.031
0.294 ± 0.016
0.299 ± 0.044
ResNet-18
0.712 ± 0.018
0.589 ± 0.018
0.510 ± 0.056
VGG-16
0.841 ± 0.027
0.867 ± 0.018
0.695 ± 0.012
ResNet-50
0.745 ± 0.008
0.700 ± 0.067
0.651 ± 0.126
EfficientNet-B0
0.739 ± 0.017
0.802 ± 0.016
0.629 ± 0.147
DenseNet-121
0.785 ± 0.040
0.788 ± 0.028
0.607 ± 0.054
TABLE VII: Reverse transfer: binary tumor detection trained on the 253-image Navoneel corpus and applied to the two large corpora. The within-corpus column is optimistic by construction – its test split is 76 images drawn from the same 253.
Model
Balanced accuracy
ConvNeXt-T
0.989
Swin-T
0.989
VGG-16
0.986
ViT-B/16
0.985
DenseNet-121
0.984
EfficientNet-B0
0.982
TABLE VIII: Tumor vs. no-tumor on the official Nickparvar test split, with the four-class head projected to binary. The last two rows are the header-only baselines of Section VI-C , which see no anatomy whatsoever.
Fig. 7: Left: confusion matrix of the strongest model on the official test split. Right: per-class recall under the three evaluations. Deduplication barely moves any class; the external corpus costs most in the tumor subtypes rather than in notumor .
Zero-shot control
leaked
clean
gap
BiomedCLIP
0.891
0.856
+ 0.035
CLIP ViT-B/16
0.865
0.966
− 0.101
TABLE IX: An attempt to separate memorization from difficulty using zero-shot models as leakage-immune controls, and why it fails. Neither control was trained on any of these corpora, so neither can memorize the leaked subset; both are exposed to whatever difficulty difference exists between the leaked and clean partitions. The two controls disagree in sign, so the difference-in-differences estimand is not identified and we do not use it.
Fig. 8: Grad-CAM for DenseNet-121 on three official-test images that have a near-twin in training (top row) and three images from the external, deduplicated corpus (bottom row); predicted-class correctness and confidence above each panel. Attribution concentrates on comparable regions in both rows and the external images are classified with equal or higher confidence. The quantitative leaked-versus-clean comparison on the official split is in Fig. 9 .
Fig. 9: SHAP attribution magnitude for leaked and clean test images. The summary statistic is the fraction of attribution mass outside the brain mask (Section VI-F ).
Data Science Program Memorial University of Newfoundland St. John’s, Canada · Department of Mathematics and Statistics Memorial University of Newfoundland St. John’s, Canada