Organizations: Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO 65211 USA · Government Degree College (Autonomous), Siddipet, Telangana 502103, India · Amar Biotech Private Limited, Kondapur, Hyderabad, Telangana 500084, India
Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer's effect on measured performance across nine architectures and three evaluation conditions. Contamination is severe at every layer: 28.8% of the dominant corpus's official test split has a near-twin in its own training split, a second corpus leaks 22.3% of its test images byte-identically, 95.5% of traceable test images share a patient with training, and file-header features containing no anatomy separate tumor from no-tumor at 0.959 balanced accuracy, at parity with fine-tuned ResNet backbones. The unexpected result is that removing every identified leaked test image leaves balanced accuracy essentially unchanged: stable performance after deduplication does not establish benchmark integrity. Our findings establish dataset integrity as a distinct, measurable axis of benchmark quality that a stable leaderboard cannot certify. For biomedical research, reported accuracy on these corpora alone does not establish that a model has learned to recognize tumors rather than exploit dataset-specific cues. We release the contaminated-file lists, recovered patient identifiers, and deduplicated splits.
Figures & tables
Fig. 1: Number of C1 test images flagged as leaked, as a function of the correlation threshold. The count is nearly flat from ρ=0.98 to ρ=0.995 , indicating the flagged pairs are true duplicates rather than an artifact of where the threshold was placed.
Fig. 2: Highest-correlation pairs across C1’s official train/test boundary, one per class. Top row: images from the official test split. Bottom row: their matches in the official training split.
Fig. 3: Best-match correlation against the training split, for every test image. The brain corpus has a distinct mode at ρ≈1 that the chest-radiograph control lacks entirely, despite nearly identical medians.
Corpus
ntrain
ntest
exact (MD5)
near ( ρ≥0.98 )
Nickparvar
5600
1600
0 (0.0%)
460 (28.75%)
SARTAJ
2870
394
88 (22.34%)
259 (65.74%)
Kermany CXR
5216
624
0 (0.0%)
0 (0.0%)
TABLE I: Train/test contamination inside the official split of each corpus. A test image counts as leaked when an image with normalized cross-correlation ≥0.98 exists in the same corpus’ training split. The chest-radiograph corpus is included as a negative control.
Query corpus
Reference corpus
n
exact
near
sartaj_all
nickparvar_all
3264
2670 (81.8%)
2705 (82.87%)
navoneel
nickparvar_all
253
66 (26.09%)
123 (48.62%)
navoneel
sartaj_all
253
0 (0.0%)
98 (38.74%)
TABLE II: Overlap between the three brain-MRI corpora. Because the Nickparvar corpus is assembled from figshare, SARTAJ and Br35H, most of SARTAJ is physically contained in it, so an uncorrected ‘cross-dataset’ evaluation between the two mostly re-scores training images.
Corpus
ntest
traceable to figshare
patients
same patient in train
Nickparvar
1600
201 (12.56%)
125
192 (95.52%)
SARTAJ
394
28 (7.11%)
11
28 (100.0%)
TABLE III: Patient-level leakage, measured by recovering cjdata.PID from the original figshare release (3064 slices from only 233 patients). The last column counts test images that have another slice of the same patient in the training split, as a fraction of the traceable subset.
Class
n
traceable
same patient in train
glioma
400
98
90
meningioma
400
26
26
notumor
400
0
0
pituitary
400
77
76
TABLE IV: Per-class breakdown of the patient-identifier recovery. The notumor class matches nothing, as it must: it originates from Br35H rather than figshare. This doubles as a check that the matcher is not matching indiscriminately.
Fig. 4: Left: the figshare source contains only 233 patients across 3064 slices. Right: of the test images whose provenance is recoverable, nearly all have another slice of the same patient in training.
Task
chance
file metadata
+ intensity
Brain: tumor vs none
0.50
0.959
0.966
Brain: 4-class
0.25
0.652
0.830
Chest X-ray (control)
0.50
0.496
0.499
Chest X-ray, 5-fold CV within train
0.50
0.992
0.997
TABLE V: Balanced accuracy of a gradient-boosted tree trained on features containing no anatomy. ‘File metadata’ is pixel dimensions, aspect ratio, file size, bytes per pixel, image mode and JPEG quantisation tables. ‘ + intensity’ adds global intensity summaries. Rows are fitted on the official training split and scored on the official test split, except the last.
Fig. 5: Balanced accuracy from features containing no anatomy, fitted on the official training split and scored on the official test split. Dashed lines are chance. The brain benchmark’s binary task is largely solvable from file headers; the chest-radiograph control is not.
Architecture
Official
Dedup.
Δ
SARTAJ (clean)
ECE in-dom.
ECE ext.
Custom CNN
0.589 ± 0.010
0.622 ± 0.008
− 0.033
0.491 ± 0.038
0.078 ± 0.004
0.185 ± 0.032
ResNet-18
0.862 ± 0.009
0.882 ± 0.011
− 0.020
0.924 ± 0.010
0.057 ± 0.008
0.081 ± 0.011
VGG-16
0.945 ± 0.002
0.946 ± 0.002
− 0.001
0.933 ± 0.039
0.035 ± 0.002
0.012 ± 0.000
ResNet-50
0.874 ± 0.009
0.894 ± 0.005
− 0.020
0.830 ± 0.021
0.053 ± 0.003
0.069 ± 0.007
EfficientNet-B0
0.929 ± 0.003
0.932 ± 0.005
− 0.003
0.912 ± 0.076
0.030 ± 0.001
0.019 ± 0.002
DenseNet-121
0.942 ± 0.002
0.944 ± 0.002
− 0.002
0.893 ± 0.011
0.036 ± 0.004
0.022 ± 0.001
TABLE VI: Four-class tumor classification trained on the Nickparvar corpus (mean ± s.d., three seeds). ‘Official’ is the corpus’ own test split; ‘deduplicated’ removes the 460 test images that recur in training; ‘SARTAJ (clean)’ is an external corpus after removing every image that overlaps the training source.
Fig. 6: Balanced accuracy of each architecture on C1’s official test split, on the deduplicated split, and on the disjoint external corpus. Error bars are standard deviation over three seeds. The ordering of architectures is not preserved across the three evaluations.
Architecture
Navoneel (held-out)
Nickparvar test
SARTAJ (clean)
Custom CNN
0.650 ± 0.031
0.294 ± 0.016
0.299 ± 0.044
ResNet-18
0.712 ± 0.018
0.589 ± 0.018
0.510 ± 0.056
VGG-16
0.841 ± 0.027
0.867 ± 0.018
0.695 ± 0.012
ResNet-50
0.745 ± 0.008
0.700 ± 0.067
0.651 ± 0.126
EfficientNet-B0
0.739 ± 0.017
0.802 ± 0.016
0.629 ± 0.147
DenseNet-121
0.785 ± 0.040
0.788 ± 0.028
0.607 ± 0.054
TABLE VII: Reverse transfer: binary tumor detection trained on the 253-image Navoneel corpus and applied to the two large corpora. The within-corpus column is optimistic by construction – its test split is 76 images drawn from the same 253.
Model
Balanced accuracy
ConvNeXt-T
0.989
Swin-T
0.989
VGG-16
0.986
ViT-B/16
0.985
DenseNet-121
0.984
EfficientNet-B0
0.982
TABLE VIII: Tumor vs. no-tumor on the official Nickparvar test split, with the four-class head projected to binary. The last two rows are the header-only baselines of Section VI-C , which see no anatomy whatsoever.
Fig. 7: Left: confusion matrix of the strongest model on the official test split. Right: per-class recall under the three evaluations. Deduplication barely moves any class; the external corpus costs most in the tumor subtypes rather than in notumor .
Zero-shot control
leaked
clean
gap
BiomedCLIP
0.891
0.856
+ 0.035
CLIP ViT-B/16
0.865
0.966
− 0.101
TABLE IX: An attempt to separate memorization from difficulty using zero-shot models as leakage-immune controls, and why it fails. Neither control was trained on any of these corpora, so neither can memorize the leaked subset; both are exposed to whatever difficulty difference exists between the leaked and clean partitions. The two controls disagree in sign, so the difference-in-differences estimand is not identified and we do not use it.
Fig. 8: Grad-CAM for DenseNet-121 on three official-test images that have a near-twin in training (top row) and three images from the external, deduplicated corpus (bottom row); predicted-class correctness and confidence above each panel. Attribution concentrates on comparable regions in both rows and the external images are classified with equal or higher confidence. The quantitative leaked-versus-clean comparison on the official split is in Fig. 9 .
Fig. 9: SHAP attribution magnitude for leaked and clean test images. The summary statistic is the fraction of attribution mass outside the brain mask (Section VI-F ).
Breast MRI is highly sensitive for detecting breast tumors, but exams contain many slices and require substantial reading time. Deep learning models often perform well on internal splits but can fail across institutions because of domain shift and dataset-origin bias. We study this failure mode for binary breast MRI tumor classification. EfficientNet-B3 and WaveViT-Small are trained using Duke Breast Cancer MRI and fastMRI, and evaluated only on the independent multi-center MAMA-MIA cohort. In a deliberately confounded setup, where label is perfectly correlated with dataset origin, external accuracy is near chance (0.5048--0.5265), despite very high recall. We then construct a mixed training set in which each class contains samples from both Duke and fastMRI, while preserving patient-level splitting, augmentation, and leakage controls. On MAMA-MIA, dataset mixing improves accuracy/F1 to 0.8463/0.8625 for WaveViT-Small and 0.8884/0.8994 for EfficientNet-B3. These results show that controlling dataset-origin bias is important for reliable breast MRI classification.
Mohammad Ali Dadrast, Hamid Usefi
Data Science Program Memorial University of Newfoundland St. John’s, Canada · Department of Mathematics and Statistics Memorial University of Newfoundland St. John’s, Canada
Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answering (VQA) benchmarks. We audit these claims and find them fundamentally compromised by data leakage at two hierarchical levels: patient-level leakage, where slides from the same case appear in both training and test folds, and institutional-level leakage, where different cases nonetheless share staining-batch and scanner signatures through a common Tissue Source Site (TSS). By tracing canonical slide, case, and TSS identifiers across major public resources, we document case level train test overlaps of 92.3~100% on TCGA-derived benchmarks, together with near-complete TSS overlap. We further demonstrate that both leakage levels are linearly decodable from foundation-model feature space, that they induce a measurable accuracy gap between leaked and audit-clean cases on a published checkpoint, and that across multiple published WSI VLMs, peak reported accuracies concentrate on the most heavily contaminated benchmarks. Therefore, the current WSI VQA evaluation cannot distinguish genuine multimodal reasoning from nearest-neighbor retrieval over memorized institutional and patient-specific artifacts. Finally, we outline concrete recommendations for contamination-free evaluation. By addressing benchmark construction, provenance disclosure, and automated overlap auditing, we aim to guide future research toward verifiable claims of progress.
Automated classification of acute lymphoblastic leukemia (ALL) from peripheral blood smear images has often reported near-perfect performance on the C-NMC 2019 dataset. We show that such results can be inflated by patient-level data leakage caused by random image-level partitioning, where cells from the same subject may appear in both training and test folds. We establish a leakage-aware benchmark under a strict subject-disjoint protocol, comparing LightGBM, RBF-SVM, EfficientNet-B0, EfficientNet-B1, and ViT-Tiny. Models are developed using three subject-disjoint folds from 73 subjects and evaluated on an external preliminary-phase test set of 1,867 images from 28 unseen subjects with zero patient overlap. Beyond discrimination, we assess calibration using expected calibration error, Brier score, and temperature scaling. Under honest evaluation, EfficientNet-B1 achieves the best performance, with AUROC 0.913, sensitivity 0.87, specificity 0.80, and calibrated ECE 0.024. Frozen-feature classifiers and ViT-Tiny show high sensitivity but poor specificity, indicating a tendency to over-predict the malignant class. A random-versus-subject-disjoint ablation shows that random splitting inflates AUROC by about 0.04 even in the conservative frozen-feature setting. These findings caution against image-level evaluation on C-NMC 2019 and provide a reproducible, calibration-aware benchmark for future work.
Nisreen Albzour
School of Systems Science and Industrial Engineering, Binghamton University, Binghamton, NY, USA