cs.CVSep 30, 2026

Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification

Authors: Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala

Organizations: Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO 65211 USA · Government Degree College (Autonomous), Siddipet, Telangana 502103, India · Amar Biotech Private Limited, Kondapur, Hyderabad, Telangana 500084, India

Abstract

Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer's effect on measured performance across nine architectures and three evaluation conditions. Contamination is severe at every layer: 28.8% of the dominant corpus's official test split has a near-twin in its own training split, a second corpus leaks 22.3% of its test images byte-identically, 95.5% of traceable test images share a patient with training, and file-header features containing no anatomy separate tumor from no-tumor at 0.959 balanced accuracy, at parity with fine-tuned ResNet backbones. The unexpected result is that removing every identified leaked test image leaves balanced accuracy essentially unchanged: stable performance after deduplication does not establish benchmark integrity. Our findings establish dataset integrity as a distinct, measurable axis of benchmark quality that a stable leaderboard cannot certify. For biomedical research, reported accuracy on these corpora alone does not establish that a model has learned to recognize tumors rather than exploit dataset-specific cues. We release the contaminated-file lists, recovered patient identifiers, and deduplicated splits.

Figures & tables

Explore similar work

CardsList
  1. Cross-Dataset Generalization in Breast MRI Tumor Classification via Class-Wise Dataset Mixing

    Jul 21, 2026Mohammad Ali Dadrast, Hamid UsefiBrain Tumor ClassificationCross-Dataset Benchmark

  2. Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks

    Jul 14, 2026Wenhao Zhang, Zhongliang Zhou, John Kang +1Whole-Slide ImagesComputational Pathology

  3. A Leakage-Aware Comparative Benchmark of Machine Learning, Deep Learning, and Transformer Models for Reliable Leukemia Detection

    Jun 23, 2026Nisreen AlbzourAcute Myeloid LeukemiaCancer Detection