Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study
Organizations: University of Missouri · Government Degree College · Amar Bio Tech Pvt Ltd
Abstract
Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifiers and a controlled evaluation protocol. Under the same protocol, the eight pretrained backbones differ by only 0.026 AUROC. In contrast, changing whether the backbone is frozen or fine-tuned changes AUROC by 0.044 on average, and changing the decision threshold changes balanced accuracy by 0.090 on average. The official test split is also measurably different from the training pool: a partition classifier distinguishes them at AUC 0.697, rising to 0.898 for normal radiographs. Most strikingly, a classifier using only file properties, with no image anatomy, reaches 0.992 balanced accuracy within the training pool but falls to 0.496 on the official test split. Validation-fitted thresholds and calibration also transfer imperfectly. These results show that a high benchmark score can support different conclusions when the split, training policy, threshold, metric, calibration, and uncertainty are not communicated with it. We end with a seven-item reporting recommendation in which each item is tied to an effect measured in the study
Figures & tables
| Architecture | Bal. Acc. | Sens. | Spec. | AUROC | ECE |
|---|---|---|---|---|---|
| Custom CNN (scratch) | 0.697 0.032 | 0.881 0.027 | 0.512 0.076 | 0.748 0.042 | 0.178 0.044 |
| ResNet-18 | 0.768 0.016 | 0.967 0.019 | 0.569 0.038 | 0.909 0.006 | 0.143 0.027 |
| VGG-16 | 0.801 0.023 | 0.996 0.001 | 0.606 0.047 | 0.957 0.007 | 0.150 0.018 |
| ResNet-50 | 0.828 0.006 | 0.984 0.000 | 0.671 0.013 | 0.949 0.005 | 0.163 0.005 |
| EfficientNet-B0 | 0.784 0.011 | 0.963 0.007 | 0.605 0.029 | 0.893 0.008 | 0.191 0.001 |
| DenseNet-121 | 0.775 0.041 | 0.992 0.006 | 0.557 0.084 | 0.946 0.014 | 0.194 0.008 |
| Question | Comparison | Observed effect |
|---|---|---|
| Architecture under fixed | eight pretrained backbones | AUROC range 0.026 |
| Backbone training | frozen vs. fine tuned | mean 0.044 AUROC; max 0.103 |
| Threshold | 0.5 vs. validation selected | mean 0.090 balanced accuracy |
| Threshold transfer | validation vs. test oracle | mean gap 0.083 balanced accuracy |
| Calibration | raw vs. validation fitted | ECE 0.172 to 0.169; worse in 9/26 runs |
| Split and file shortcut | training pool CV vs. official test | 0.992 to 0.496 balanced accuracy |
| Study | Method | Evaluation setup | Reported score |
| ( Stephen et al., 2019 ) | CNN from scratch | pooled, re-split | 0.937 acc. |
| ( Kundu et al., 2021 ) | 3-model ensemble | pooled, 5-fold CV | 0.988 acc. |
| This work: untouched official test split; threshold selected on validation | |||
| ViT-B/16 | fine-tuned | official test | 0.858 acc. |
| ViT-B/16 | fine-tuned | official test | 0.811 bal. acc. |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Architecture | Par. (M) | Bal. Acc. | Sens. | Spec. | F1 | AUROC | AUPRC | ECE |
|---|---|---|---|---|---|---|---|---|
| Custom CNN | 0.5 | 0.697 0.032 | 0.881 0.027 | 0.512 0.076 | 0.811 0.015 | 0.748 0.042 | 0.787 0.048 | 0.178 0.044 |
| ResNet-18 | 11.2 | 0.768 0.016 | 0.967 0.019 | 0.569 0.038 | 0.870 0.009 | 0.909 0.006 | 0.930 0.004 | 0.143 0.027 |
| VGG-16 | 134.3 | 0.801 0.023 | 0.996 0.001 | 0.606 0.047 | 0.893 0.010 | 0.957 0.007 | 0.966 0.005 | 0.150 0.018 |
| ResNet-50 | 23.5 | 0.828 0.006 | 0.984 0.000 | 0.671 0.013 | 0.903 0.003 | 0.949 0.005 | 0.963 0.006 | 0.163 0.005 |
| EfficientNet-B0 | 4.0 | 0.784 0.011 | 0.963 0.007 | 0.605 0.029 | 0.876 0.004 | 0.893 0.008 | 0.921 0.010 | 0.191 0.001 |
| DenseNet-121 | 7.0 | 0.775 0.041 | 0.992 0.006 | 0.557 0.084 | 0.880 0.019 | 0.946 0.014 | 0.961 0.011 | 0.194 0.008 |
| Architecture | Balanced accuracy | F1 | AUROC | ||||||
|---|---|---|---|---|---|---|---|---|---|
| frozen | f.-t. | frozen | f.-t. | frozen | f.-t. | ||||
| ResNet-18 | 0.747 | 0.768 0.016 | 0.021 | 0.843 | 0.870 0.009 | 0.026 | 0.846 | 0.909 0.006 | 0.063 |
| VGG-16 | 0.777 | 0.801 0.023 | 0.024 | 0.872 | 0.893 0.010 | 0.021 | 0.906 | 0.957 0.007 | 0.052 |
| ResNet-50 | 0.782 | 0.828 0.006 | 0.046 | 0.869 | 0.903 0.003 | 0.034 | 0.909 | 0.949 0.005 | 0.040 |
| EfficientNet-B0 | 0.697 | 0.784 0.011 | 0.087 | 0.819 | 0.876 0.004 | 0.057 | 0.790 | 0.893 0.008 | 0.103 |
| DenseNet-121 | 0.835 | 0.775 0.041 | 0.061 | 0.903 | 0.880 0.019 | 0.023 | 0.946 | 0.946 0.014 | 0.000 |
| Architecture | Balanced accuracy | F1 | ECE | |||
|---|---|---|---|---|---|---|
| default | tuned | default | tuned | default | tuned | |
| Custom CNN | 0.652 0.052 | 0.697 0.032 | 0.809 0.013 | 0.811 0.015 | 0.139 0.079 | 0.178 0.044 |
| ResNet-18 | 0.723 0.028 | 0.768 0.016 | 0.852 0.008 | 0.870 0.009 | 0.119 0.012 | 0.143 0.027 |
| VGG-16 | 0.738 0.034 | 0.801 0.023 | 0.865 0.015 | 0.893 0.010 | 0.169 0.022 | 0.150 0.018 |
| ResNet-50 | 0.709 0.010 | 0.828 0.006 | 0.852 0.004 | 0.903 0.003 | 0.147 0.003 | 0.163 0.005 |
| EfficientNet-B0 | 0.642 0.006 | 0.784 0.011 | 0.823 0.002 | 0.876 0.004 | 0.248 0.004 | 0.191 0.001 |
| Threshold fitted on validation (Youden) | 0.787 |
|---|---|
| Threshold matched to the target prior | 0.863 |
| Threshold fitted on test (oracle) | 0.870 |
| prior matching recovers 92% of the oracle gap | |
| (b) Reliability and resolution decomposition of the Brier score | ||
|---|---|---|
| uncalibrated | temp.-scaled | |
| Reliability ( better) | 0.0655 | 0.0705 |
| Resolution ( better) | 0.1078 | 0.1173 |
| (c) Proxy -distance between partitions | ||
| Training pool vs. test, normal only | 1.226 | |
| Training pool vs. test, pneumonia only | 0.395 | |
| Architecture | single seed | 3-seed ensemble | |
|---|---|---|---|
| ViT-B/16 | 0.9834 | 0.9865 | 0.0031 |
| Swin-T | 0.9750 | 0.9798 | 0.0048 |
| VGG-16 | 0.9660 | 0.9650 | 0.0010 |
| ResNet-50 | 0.9599 | 0.9631 | 0.0032 |
| ResNet-18 | 0.9305 | 0.9448 | 0.0143 |
| EfficientNet-B0 | 0.9213 | 0.9344 | 0.0131 |