Beyond the Model: The Critical Role of Data Filtering in Clinical Machine Learning
Authors: Noah Subedar, Colin Campbell, Wenjing Zhang, Dan Perri, Sarah Culgin, Andrew Hamilton-Wright
Organizations: School of Computer Science, University of Guelph, Ontario, Canada · St. Joseph’s Healthcare Hamilton, Hamilton, Ontario, Canada · The Research Institute of St. Joe’s Hamilton, Hamilton, Ontario, Canada · McMaster Univeristy, Hamilton, Ontario, Canada
Machine learning (ML) studies using clinical data often rely on preprocessing and filtering pipelines before model development. The filtering decisions made in these pipelines can alter the dataset's statistical structure and may artificially reduce or increase the complexity of the prediction task. We argue that filtering choices should be treated as part of the scientific method rather than as a routine preprocessing step. We further discuss the need for explainable and transparent preprocessing pipelines that allow researchers to understand why specific filtering choices are made and how these choices affect the resulting data distribution and model performance. All of the source code for this work is available on GitHub.
Figures & tables
Feature
Abbreviation
Valid Range
Unit
Heart Rate
HR
[1,600]
bpm
Systolic Blood Pressure
SBP
[1,400]
mmHg
Diastolic Blood Pressure
DBP
[1,300]
mmHg
Mean Blood Pressure
MBP
[1,300]
mmHg
Respiration Rate
RR
[1,70]
breaths/min
Temperature
Temp
[21,50]
°C
TABLE I : Vital sign features and physiologically validated ranges used for outlier filtering.
Filter
Aggregation Stage
Level
Threshold
Heart Rate
Pre
Value
[1,600] bpm
Systolic BP
Pre
Value
[1,400] mmHg
Diastolic BP
Pre
Value
[1,300] mmHg
Mean BP
Pre
Value
[1,300] mmHg
Respiration Rate
Pre
Value
[1,70] breaths/min
Temperature
Pre
Value
[21,50] °C
TABLE II : Summary of filtering strategies
Fig. 1 : Centroid density of the raw MIMIC-III dataset
Fig. 2 : Heatmap of mean hourly observation counts per vital sign across MIMIC-III ICU stays (unfiltered, raw).
Fig. 3 : Centroid deviation by vital from baseline after the all-vitals outlier filter (mean aggregation). Grouped bars show the shifts for ICU and mortality positive/negative classes.
Fig. 4 : Empirical Cumulative Distribution Functions for 4 vital representatives showing the density distribution across the 4 different classes for the fill missing data filter and raw dataset.
Fig. 5 : Filter impact on mortality prediction under mean aggregation. Each bar shows the deviation in test accuracy (left) and macro F1 (right) relative to the raw unfiltered baseline, averaged over four random seeds.
Row
No. of Records
Max
Mean
Median
Std
a
1,158,296
9,999,999
98.6
86
9,291.5
b
1,157,789
459
90
86
24.4
c
45,985
555,652
105.3
87.1
2,590.9
d
45,989
251.5
93.187
87.1
25.5
TABLE III : Heart rate summary statistics of MIMIC-III across four levels of aggregation. Values are in beats per minute (BPM). Row (a) is the raw dataset; row (b) applies the project’s physiological plausibility range; rows (c) and (d) are the results of applying mean aggregation over the first 24 hours of the stay before and after the heart-rate outlier filter.
Fig. 6 : Centroid deviation from baseline per vital sign after the high-invalid-data filter (mean aggregation). Grouped bars show the four outcome sub-populations: ICU positive/negative and mortality positive/negative.
Fig. 7 : McNemar significance test for mortality prediction under mean aggregation. Bars show −log10(p) for each filter versus the raw baseline. The dashed red lines mark our pre-determined p=0.05 and p2=0.0001 threasholds.
Fig. 8 : Filter impact on ICU prediction under mean aggregation. Each bar shows the deviation in test accuracy relative to the raw unfiltered baseline, averaged over four random seeds
Saw Swee Hock School of Public Health, National University of Singapore, Singapore 117549 · Data Science and Artificial Intelligence Lab, Singapore General Hospital, Singapore 169608 · Duke-NUS Medical School, Singapore 169857 +5