Beyond the Model: The Critical Role of Data Filtering in Clinical Machine Learning
Authors: Noah Subedar, Colin Campbell, Wenjing Zhang, Dan Perri, Sarah Culgin, Andrew Hamilton-Wright
Organizations: School of Computer Science, University of Guelph, Ontario, Canada · St. Joseph’s Healthcare Hamilton, Hamilton, Ontario, Canada · The Research Institute of St. Joe’s Hamilton, Hamilton, Ontario, Canada · McMaster Univeristy, Hamilton, Ontario, Canada
Machine learning (ML) studies using clinical data often rely on preprocessing and filtering pipelines before model development. The filtering decisions made in these pipelines can alter the dataset's statistical structure and may artificially reduce or increase the complexity of the prediction task. We argue that filtering choices should be treated as part of the scientific method rather than as a routine preprocessing step. We further discuss the need for explainable and transparent preprocessing pipelines that allow researchers to understand why specific filtering choices are made and how these choices affect the resulting data distribution and model performance. All of the source code for this work is available on GitHub.
Figures & tables
Feature
Abbreviation
Valid Range
Unit
Heart Rate
HR
[1,600]
bpm
Systolic Blood Pressure
SBP
[1,400]
mmHg
Diastolic Blood Pressure
DBP
[1,300]
mmHg
Mean Blood Pressure
MBP
[1,300]
mmHg
Respiration Rate
RR
[1,70]
breaths/min
Temperature
Temp
[21,50]
°C
TABLE I : Vital sign features and physiologically validated ranges used for outlier filtering.
Filter
Aggregation Stage
Level
Threshold
Heart Rate
Pre
Value
[1,600] bpm
Systolic BP
Pre
Value
[1,400] mmHg
Diastolic BP
Pre
Value
[1,300] mmHg
Mean BP
Pre
Value
[1,300] mmHg
Respiration Rate
Pre
Value
[1,70] breaths/min
Temperature
Pre
Value
[21,50] °C
TABLE II : Summary of filtering strategies
Fig. 1 : Centroid density of the raw MIMIC-III dataset
Fig. 2 : Heatmap of mean hourly observation counts per vital sign across MIMIC-III ICU stays (unfiltered, raw).
Fig. 3 : Centroid deviation by vital from baseline after the all-vitals outlier filter (mean aggregation). Grouped bars show the shifts for ICU and mortality positive/negative classes.
Fig. 4 : Empirical Cumulative Distribution Functions for 4 vital representatives showing the density distribution across the 4 different classes for the fill missing data filter and raw dataset.
Fig. 5 : Filter impact on mortality prediction under mean aggregation. Each bar shows the deviation in test accuracy (left) and macro F1 (right) relative to the raw unfiltered baseline, averaged over four random seeds.
Row
No. of Records
Max
Mean
Median
Std
a
1,158,296
9,999,999
98.6
86
9,291.5
b
1,157,789
459
90
86
24.4
c
45,985
555,652
105.3
87.1
2,590.9
d
45,989
251.5
93.187
87.1
25.5
TABLE III : Heart rate summary statistics of MIMIC-III across four levels of aggregation. Values are in beats per minute (BPM). Row (a) is the raw dataset; row (b) applies the project’s physiological plausibility range; rows (c) and (d) are the results of applying mean aggregation over the first 24 hours of the stay before and after the heart-rate outlier filter.
Fig. 6 : Centroid deviation from baseline per vital sign after the high-invalid-data filter (mean aggregation). Grouped bars show the four outcome sub-populations: ICU positive/negative and mortality positive/negative.
Fig. 7 : McNemar significance test for mortality prediction under mean aggregation. Bars show −log10(p) for each filter versus the raw baseline. The dashed red lines mark our pre-determined p=0.05 and p2=0.0001 threasholds.
Fig. 8 : Filter impact on ICU prediction under mean aggregation. Each bar shows the deviation in test accuracy relative to the raw unfiltered baseline, averaged over four random seeds
We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments suggest that with enough compute, the best data filter is no data filter. We find that sufficiently trained large parameter models not only tolerate low-quality and distractor data, but in fact benefit from nominally ``poor'' data.
Christopher Mohri, John Duchi, Tatsunori Hashimoto
Department of Computer Science Stanford University · Departments of Statistics and Electrical Engineering Stanford University
Postoperative adverse events, including mortality and morbidity, remain a major global burden, many of which are preventable through early identification of high-risk patients and targeted perioperative care. Accurate risk stratification is therefore essential. With the growing availability of large-scale electronic health records (EHRs), machine learning (ML) provides a data-driven approach to model complex clinical patterns. However, existing studies vary widely in design, and methodological practices remain fragmented. This scoping review characterizes ML pipelines for surgical risk stratification and outcome prediction using EHR data. We reviewed 190 studies covering the ML workflow, including data preprocessing, algorithm selection, model evaluation, and explainability. Most studies relied on single-center private datasets with limited data modalities, while the scarcity of open-access surgical datasets constrained reproducibility and generalizability. Reporting of key preprocessing steps, including missing data handling, feature selection, and class imbalance, was often incomplete. Conventional ML models and simple neural networks predominated, whereas deep learning and multimodal approaches remained uncommon. Benchmark datasets and standardized evaluation protocols were largely absent, hindering cross-study comparisons. Only about one-third of studies incorporated explainability methods. This review identifies methodological gaps limiting clinically robust postoperative ML tools and provides a structured reference to support more rigorous, reproducible, and clinically meaningful ML development for perioperative care.
Yizhi Dong, Yuhe Ke, Hairil Rizal Abdullah +4
Saw Swee Hock School of Public Health, National University of Singapore, Singapore 117549 · Data Science and Artificial Intelligence Lab, Singapore General Hospital, Singapore 169608 · Duke-NUS Medical School, Singapore 169857 +5
Prior work evaluates code generation bias primarily through simple conditional statements, which represent only a narrow slice of real-world programming and reveal solely overt, explicitly encoded bias. We demonstrate that this approach dramatically underestimates bias in practice by examining a more realistic task: generating machine learning (ML) pipelines. Testing both code-specialized and general-instruction large language models, we find that generated pipelines exhibit significant bias during feature selection. Sensitive attributes appear in 87.7% of cases on average, despite models demonstrably excluding irrelevant features (e.g., including "race" while dropping "favorite color" for credit scoring). This bias is substantially more prevalent than that captured by conditional statements, where sensitive attributes appear in only 59.2% of cases. These findings are robust across prompt mitigation strategies, varying numbers of attributes, and different pipeline difficulty levels. Our results challenge simple conditionals as valid proxies for bias evaluation and suggest current benchmarks underestimate bias risk in practical deployments.
Minh Duc Bui, Xenia Heilmann, Mattia Cerrato +2
Johannes Gutenberg University Mainz, Germany · Universidad Iberoamericana, Ciudad de Mexico · University of Colorado Boulder, USA