MedFeat: Model-Aware and Explainability-Driven Feature Engineering with LLMs for Tabular Prediction
Organizations: Microsoft Research · University of Oxford
Abstract
In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods. LLMs are increasingly used to automate this process, acting as domain experts that propose diverse feature transformations to boost downstream performance. However, the feature generation process of existing LLM-based methods is agnostic to the downstream learner: the LLM receives no signal about which features currently drive predictions or where the model's representational capacity falls short, so proposals are neither targeted to promising regions of the feature space nor tailored to the learner's inductive bias. This shortcoming is amplified in healthcare data, which simultaneously exhibits class imbalance, heterogeneous feature spaces, and strict interpretability requirements. In this paper, we propose MedFeat, the first feature engineering framework inspired by the workflow of machine learning practitioners, leveraging model-awareness and feature importance signals to iteratively guide feature discovery for clinical tabular learning. We evaluate MedFeat on a broad range of challenging real-world clinical tasks and show that it statistically significantly outperforms state-of-the-art baselines, with an average F1 improvement of more than 10% over the baseline across models with distinct inductive biases.
Figures & tables
| Method | Mortality-ICU | Mortality-10yrs | Heart Failure | Mortality-Ward | Cog-Impairment | Discharge | Readmission | Avg. Rank | |||||||
| F1 | AUROC | F1 | AUROC | F1 | AUROC | F1 | AUROC | F1 | AUROC | F1 | AUROC | F1 | AUROC | ||
| XGBoost | |||||||||||||||
| Baseline | 9.5 0.9 | 75.6 1.3 | 23.6 0.8 | 68.3 1.9 | 12.9 0.6 | 67.4 1.2 | 1.4 0.2 | 71.0 1.4 | 55.3 1.7 | 73.8 1.6 | 39.8 0.4 | 71.0 0.5 | 29.7 0.5 | 58.7 1.7 | 4.1 |
| OpenFE | 10.2 0.5 | 73.8 0.5 | 25.0 1.8 | 69.2 1.9 | 12.2 0.4 | 65.6 1.0 | 6.1 4.1 | 74.8 1.1 | 56.7 1.1 | 74.7 1.0 | 38.1 0.8 | 68.9 1.1 | 29.2 0.3 | 59.0 0.4 | 3.6 |
| CAAFE | 10.0 0.7 | 75.8 1.4 | 24.9 1.6 | 68.2 2.1 | 12.8 0.6 | 67.9 0.7 | 1.5 0.5 | 68.3 3.2 | 55.6 1.3 | 74.1 1.4 | 40.2 0.8 | 71.1 0.7 | 30.0 0.2 | 60.4 0.7 | 3.2 |
| FeatLLM | N/A | 50.0 0.0 | N/A | 55.9 5.5 | N/A | 50.0 0.0 | N/A | 50.0 0.0 | N/A | 51.3 1.9 | N/A | 50.0 0.0 | N/A | 50.0 0.0 | 6.0 |
| Model | F1 | AUROC | AUPRC | ||||||
| MedFeat | Baseline | % [min, max] | MedFeat | Baseline | % [min, max] | MedFeat | Baseline | % [min, max] | |
| Logistic Reg. | 26.6 1.1 | 25.8 1.1 | 73.0 1.3 | 72.2 1.2 | 23.5 0.9 | 22.6 0.9 | |||
| XGBoost | 27.8 1.1 | 27.3 1.0 | 75.3 1.4 | 74.8 1.4 | 27.5 1.8 | 27.0 1.7 | |||
| Overall | 27.2 1.1 | 26.6 1.0 | 74.1 1.4 | 73.5 1.3 | 25.5 1.4 | 24.8 1.3 | |||
| Ablation Mode | F1 ( %) | AUROC ( %) | AUPRC ( %) |
| w/o SHAP Importance | 25.3 1.0 (–4.5%) | 71.3 1.2 (–1.5%) | 22.8 0.9 (–4.2%) |
| w/o Model-Awareness | 25.6 1.0 (–3.4%) | 71.0 1.2 (–1.9%) | 23.2 1.0 (–2.5%) |
| w/o Island Sampling | 25.8 1.1 (–2.6%) | 71.1 1.3 (–1.8%) | 23.3 1.0 (–2.1%) |
| w/o Memory | 25.8 1.0 (–2.6%) | 71.1 1.2 (–1.8%) | 23.3 1.0 (–2.1%) |
| CAAFE | OCTree | FeatLLM | MedFeat (Ours) | ||||||||
| LLM Backbone | F1 | AUROC | AUPRC | F1 | AUROC | AUPRC | AUROC | AUPRC | F1 | AUROC | AUPRC |
| gpt-4o-2024-11-20 | 25.1 0.9 | 70.8 1.3 | 22.5 1.1 | 25.4 1.0 | 71.0 1.2 | 22.5 0.9 | 52.2 1.6 | 12.2 0.7 | 26.5 1.0 | 72.4 1.4 | 23.8 0.9 |
| gpt-5.4 | 25.2 1.0 | 70.7 1.2 | 22.5 1.0 | 25.8 1.0 | 70.6 0.9 | 21.8 0.8 | 53.8 2.8 | 13.3 1.4 | 26.5 0.9 | 72.2 1.2 | 24.8 1.3 |
| gemini-3.1-pro-preview | 25.2 0.8 | 70.8 1.3 | 22.5 0.9 | 25.8 1.0 | 70.8 1.1 | 21.9 0.8 | 53.5 2.3 | 13.0 0.9 | 25.9 1.0 | 72.1 1.2 | 23.6 1.0 |
| qwen-3.6-27B-FP8 | 25.0 0.9 | 70.5 1.1 | 22.4 0.9 | 25.8 1.0 | 70.7 0.9 | 21.8 0.8 | 52.2 2.2 | 12.3 1.0 | 26.2 1.0 | 71.5 1.1 | 23.9 1.0 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Range [Min, Max] |
| max_depth | [2, 7] |
| n_estimators | [100, 2000] |
| min_child_weight | [10, 100] |
| max_delta_step | [0, 10] |
| subsample | [0.5, 1.0] |
| learning_rate | [0.01, 0.5] |
| Parameter | Range [Min, Max] |
| C | [ , ] |
| Dataset | Task | N | Ratio (+:-) | Task Description | |
| IORD | Mortality-Ward | 92,237 | 31 | 1:313 | Predict whether the patient will die within 1 day. |
| Discharge | 92,237 | 31 | 1:4 | Predict whether the patient will be discharged within 1 day. | |
| MIMIC | Mortality-ICU | 90,257 | 16 | 1:47 | Predict whether the patient will die within 24 hours. |
| Readmission | 74,044 | 17 | 1:4 | Predict whether the patient will be readmitted within 30 days. | |
| Heart Failure | 74,044 | 17 | 1:17 | Predict whether the patient will develop heart failure within 30 days. | |
| HRS | Mortality-10yrs | 8,899 | 41 | 1:9 | Predict whether the participant will die within 10 years. |
| Feature | Type |
| Age | Numerical |
| Sex | Categorical |
| Index of multiple deprivation score | Numerical |
| Admission method code | Categorical |
| Hours in hospital after admission | Numerical |
| Number of unique medicines received last 24h | Numerical |
| Feature | Type |
| Age at ICU admission | Numerical |
| Sex | Categorical |
| Admission type | Categorical |
| Admission location | Categorical |
| Insurance | Categorical |
| Language | Categorical |
| Feature | Type |
| Age at ICU admission | Numerical |
| Sex | Categorical |
| Admission type | Categorical |
| Admission location | Categorical |
| Insurance | Categorical |
| Language | Categorical |
| Feature | Type |
| Respondent gender | Categorical |
| Respondent race | Categorical |
| Respondent years of education | Numerical |
| Respondent veteran status | Binary |
| Respondent marital status at wave | Categorical |
| Respondent age (years) at interview at wave | Numerical |
| MedFeat vs. | mean | Holm-adj. |
| Baseline | 1.71 | 0.039* |
| OpenFE | 7.83 | 0.039* |
| CAAFE | 1.61 | 0.039* |
| FeatLLM | 20.18 | 0.039* |
| OCTree | 1.34 | 0.039* |
| Method | Cog-Impairment | Mortality-10yrs | Mortality-Ward | Discharge | Mortality-ICU | Heart Failure | Readmission |
| Baseline | 56.3 1.9 | 25.3 0.8 | 3.6 0.4 | 31.3 0.8 | 11.6 1.0 | 7.4 0.2 | 18.8 0.4 |
| OpenFE | 48.3 3.0 | 18.6 5.6 | 1.7 0.8 | 26.2 1.3 | 7.1 4.2 | 10.7 8.0 | 18.2 0.4 |
| CAAFE | 56.3 1.9 | 25.4 0.9 | 3.6 0.4 | 31.3 0.8 | 11.8 1.1 | 7.6 0.4 | 19.0 0.4 |
| FeatLLM | 34.2 2.6 | 12.8 3.0 | 2.1 1.3 | 17.4 0.2 | 2.1 0.0 | 4.5 0.0 | 15.7 0.0 |
| OCTree | 56.3 1.9 | 25.0 0.9 | 3.4 0.4 | 31.3 0.8 | 12.8 1.1 | 7.6 0.3 | 19.0 0.3 |
| MedFeat | 58.8 1.6 | 26.3 1.1 | 3.9 0.7 | 34.6 0.7 | 13.6 1.1 | 7.8 0.4 | 19.9 0.5 |
| Method | Cog-Impairment | Mortality-10yrs | Mortality-Ward | Discharge | Mortality-ICU | Heart Failure | Readmission |
| Baseline | 54.9 1.9 | 19.6 2.5 | 1.7 0.3 | 37.0 0.6 | 18.1 2.7 | 8.2 0.6 | 20.3 0.5 |
| OpenFE | 55.2 1.3 | 20.8 2.6 | 2.6 1.9 | 34.8 0.9 | 12.1 0.9 | 7.4 0.4 | 19.7 0.3 |
| CAAFE | 55.1 1.6 | 19.7 2.2 | 2.2 1.2 | 37.1 0.7 | 16.2 2.3 | 8.3 0.5 | 20.6 0.6 |
| FeatLLM | 31.0 1.1 | 11.5 2.0 | 0.3 0.0 | 17.3 0.0 | 2.1 0.0 | 4.5 0.0 | 15.7 0.0 |
| OCTree | 55.5 1.3 | 19.1 2.4 | 1.9 0.5 | 37.0 0.5 | 18.0 2.1 | 8.3 0.2 | 20.4 0.3 |
| MedFeat | 57.2 0.9 | 21.7 1.5 | 2.4 0.6 | 37.4 0.5 | 20.3 1.9 | 8.9 0.4 | 20.9 0.6 |
| Metric | Baseline | MedFeat | |
| Deaths caught (TP) | 36.0 | 39.2 | +3.2 |
| Deaths missed (FN) | 23.0 | 19.8 | –3.2 |
| False alarms (FP) | 5225.6 | 4318.2 | –907.4 |
| F1 | 1.4 | 1.8 | +0.4 |
| Method | Mortality-Ward | Mortality-10yrs | Mortality-ICU | |||
| F1 | AUROC | F1 | AUROC | F1 | AUROC | |
| Raw RF | 2.6 0.3 | 72.2 3.8 | 28.1 2.1 | 74.0 2.2 | 10.9 0.4 | 79.4 1.5 |
| MedFeat + GI | 2.7 0.3 | 73.8 3.2 | 29.4 3.1 | 74.6 2.4 | 11.4 1.9 | 79.7 1.1 |
| MedFeat + SHAP | 2.8 0.3 | 74.5 3.3 | 29.8 1.4 | 74.7 1.9 | 11.5 0.3 | 79.9 0.5 |
| Task | Baseline | MedFeat | ||
| F1 | AUROC | F1 | AUROC | |
| Mortality-ICU | 9.6 1.0 | 78.4 2.7 | 10.4 1.0 | 80.3 1.0 |
| Mortality-10yrs | 19.5 1.3 | 60.7 1.8 | 27.7 2.5 | 73.3 3.7 |
| Heart Failure | 13.8 0.5 | 65.5 1.4 | 13.9 0.6 | 66.2 1.4 |
| Mortality-Ward | 1.8 0.5 | 72.8 1.7 | 2.0 0.5 | 79.2 3.8 |
| Cog-Impairment | 58.9 1.3 | 76.4 1.4 | 59.4 1.3 | 77.1 1.5 |
| Task | Acc. | Rej. | Inv. | % Runs |
| Cog-Impairment | 6.0 | 4.8 | 0.1 | 100% |
| Mortality-10yrs | 4.1 | 5.3 | 0.2 | 90% |
| Mortality-Ward | 3.4 | 4.2 | 0.0 | 80% |
| Discharge | 4.5 | 4.0 | 0.3 | 90% |
| Mortality-ICU | 4.3 | 3.7 | 0.2 | 100% |
| Heart Failure | 6.2 | 4.1 | 0.1 | 100% |
| Task | Pattern | Example | Ref. |
| Mortality-Ward | Clinical scores | Early Warning Score (EWS) | Williams (2022) |
| Mortality-10yrs | Behavioral risk | Alcohol depression | Yan et al. (2025) |
| Heart Failure | Demographics + vitals | Age heart rate change | Ostrowska et al. (2024) |
| Readmission | Vital sign instability | Var(SpO 2 ) Var(respiratory rate) | Nguyen et al. (2017) |
| Method | F1 | AUROC | AUPRC |
| CAAFE | 25.1 0.7 | 70.6 1.2 | 22.4 1.0 |
| FeatLLM | – | 58.9 4.4 | 15.1 1.8 |
| MedFeat | 26.2 1.0 | 71.5 1.1 | 23.9 1.0 |
| Mortality-ICU | Mortality-10yrs | Heart Failure | Mortality-Ward | Cog-Impairment | Discharge | Readmission | |
| # features | 21 | 84 | 21 | 28 | 84 | 28 | 21 |
| XGB F1 | 10.1 0.5 | 25.2 2.1 | 13.7 0.4 ‡ | 1.5 0.3 | 56.8 1.3 | 40.5 0.5 | 29.5 1.0 † |
| XGB AUC | 76.5 1.4 | 68.8 2.1 | 65.5 1.9 † | 74.6 2.7 | 74.3 1.4 | 71.2 0.6 | 56.3 1.0 † |
| LR F1 | 9.2 0.8 | 27.5 3.7 † | 14.4 0.2 ‡ | 2.3 0.7 | 57.5 1.1 | 38.5 0.5 | 30.1 0.8 ‡ |
| LR AUC | 77.7 1.2 | 74.9 1.8 † | 68.1 1.1 ‡ | 84.0 1.4 | 75.7 1.0 | 69.6 0.2 | 57.4 0.8 † |
| Mortality-Ward | Discharge | Heart Failure | Mean | |
| 0.000 | +0.52 | +0.66 | +0.28 | +0.49 |
| 0.005 | +0.38 | +0.66 | +0.38 | +0.47 |
| 0.010 | +0.40 | +0.90 | +0.70 | +0.67 |
| 0.020 | +0.28 | +0.74 | +0.76 | +0.59 |