Large pretrained clinical models provide a practical way to reuse learned prior knowledge across hospitals by adapting models to them. In practice, a target hospital may only have a small labeled patient cohort, a setting commonly referred to as few-shot adaptation. This requires making multiple decisions, such as which pretrained model to adapt, how much of the model to update, and which patients to use. Nevertheless, this process faces two primary challenges. First, the best adaptation strategy varies across clinical tasks. Second, evaluating and comparing candidate strategies becomes unreliable due to the small patient cohort. In this work, we introduce AutoAdapt with two core designs to deal with these challenges. The Adapter defines an extensible space of adaptation recipes, and the Automator forms a weighted recipe combination from evidence within the adaptation patients. We propose a reliability rule to ensure that only the most effective strategy on most available patients will be selected. These selected strategies then form a combination for effective few-shot adaptation. We conduct extensive experiments across critical care, emergency care, and diagnostic datasets, and the results show that AutoAdapt consistently achieves state-of-the-art performance using only a few patients for adaptation.
Figures & tables
Figure 1: Independent patient support in the MIMIC-IV cohorts.
Figure 2: Overview of AutoAdapt . The Adapter defines a recipe through its source representation, update capacity, target data policy, and input route. The Automator scores complete recipes using held-out adaptation patients. A shared evidence map forms a compact weighted combination, while the reliability check protects a predefined default recipe when the comparison is uncertain. Participating recipes are refit on the adaptation cohort before evaluation begins.
Method
Ventr.
Crit. rhythm
Oxy. crisis
Short VT
Pause/ asystole
ST shift
Mean
Fixed recipe baselines
Frozen source encoder + target head
.512
.556
.515
.516
.567
.622
.548
Full fine-tuning
.642
.663
.560
.677
.654
.731
.655
LoRA
.640
.662
.557
.672
.656
.716
.651
Existing adaptation methods
PRAM ( Jeong et al., 2026 )
.517
.515
.516
.545
.487
.518
.516
Table 1: MIMIC-III → ALOTT few-shot comparison on six clinically central tasks selected for clinical relevance, patient support, waveform coverage, and task difficulty. Values are patient-level AUROC. Every method uses the same 32 positive and 32 negative adaptation patients per task.
Method
Acute
Circ.
SOFA
Delir.
AKI
Mort.
Mean
Fixed recipe baselines
Frozen source encoder + target head
.545
.596
.640
.547
.434
.637
.567
Full fine-tuning
.538
.585
.591
.519
.465
.665
.560
LoRA
.564
.586
.606
.513
.477
.685
.572
Existing-work baselines
PRAM ( Jeong et al., 2026 )
.574
.605
.669
.548
.468
.649
.585
Table 2: MIMIC-IV scarce-label stress test using patient-level AUROC averaged over five evaluation runs. Every method uses the same adaptation and evaluation patients in each run. Underlining marks the strongest non-AutoAdapt result and boldface marks the highest result.
Figure 3: Checkpoint and recipe analysis on MIMIC-IV. Panel A tests source checkpoint and head pairings. Panel B compares a fixed recipe, naive adaptation-only selection, and the reliability-aware Automator. Panel C adds the Adapter and Automator components cumulatively.
Figure 4: Cross-dataset comparisons. Panel A reports MC-MED macro AUROC over four tasks with sufficient patient support, comparing local EHR, waveform transfer, fixed fusion, and AutoAdapt . Panel B reports macro AUROC over five shared diagnostic ECG tasks in CPSC and Ningbo. Methods use the same target folds within each dataset. The reference line marks chance AUROC.
Figure 5: Checkpoint routing under distribution shift. Panels A and B show how transfer gain changes with task distance and waveform distance. Panel C compares routing rules. Every point uses the same target patients. Evaluation labels measure transfer gain but do not select a checkpoint.
Figure 6: Stability on ALOTT new-pressor prediction as the adaptation panel grows from one to 32 positive patients, with an equal number of negative patients. Panel A reports evaluation AUROC over ten repeated panels against two fixed baselines. Panel B reports the number of distinct selections and the fraction using a recipe combination.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Prior reference
AutoAdapt
Gain
AF vs sustained VT/VF alarm proxy
.727
.793
+.066
Critical rhythm vs short VT severity proxy
.612
.655
+.043
Critical rhythm vs ST-shift severity proxy
.663
.697
+.033
Critical rhythm vs VPC severity proxy
.597
.647
+.050
Critical vs noncritical electrical proxy
.569
.617
+.047
Pause vs sustained VT/VF proxy
.595
.653
+.057
Appendix
Table 3: The fourteen ALOTT tasks not included in the six-task main panel. Values are patient-level AUROC on the same evaluation cohorts used for the main comparison.
All eligible
Adaptation/run
Evaluation/run
Task
O/G/H (h)
Patients
Positive
Patients
Positive
Patients
Positive
Acute
6/0/6
87
30
17–18
6
69–70
24
Circ.
6/0/6
139
26
27–29
5–6
110–112
20–21
SOFA
6/6/24
147–148
25–26
29–30
5
117–119
20–21
Delir.
6/3/9
91
22
17–19
4–5
72–74
17–18
AKI
6/1/23
46–47
19
9–10
4
36–37
15
Appendix
Table 4: Independent patient support in the five MIMIC-IV runs. Ranges show variation across patient partitions. Evaluation counts include patients with a usable prediction. Waveform row counts are omitted because repeated measurements are not independent events.
Checkpoint
MIMIC-III training state
Saved model
Target use
Pretrained
Masked time and frequency pretraining at selected step 1,300
37,699,612
Encoder with target fitted head
Deterioration
Pretrained start with escalation or death supervision at epoch 8
37,700,125
Encoder with aligned acute source head
Circulatory
Pretrained start with five hemodynamic labels at epoch 5
37,702,177
Encoder with aligned pressor source head
Respiratory
Pretrained start with five respiratory labels at epoch 15
37,702,177
Encoder only for main targets
Appendix
Table 5: Source checkpoint library. Specialized checkpoints start from the general pretrained model and receive supervised MIMIC-III training. A source classifier is reused only when its label matches the target contract.
Target
Setting and input
Positive support
Purpose
ALOTT
ICU with EHR and bedside waveforms
16–779
Primary 64-patient transfer
MIMIC-IV
ICU waveform with EHR extension
8–30
Scarce label stress test
MC-MED
ED with EHR and bedside waveforms
23–162
Reliability of the default recipe
CPSC and Ningbo
Diagnostic 12 lead ECG
Native folds
Acquisition shift
Appendix
Table 6: Target datasets and evaluation scope. Support gives positive evaluation patients after waveform quality control for ALOTT and MC-MED. The MIMIC-IV entry gives the full eligible range.
Method
Target labels
Encoder update
Target operation
Traditional baselines
Source head
No
None
Reuse the original source classifier unchanged
Frozen + head
Yes
None
Fit a new logistic head on frozen embeddings
Frozen + cosine prototype
Yes
None
Compare with adaptation class centroids using cosine similarity
Frozen + Euclidean prototype
Yes
None
Compare with adaptation class centroids using Euclidean distance
Frozen + cosine 5-NN
Yes
None
Vote among five nearby adaptation examples
Appendix
Table 7: Target adaptation performed by each comparison. Candidate models and combination weights use labels from the declared adaptation patients.
D2 method
Min. pos.
Preferred pos.
Rank
Updated params.
Aligned source head
0
0
0
0
Frozen + target head
4
4
1
head only
LoRA
4
8
2
197,121
Partial fine tuning
4
8
3
6,306,305
Full fine tuning
4
20
4
26.81M
Appendix
Table 8: Fixed Automator thresholds. Minimum support controls eligibility. A challenger below the preferred count also needs a favorable paired confidence bound. Complexity resolves task-utility differences within δQ . A source head is reused only for an aligned label.
Source or target
Concept set
Pretrained source
∅
Deterioration source
death, mortality, organ support, respiratory, circulatory, renal, ventilation, pressor, RRT
Table 9: Clinical concepts used for task distance. Tokens are fixed before analysis. The general pretrained encoder has no outcome specific concept prior, so its set is empty. It remains eligible through distribution and adaptation evidence.