Large pretrained clinical models provide a practical way to reuse learned prior knowledge across hospitals by adapting models to them. In practice, a target hospital may only have a small labeled patient cohort, a setting commonly referred to as few-shot adaptation. This requires making multiple decisions, such as which pretrained model to adapt, how much of the model to update, and which patients to use. Nevertheless, this process faces two primary challenges. First, the best adaptation strategy varies across clinical tasks. Second, evaluating and comparing candidate strategies becomes unreliable due to the small patient cohort. In this work, we introduce AutoAdapt with two core designs to deal with these challenges. The Adapter defines an extensible space of adaptation recipes, and the Automator forms a weighted recipe combination from evidence within the adaptation patients. We propose a reliability rule to ensure that only the most effective strategy on most available patients will be selected. These selected strategies then form a combination for effective few-shot adaptation. We conduct extensive experiments across critical care, emergency care, and diagnostic datasets, and the results show that AutoAdapt consistently achieves state-of-the-art performance using only a few patients for adaptation.
Figures & tables
Figure 1: Independent patient support in the MIMIC-IV cohorts.
Figure 2: Overview of AutoAdapt . The Adapter defines a recipe through its source representation, update capacity, target data policy, and input route. The Automator scores complete recipes using held-out adaptation patients. A shared evidence map forms a compact weighted combination, while the reliability check protects a predefined default recipe when the comparison is uncertain. Participating recipes are refit on the adaptation cohort before evaluation begins.
Method
Ventr.
Crit. rhythm
Oxy. crisis
Short VT
Pause/ asystole
ST shift
Mean
Fixed recipe baselines
Frozen source encoder + target head
.512
.556
.515
.516
.567
.622
.548
Full fine-tuning
.642
.663
.560
.677
.654
.731
.655
LoRA
.640
.662
.557
.672
.656
.716
.651
Existing adaptation methods
PRAM ( Jeong et al., 2026 )
.517
.515
.516
.545
.487
.518
.516
Table 1: MIMIC-III → ALOTT few-shot comparison on six clinically central tasks selected for clinical relevance, patient support, waveform coverage, and task difficulty. Values are patient-level AUROC. Every method uses the same 32 positive and 32 negative adaptation patients per task.
Method
Acute
Circ.
SOFA
Delir.
AKI
Mort.
Mean
Fixed recipe baselines
Frozen source encoder + target head
.545
.596
.640
.547
.434
.637
.567
Full fine-tuning
.538
.585
.591
.519
.465
.665
.560
LoRA
.564
.586
.606
.513
.477
.685
.572
Existing-work baselines
PRAM ( Jeong et al., 2026 )
.574
.605
.669
.548
.468
.649
.585
Table 2: MIMIC-IV scarce-label stress test using patient-level AUROC averaged over five evaluation runs. Every method uses the same adaptation and evaluation patients in each run. Underlining marks the strongest non-AutoAdapt result and boldface marks the highest result.
Figure 3: Checkpoint and recipe analysis on MIMIC-IV. Panel A tests source checkpoint and head pairings. Panel B compares a fixed recipe, naive adaptation-only selection, and the reliability-aware Automator. Panel C adds the Adapter and Automator components cumulatively.
Figure 4: Cross-dataset comparisons. Panel A reports MC-MED macro AUROC over four tasks with sufficient patient support, comparing local EHR, waveform transfer, fixed fusion, and AutoAdapt . Panel B reports macro AUROC over five shared diagnostic ECG tasks in CPSC and Ningbo. Methods use the same target folds within each dataset. The reference line marks chance AUROC.
Figure 5: Checkpoint routing under distribution shift. Panels A and B show how transfer gain changes with task distance and waveform distance. Panel C compares routing rules. Every point uses the same target patients. Evaluation labels measure transfer gain but do not select a checkpoint.
Figure 6: Stability on ALOTT new-pressor prediction as the adaptation panel grows from one to 32 positive patients, with an equal number of negative patients. Panel A reports evaluation AUROC over ten repeated panels against two fixed baselines. Panel B reports the number of distinct selections and the fraction using a recipe combination.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Prior reference
AutoAdapt
Gain
AF vs sustained VT/VF alarm proxy
.727
.793
+.066
Critical rhythm vs short VT severity proxy
.612
.655
+.043
Critical rhythm vs ST-shift severity proxy
.663
.697
+.033
Critical rhythm vs VPC severity proxy
.597
.647
+.050
Critical vs noncritical electrical proxy
.569
.617
+.047
Pause vs sustained VT/VF proxy
.595
.653
+.057
Appendix
Table 3: The fourteen ALOTT tasks not included in the six-task main panel. Values are patient-level AUROC on the same evaluation cohorts used for the main comparison.
All eligible
Adaptation/run
Evaluation/run
Task
O/G/H (h)
Patients
Positive
Patients
Positive
Patients
Positive
Acute
6/0/6
87
30
17–18
6
69–70
24
Circ.
6/0/6
139
26
27–29
5–6
110–112
20–21
SOFA
6/6/24
147–148
25–26
29–30
5
117–119
20–21
Delir.
6/3/9
91
22
17–19
4–5
72–74
17–18
AKI
6/1/23
46–47
19
9–10
4
36–37
15
Appendix
Table 4: Independent patient support in the five MIMIC-IV runs. Ranges show variation across patient partitions. Evaluation counts include patients with a usable prediction. Waveform row counts are omitted because repeated measurements are not independent events.
Checkpoint
MIMIC-III training state
Saved model
Target use
Pretrained
Masked time and frequency pretraining at selected step 1,300
37,699,612
Encoder with target fitted head
Deterioration
Pretrained start with escalation or death supervision at epoch 8
37,700,125
Encoder with aligned acute source head
Circulatory
Pretrained start with five hemodynamic labels at epoch 5
37,702,177
Encoder with aligned pressor source head
Respiratory
Pretrained start with five respiratory labels at epoch 15
37,702,177
Encoder only for main targets
Appendix
Table 5: Source checkpoint library. Specialized checkpoints start from the general pretrained model and receive supervised MIMIC-III training. A source classifier is reused only when its label matches the target contract.
Target
Setting and input
Positive support
Purpose
ALOTT
ICU with EHR and bedside waveforms
16–779
Primary 64-patient transfer
MIMIC-IV
ICU waveform with EHR extension
8–30
Scarce label stress test
MC-MED
ED with EHR and bedside waveforms
23–162
Reliability of the default recipe
CPSC and Ningbo
Diagnostic 12 lead ECG
Native folds
Acquisition shift
Appendix
Table 6: Target datasets and evaluation scope. Support gives positive evaluation patients after waveform quality control for ALOTT and MC-MED. The MIMIC-IV entry gives the full eligible range.
Method
Target labels
Encoder update
Target operation
Traditional baselines
Source head
No
None
Reuse the original source classifier unchanged
Frozen + head
Yes
None
Fit a new logistic head on frozen embeddings
Frozen + cosine prototype
Yes
None
Compare with adaptation class centroids using cosine similarity
Frozen + Euclidean prototype
Yes
None
Compare with adaptation class centroids using Euclidean distance
Frozen + cosine 5-NN
Yes
None
Vote among five nearby adaptation examples
Appendix
Table 7: Target adaptation performed by each comparison. Candidate models and combination weights use labels from the declared adaptation patients.
D2 method
Min. pos.
Preferred pos.
Rank
Updated params.
Aligned source head
0
0
0
0
Frozen + target head
4
4
1
head only
LoRA
4
8
2
197,121
Partial fine tuning
4
8
3
6,306,305
Full fine tuning
4
20
4
26.81M
Appendix
Table 8: Fixed Automator thresholds. Minimum support controls eligibility. A challenger below the preferred count also needs a favorable paired confidence bound. Complexity resolves task-utility differences within δQ . A source head is reused only for an aligned label.
Source or target
Concept set
Pretrained source
∅
Deterioration source
death, mortality, organ support, respiratory, circulatory, renal, ventilation, pressor, RRT
Table 9: Clinical concepts used for task distance. Tokens are fixed before analysis. The general pretrained encoder has no outcome specific concept prior, so its set is empty. It remains eligible through distribution and adaptation evidence.
Pretrained biomedical vision-language models achieve strong zero-shot performance in biomedical image classification. However, downstream biomedical classification often depends on subtle visual differences between classes that may not be fully captured by pretrained representations. Few-shot adaptation addresses this mismatch by optimizing a task-specific predictor on a small labeled support set. Because the selected examples capture only part of the visual variation within the target classes, the adapted predictions can depend strongly on their composition. We propose Prompt-Anchored Residual Adaptation (PARA), which retains the frozen prompt prediction as a support-invariant semantic anchor and incorporates a visual prediction learned from the support set through an anchor-relative residual. The residual step is computed in a closed form from frozen support embeddings using anchor discrepancy and support agreement. Support-set dependence also limits evaluation: comparisons are fair within a shared draw but remain conditional on its composition. To obtain more reliable comparisons, we introduce a repeated-support protocol that separates support-selection variation from optimization randomness and reports both average and worst-20% performance. PARA achieves state-of-the-art performance in both few-shot classification and base-to-novel generalization.
Jingxuan Kang, Qianying Yue, Che Liu +1
Imperial College London · The Chinese University of Hong Kong
Medical Vision-Language Models (VLMs) exhibit strong zero-shot performance, yet their effectiveness still declines on out-of-distribution (OOD) data due to domain shifts and class bias inherited from large-scale pretraining. Existing few-shot adaptation methods typically introduce additional trainable components, which can be unstable in extremely low-data regimes (e.g., 1-shot), and lack robustness on different medical data. We present TCLA, a purely training-free few-shot adaptation method for Medical VLMs, which is fast and model-agnostic. TCLA corrects inference logits based on a small set of support samples, boosting pretrained VLMs performance by improving inter-class deconfusion and reducing domain shift. Extensive experiments on nine datasets across multiple medical imaging modalities including X-ray, Ultrasound, MRI, CT, Histopathology, demonstrate that TCLA consistently improves OOD performance of Medical VLMs and, in most of cases, outperforms existing training-based adaptation methods.
Tianyou Jiang, Ziyu Zhou
University of Bern, Bern, Switzerland · Shanghai Jiao Tong University, Shanghai, China
Uncertainty estimation for medical vision--language models (VLMs) using conformal prediction has gained increasing attention due to its distribution-free coverage guarantees. However, standard conformal prediction relies on exchangeability between calibration and test data and typically requires a sufficiently large calibration set to obtain reliable coverage. These assumptions are difficult to satisfy in few-shot transfer settings, where only a small labeled support set is available to adapt a pretrained VLM to a new medical task, while an unlabeled query set is used for evaluation. Supervised fine-tuning on the support set changes the model parameters and consequently shifts the nonconformity score distribution, breaking exchangeability between calibration and query samples and leading to unreliable coverage under distribution shift. Existing transductive conformal adaptation methods often preserve validity by avoiding supervised updates. While this helps maintain conformal assumptions, it underutilizes the scarce labeled support data and limits task adaptation, which is the primary objective in few-shot learning. In this setting, conformal prediction should serve as an uncertainty estimation layer that supports the adapted model, rather than preventing adaptation itself. To this end, we propose AlignCP, a framework that reconciles supervised few-shot adaptation with conformal uncertainty estimation under non-exchangeability. AlignCP learns a reweighted calibration distribution that reduces the score-level discrepancy between the labeled support set and the unlabeled query set. By aligning the one-dimensional nonconformity score distributions, AlignCP aims to close the coverage gap induced by adaptation without requiring query labels.
Xuan Cuong Ngo, Ngan Le
University of Arkansas Fayetteville, Arkansas, USA