Pretrained biomedical vision-language models achieve strong zero-shot performance in biomedical image classification. However, downstream biomedical classification often depends on subtle visual differences between classes that may not be fully captured by pretrained representations. Few-shot adaptation addresses this mismatch by optimizing a task-specific predictor on a small labeled support set. Because the selected examples capture only part of the visual variation within the target classes, the adapted predictions can depend strongly on their composition. We propose Prompt-Anchored Residual Adaptation (PARA), which retains the frozen prompt prediction as a support-invariant semantic anchor and incorporates a visual prediction learned from the support set through an anchor-relative residual. The residual step is computed in a closed form from frozen support embeddings using anchor discrepancy and support agreement. Support-set dependence also limits evaluation: comparisons are fair within a shared draw but remain conditional on its composition. To obtain more reliable comparisons, we introduce a repeated-support protocol that separates support-selection variation from optimization randomness and reports both average and worst-20% performance. PARA achieves state-of-the-art performance in both few-shot classification and base-to-novel generalization.
Figures & tables
Figure 1: Lowest and highest Accuracy across 20 support draws for three prompt learning methods on COVID-19.
Figure 2: Overview of PARA. During training, the support set is used to optimize the LoRA parameters. At inference, all support examples form the adapted class prototypes, and a query image is scored by both the frozen prompt branch and the adapted visual branch. Their difference is scaled by η(S) and added to the anchor prediction, where η(S) aggregates class-wise anchor discrepancy and support agreement computed from frozen support embeddings. Panel (a) details prototype construction during training and inference, while panel (b) illustrates how the two geometry quantities jointly determine the residual step.
Accuracy
Macro-F1
Method
K=1
K=2
K=4
K=8
K=16
Avg.
K=1
K=2
K=4
K=8
K=16
Avg.
Zero-shot Method
BiomedCLIP
54.35
44.55
CLIP-Based Adaptation Methods
Tip-Adapter
55.12 ± 0.60
55.99 ± 0.81
57.16 ± 1.10
59.22 ± 1.45
61.91 ± 1.42
57.88 ± 1.08
45.46 ± 0.53
46.24 ± 0.78
47.63 ± 0.89
49.89 ± 1.25
52.91 ± 1.52
48.43 ± 1.00
CLIP-LoRA
54.82 ± 1.18
54.82 ± 1.09
55.23 ± 0.91
56.15 ± 1.15
59.71 ± 1.40
56.15 ± 1.15
46.53 ± 0.92
47.07 ± 0.86
47.77 ± 0.76
50.46 ± 0.88
55.33 ± 1.20
49.43 ± 0.92
Table 1: Comparison with state-of-the-art methods on 11 biomedical datasets (%). Entries report repeated-support means ± average within-dataset support-draw standard deviations. BiomedCLIP is support-invariant; best means are bold.
Shots per class ( K )
Method
1
2
4
8
16
Avg.
Zero-shot Method
BiomedCLIP
54.35
CLIP-Based Adaptation Methods
Tip-Adapter
54.27
54.86
55.59
57.13
59.71
56.31
CLIP-LoRA
53.06
53.31
53.92
54.50
57.67
54.49
Table 2: Performance over the worst 20% of support draws. Accuracy CVaR 20 (%) is averaged across 11 biomedical datasets. BiomedCLIP is invariant to support selection. The best adaptation result in each column is shown in bold.
Accuracy
Macro-F1
Mean ± SD
CVaR 20
Mean ± SD
CVaR 20
Method
Base
Novel
HM
Base
Novel
HM
Base
Novel
HM
Base
Novel
HM
BiomedCLIP
55.87
79.91
65.76
55.87
79.91
65.76
49.09
68.77
57.29
49.09
68.77
57.29
CLIP-LoRA
78.00 ± 2.23
72.56 ± 4.47
75.18 ± 2.76
74.62
66.40
70.27
76.41 ± 2.19
66.43 ± 3.53
71.07 ± 2.47
73.09
61.48
66.78
CoOp
72.54 ± 2.75
68.59 ± 5.07
70.51 ± 3.47
68.38
62.10
65.09
70.94 ± 2.58
56.00 ± 4.40
62.59 ± 3.31
66.98
49.75
57.09
CoCoOp
71.20 ± 2.94
68.37 ± 5.56
69.76 ± 3.69
66.82
60.52
63.51
69.54 ± 2.72
54.48 ± 5.03
61.10 ± 3.63
65.42
47.47
55.02
Table 3: Base-to-novel generalization over 20 support selections on 10 datasets (%). Adaptation uses 16-shot base-class support. SD is averaged across datasets. For both mean and CVaR 20 , HM is computed from the displayed Base and Novel values; HM SD is computed from the corresponding draw-wise harmonic means.
Figure 3: Accuracy performance profiles for PARA and the five baselines used in paired significance tests across 11 datasets. A curve gives the fraction of datasets on which a method is within a factor τ of the lowest error among the methods shown ( Dolan and Moré 2002 ) .
Figure 4: The base–novel trade-off. Base and novel Accuracy (%) averaged over 10 datasets and 20 support selections.
Variant
Mean Acc.
CVaR 20
SD
w/o Prototype Anchoring
66.66
61.70
3.27
w/o Prompt-Anchored Residual
64.98
57.75
4.84
w/o Support Geometry
64.78
60.17
3.17
PARA (Full)
66.95
62.41
3.06
Table 4: Component ablation. SD denotes the standard deviation across support draws.
Table A1: Dataset overview and the fixed split at the image level, stratified by class.
Figure A1: Few-shot Accuracy for each dataset. The Average panel is the mean across the 11 datasets at each shot level. Each adaptation point is the mean over 20 support draws after averaging optimizer seeds within each draw. BiomedCLIP denotes the zero-shot anchor at K=0 , which is independent of the sampled support set.
Figure A2: Few-shot Macro-F1 for each dataset under the same protocol as Figure A1 . The Average panel is the mean across the 11 datasets at each shot level.
Accuracy
Mean
CVaR 20
Baseline
Δ
Wins
pH
Δ
Wins
pH
Tip-Adapter
9.07
9/11
0.0244
6.09
9/11
0.0645
CLIP-LoRA
10.80
11/11
0.0117
7.91
10/11
0.0195
TaskRes
4.44
10/11
0.0244
6.65
10/11
0.0137
CLAP
6.57
11/11
0.0117
9.91
11/11
0.0117
Appendix
Table A2: Paired comparisons between PARA and all 12 adaptation baselines across 11 datasets. Δ is the mean paired difference in percentage points, Wins counts positive differences, and pH is adjusted with Holm’s method over the 12 baselines separately for each summary. (a) Accuracy.
Macro-F1
Mean
CVaR 20
Baseline
Δ
Wins
pH
Δ
Wins
pH
Tip-Adapter
12.91
11/11
0.0117
10.17
11/11
0.0117
CLIP-LoRA
11.90
11/11
0.0117
9.09
11/11
0.0117
TaskRes
2.37
9/11
0.0391
3.76
9/11
0.0244
CLAP
4.34
10/11
0.0156
6.73
11/11
0.0117
Appendix
Table A3: Paired comparisons between PARA and all 12 adaptation baselines (continued). (b) Macro-F1.
Dataset
BiomedCLIP
Tip-Adapter
CLIP-LoRA
TaskRes
CLAP
LP++
SS-Text+
TAMP
PARA
Accuracy: Mean
BTMRI
62.53
65.79
61.95
74.49
72.50
74.19
71.36
73.54
76.93
BUSI
54.95
56.30
57.98
59.89
59.57
60.03
56.70
59.81
63.31
CHMNIST
30.43
41.86
44.95
67.88
66.34
68.69
71.63
67.55
66.96
COVID-19
66.46
71.40
62.89
70.16
66.51
69.95
67.40
69.45
75.50
CTKidney
57.45
60.47
65.38
72.01
70.13
71.92
59.55
69.11
75.22
Appendix
Table A4: Few-shot results for each dataset across all methods in the main comparison (%). Each value is averaged over the five shot levels. Mean and CVaR 20 are shown in separate blocks. Bold marks the best result across all methods. (a) Frozen and CLIP-based adaptation methods.
Dataset
CoOp
CoCoOp
ProGrad
BiomedCoOp
vMFCoOp
PARA
Accuracy: Mean
BTMRI
72.93
68.06
72.92
74.24
74.66
76.93
BUSI
58.16
56.36
59.51
59.10
59.10
63.31
CHMNIST
64.30
58.07
66.77
68.69
66.53
66.96
COVID-19
69.96
63.84
65.20
73.78
70.97
75.50
CTKidney
68.03
64.32
67.76
69.24
68.80
75.22
Appendix
Table A5: Few-shot results for each dataset (continued). (b) Prompt-learning methods; PARA is repeated for comparison.
Biomedical Vision--Language Models (VLMs) have shown remarkable promise in few-shot medical diagnosis but face a critical bottleneck: \textit{fragility to prompt variations}.Existing adaptation frameworks typically optimize visual and textual prompts as independent streams, relying on ideal ``Golden Prompts''. In clinical reality, where descriptions are often noisy and heterogeneous, this modality isolation leads to unstable cross-modal alignment. To address this, we propose BiomedAP, a vision-informed dual-anchor framework with gated cross-modal fusion.BiomedAP enforces synergistic alignment through two mechanisms: (1) Gated Cross-Modal Fusion, which enables layer-wise interaction between modalities, acting as a dynamic noise regulator to suppress irrelevant textual cues; and (2) a Dual-Anchor Constraint that regularizes learnable prompts toward stable semantic centroids derived from both expert templates (High Anchors) and few-shot visual prototypes (Low Anchors). Extensive experiments across 11 benchmarks demonstrate that BiomedAP consistently surpasses baselines, achieving competitive few-shot accuracy and markedly enhanced robustness under prompt perturbations. Our code is available at: https://github.com/tongdiedie/BiomedAP. Keywords: Vision-Language Models; Prompt Learning; Parameter-Efficient Fine-Tuning; Few-shot Learning
Accurate biomedical image classification under low-resource conditions remains challenging due to limited annotations, subtle inter-class visual differences, and complex disease semantics. While vision--language models offer a promising foundation for mitigating data scarcity, their effective adaptation in biomedical settings is constrained by the need for parameter-efficient tuning alongside fine-grained and semantically consistent representation learning. In this work, we propose Multi-View Synergistic Learning (MVSL), a unified framework that addresses these challenges by jointly considering adaptation paradigms, representation granularity, and disease semantic relationships. MVSL decouples the adaptation of visual and textual encoders to respect their distinct representational characteristics, enabling more stable and effective parameter-efficient fine-tuning. It further introduces multi-granularity contrastive learning to explicitly model both global image semantics and localized lesion-level evidence, improving fine-grained discrimination for visually similar disease categories. In addition, MVSL preserves disease-level semantic structure by incorporating structured supervision derived from large language models, which constrains textual representations at the class level and indirectly regularizes visual embeddings through cross-modal alignment. Together, these components enable more stable cross-modal alignment and improved discrimination under limited supervision. Extensive experiments on 11 public biomedical datasets spanning 9 imaging modalities and 10 anatomical regions demonstrate that MVSL consistently outperforms state-of-the-art methods in few-shot and zero-shot classification settings.
Xiaoliu Luo, Minxue Xiao, Ting Xie +5
Chongqing University of Technology · Hebei University of Technology · Nanyang Technological University +3
Medical Vision-Language Models (VLMs) exhibit strong zero-shot performance, yet their effectiveness still declines on out-of-distribution (OOD) data due to domain shifts and class bias inherited from large-scale pretraining. Existing few-shot adaptation methods typically introduce additional trainable components, which can be unstable in extremely low-data regimes (e.g., 1-shot), and lack robustness on different medical data. We present TCLA, a purely training-free few-shot adaptation method for Medical VLMs, which is fast and model-agnostic. TCLA corrects inference logits based on a small set of support samples, boosting pretrained VLMs performance by improving inter-class deconfusion and reducing domain shift. Extensive experiments on nine datasets across multiple medical imaging modalities including X-ray, Ultrasound, MRI, CT, Histopathology, demonstrate that TCLA consistently improves OOD performance of Medical VLMs and, in most of cases, outperforms existing training-based adaptation methods.
Tianyou Jiang, Ziyu Zhou
University of Bern, Bern, Switzerland · Shanghai Jiao Tong University, Shanghai, China