Scientific image segmentation methods rely on extensive annotation and task-specific training, limiting adaptation across imaging modalities and experimental conditions. We present a training-free, one-shot framework that specializes vision foundation models using a single annotated reference image. The framework combines DINOv3 representations with background-adaptive feature orthogonalization to suppress artifact-related feature directions, after which cosine similarity localizes candidate regions for SAM segmentation. We evaluate the framework on red-blood-cell microscopy, structured-illumination pool boiling, and chest radiography. Relative to the strongest baseline, the proposed method improves mean IoU by 5.91% and 78.62% on the microscopy and pool-boiling datasets, respectively, while achieving comparable performance on chest radiographs. These results demonstrate that one-shot reference conditioning can adapt general-purpose vision models to specialized scientific segmentation tasks.
Figures & tables
Figure 1 : Generalist-to-specialist one-shot segmentation across scientific imaging domains. A single annotated reference defines each target concept, and the same frozen DINOv3–SAM framework localizes and segments corresponding structures without task-specific training.
Figure 2 : End-to-end one-shot segmentation framework. DINOv3 features from the target image are compared with a foreground prototype obtained from the annotated reference. The resulting similarity map is thresholded to localize candidate regions. A distance transform identifies interior point prompts, while connected regions provide bounding-box prompts for SAM2, which produces the final segmentation masks.
Dataset
Imaging Modality
Method
Mean Target IoU
Mean Target Dice (F1)
PathOlOgics (RBCs)
Microscopy
SAM2
0.8670
0.9201
PathOlOgics (RBCs)
Microscopy
GF-SAM
0.8720
0.9297
PathOlOgics (RBCs)
Microscopy
INSID3
0.7670
0.8644
PathOlOgics (RBCs)
Microscopy
Ours
0.9235
0.9591
Pool Boiling
SI Reconstructed
SAM2
0.0600
0.1127
Pool Boiling
SI Reconstructed
GF-SAM
0.1646
0.2761
Table 1 : One-shot segmentation performance across scientific imaging domains. Best results are shown in bold and second-best results are underlined.
Figure 3 : Qualitative segmentation comparison on structured-illumination pool-boiling data. The proposed method recovers more of the required annotated features than SAM2 and GF-SAM while accurately reducing the enlarged and merged regions produced by INSID3.
Figure 4 : Effect of background-adaptive feature orthogonalization on an RMS-reconstructed pool-boiling frame. Compared with unprojected DINOv3 features and the Gaussian-noise projection, the proposed method suppresses background responses while preserving target-bubble activation.
Segmenting a new biomedical dataset usually means a domain-specific model trained on substantial annotation, or a foundation model steered at inference time. We present Exemplar, a few-shot segmenter that fuses a frozen DINOv3 backbone with a fixed bank of classical native-resolution filter responses in one lightweight head, fitted from the support masks alone. In the few-mask, native-resolution regime, classical priors and frozen self-supervised features are complementary: fused in one head, a single fixed configuration spans eleven biomedical imaging datasets. Under the same head, the classical bank alone reaches 0.693 on the eleven-dataset panel, scored by foreground intersection-over-union or centreline Dice, and the frozen features alone 0.672; the bank leads on seven of the eleven and the features on the rest, and fused they reach 0.782. Against five forward-pass few-shot methods, Exemplar leads in 54 of 55 method-dataset comparisons, 52 of them significant after Holm correction. From a single annotated mask it reaches 0.703 on the same panel, against 0.682 for a from-scratch nnU-Net trained on that same mask. At eight masks nnU-Net overtakes it on the panel mean, chiefly on centreline agreement, but takes 16-77x longer to fit.
Michal Průšek, Adam Novozámský, Filip Šroubek
The Czech Academy of Sciences, Institute of Information Theory and Automation, Czechia
Medical image segmentation remains difficult to scale because high-performing methods typically rely on dense expert annotations and task-specific training. We introduce GazeRefine, a training-free framework that uses gaze as an inference-time prompt for zero-shot medical image segmentation. Sparse, duration-weighted fixations are converted into foreground and background priors that initialize semantic prototypes in frozen DINOv3 feature space. These prototypes are iteratively refined through foreground-background discrimination, feature-space affinity propagation, and anchoring to the initial gaze guidance, allowing segmentation to extend beyond directly fixated regions while limiting semantic drift. GazeRefine requires no segmentation masks, fine-tuning, adapters, prompt encoders, or gradient updates. We evaluate the method on gaze-annotated polyp segmentation and prostate MRI segmentation. The results show strong performance on colonoscopy images and competitive performance on prostate MRI, supporting gaze-guided prototype refinement as a promising approach for segmentation-label-efficient, human-in-the-loop medical image segmentation. Our tools and code can be found in the following repository: https://github.com/MohammedOussamaBEN/GazeRefine.git
Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri +10
Université Sorbonne Paris Nord · Northwestern University · VSB-Technical University of Ostrava
Few-shot medical image segmentation (FS-MIS) aims to segment novel regions of interest (ROIs) from a few annotated support examples. Despite rapid progress, existing FS-MIS solutions span diverse paradigms but are evaluated under inconsistent settings, leaving their relative effectiveness unclear. We introduce FAME, a unified benchmark for evaluating FS-MIS solutions, covering specialists, SAM-based methods, CLIP-based methods, and MLLM-based methods. FAME contains 14,958 test samples across 7 anatomical sites, 9 imaging modalities, and 14 ROI categories, and evaluates models under zero-shot and ten-shot settings with additional assessment of target-absence recognition and generalization under covariate and semantic shifts. Our evaluation reveals several findings. First, effective few-shot segmentation depends on how models exploit support examples: direct visual adaptation generally outperforms prompt-based strategies. Second, increasing support examples improves performance only when models can effectively utilize them. Third, semantic transfer remains substantially more challenging than imaging-domain adaptation, and strong localization ability does not necessarily imply reliable target-absence recognition. We hope FAME provides a comprehensive understanding of current FS-MIS solutions and facilitates the development of more effective and reliable few-shot medical segmentation methods.