Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining distribution, yet fail on specialized domains where the required discriminative features were never learned, a regime we term distant out-of-distribution (OOD). Standard adaptation methods cannot overcome this representational absence because they operate within the encoder's existing feature space. However, VLMs retain a robust descriptive capacity even when discrimination collapses: a model that cannot classify a medical scan can still articulate its visual patterns. Exploiting this asymmetry, we introduce Inductive Visual Logic (IVL), a training-free framework that constructs classification knowledge from the model's surviving descriptive ability. IVL extracts visual traits from few-shot support images through dual-mode prompting, combining semantic descriptions with primitive visual observations, and organizes them into per-class trait dictionaries. At inference, hierarchical filtering identifies spatially grounded trait evidence for classification. Across multiple distant-OOD benchmarks, IVL achieves the highest aggregate accuracy under two VLM backbones while producing interpretable, trait-traceable predictions.
Figures & tables
Figure 1 : Distant-OOD vs. near-OOD characterization. (a) Two-dimensional projection based on feature-space divergence (MK-MMD) and latent cluster-label alignment (NMI). Distant-OOD datasets fall outside the low-divergence, high-alignment quadrant where parametric methods succeed. (b) Radar chart of accuracy gains after SFT across datasets, showing that fine-tuning yields substantial improvements only in near-OOD settings but only marginal gains on specialized domains.
Figure 2 : Reasoning comparison between gradient-based methods and IVL. Visual-RFT produces answers before articulating reasoning, deviating from human cognitive patterns. IVL mirrors human inductive learning by first observing visual traits, comparing them against a structured traits dictionary, refining evidence with prior knowledge, and arriving at an evidence-grounded classification decision.
Figure 3 : Offline Trait Building : (a) The system extracts semantic ( tiS ) and low-level ( tiL ) traits from support images (b) clusters them per class and (c) assigns canonical names, and stores them in database T . Inference : Given test image xtest , hierarchical filtering is performed: (d) compute sg=E^T⊤v^global to select top- k1 traits T(1) , (e) compute image patch-trait attention scores for localized grounding while selecting top- k2 traits T(2) and (f) extracting top-3 regions via NMS, and (g) prompt the VLM with the original image, cropped regions, and the refined trait set T(2) for final classification.
Figure 4 : Localized Trait Refinement. Given the test image and the globally filtered trait list T(1) , the VLM produces cross-modal attention maps from the grounding heads {L1,L2,L3} . After averaging and spatial smoothing, each trait’s attention map is flattened into a P -dimensional patch-level vector. The localization score sil is computed as the maximum attention over patches, and the top- k2 traits are selected to form T(2) .
Dataset
Zero-Shot
MajorVote
ICL
CuPL [ 30 ]
Menon et al . [ 23 ]
V-RFT [ 18 ]
SFT+LoRA
SFT
Tip-Adapter [ 45 ]
CoOp [ 49 ]
IVL (Ours)
Qwen2.5-VL-7B
Pokémon
47.47
48.23
22.00
15.28
25.77
48.77
45.59
45.52
9.72
20.01
50.30
Retinal OCT
14.36
16.39
13.75
12.50
11.46
14.00
13.69
14.77
12.50
13.39
20.83
WM811k
13.13
13.73
12.41
12.77
12.29
9.52
9.16
12.41
11.08
9.30
16.51
8-shot
Pokémon
–
–
34.24
–
–
49.09
50.09
49.94
6.94
29.17
52.20
Table 1 : Experimental results on distant-OOD benchmarks. Zero-shot and description-based methods (CuPL, Menon) use 0-shot. Bold : best per dataset within each backbone. Underline : second best. All few-shot results use n=1 unless marked 8-shot.
Bottle
Cable
Capsule
Carpet
Grid
Hazelnut
Leather
Metal Nut
Pill
Screw
Tile
Toothbrush
Transistor
Wood
Zipper
Mean
Semantic
39.2
14.7
35.7
36.9
34.3
52.7
28.0
26.7
5.9
32.3
57.7
51.7
15.4
40.2
17.0
32.6
Low-level
40.5
30.3
35.7
50.8
36.6
46.7
39.5
31.2
10.5
30.3
52.6
72.5
14.4
47.9
29.3
37.9
Dual
46.8
45.4
44.4
60.4
43.1
72.4
48.3
40.9
22.0
33.1
65.8
65.0
11.6
65.8
26.6
46.1
Δ
+6.3
+15.1
+8.7
+9.6
+6.5
+19.7
+8.8
+9.7
+11.5
+0.8
+8.1
−7.5
−3.8
+17.9
−2.7
+7.2
Table 2 : Prompt design ablation on MVTec AD (Qwen2.5-VL, 1-shot, 3-run mean per mode). Semantic: knowledge-driven prompting. Low-level: observation-driven prompting. Δ : dual-mode gain over best single mode.
Component
Configuration
Acc. (%)
Δ
Extraction ( Sec. 3.5 )
Semantic only
32.6
−6.2
Low-level only
27.8
−11.1
Filtering ( Sec. 3.6 )
Global only
39.6
+0.7
Local only
37.5
−1.4
No filtering
36.8
−2.1
Quality ( Sec. 3.5 )
w/o salience
36.3
−2.6
Table 3 : Component ablation on the Pokémon benchmark (144-sample balanced subset, Qwen2.5-VL, 1-shot). Δ : change relative to full pipeline.
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Symbol
Description
Problem Setting
f
Generative VLM (composed of ϕv and ϕl )
ϕv
Vision encoder
ϕl
Language decoder
Ptrain
Pretraining distribution
Ptarget
Target distribution
D={(xi,yi)}i=1n×C
Few-shot support set
Appendix
Table 4 : Summary of notation.
Dataset
MK-MMD
NMI
0-shot (%)
ΔSFT (%)
Regime
Reference
ImageNet
—
0.881
—
—
(reference)
Near-OOD (low MK-MMD, high NMI)
Pets-37
0.174
0.651
80.9
+6.3 †
Near-OOD
Cars-196
0.289
0.750
56.2
+27.3
Near-OOD
Flowers-102
0.310
0.975
72.4
+23.9
Near-OOD
Appendix
Table 5 : Representational absence metrics and adaptation outcomes across all evaluated datasets (Qwen2.5-VL-7B). Regime assignment is determined by MK-MMD and NMI thresholds; ΔSFT is reported for validation. Datasets are grouped by regime and sorted by MK-MMD within each group.
Figure 5 : Two-stage OOD taxonomy under Qwen2.5-VL ( left ) and LLaVA-1.6 ( right ). Thresholds are determined independently by two-stage Otsu applied to each encoder’s own metric distribution: τMMD=0.49 for both; τNMI=0.42 (Qwen2.5-VL) and τNMI=0.45 (LLaVA-1.6).
Zero-Shot
SFT
CoOp
IVL (Ours)
1-shot
18.26
28.14
2.93
31.34
8-shot
—
29.13
5.63
33.18
Appendix
Table 6 : Adaptation results on Pets-37 under LLaVA-1.6-Mistral-7B. The dataset is assigned to the distant-OOD regime by the MK-MMD/NMI framework under two-stage Otsu (NMI =0.379<τNMI=0.45 ).
Dataset
Backbone
ZS
SFT
IVL
Pets-37
Qwen2.5-VL
80.9
87.2
79.8
ImageNet-A
Qwen2.5-VL
79.4
82.0
78.9
Appendix
Table 7 : Near-OOD control cases (1-shot). When the frozen encoder preserves the target semantics, SFT improves over zero-shot and IVL provides no structural advantage, consistent with Remark 1 of the main paper. ImageNet-A uses a random 20-class subset.
τNMI
Datasets unchanged
0.30
22/22
0.36
22/22
0.42 (Otsu)
22/22
0.48
22/22
0.54
22/22
0.60
22/22
Appendix
Table 8 : Threshold robustness (Qwen2.5-VL). With τMMD fixed at its Otsu value 0.49 , the NMI threshold τNMI is swept across a band bracketing its Otsu value 0.42 . All 22 datasets retain identical regime assignments, confirming a natural separation in the data.
Figure 6 : Two-stage OOD taxonomy under Qwen2.5-VL with ImageNet-1K ( left ) and ReLAION-400M ( right ) as the MK-MMD reference. Because NMI never involves the reference, each dataset keeps its vertical position across panels and moves only horizontally. No dataset crosses a regime boundary. Only τMMD shifts ( 0.49→0.45 ); τNMI=0.42 is identical under both references.
Dataset
Category
MK-MMD
#Cls
Traits/Image
Avg. Len
MeanNormIDF
Oxford Pets
Near-OOD
0.174
37
12.5±3.5
2.2
0.890
Stanford Cars
Near-OOD
0.289
196
16.7±4.9
2.2
0.897
Flowers-102
Near-OOD
0.310
102
14.6±3.7
2.3
0.919
FGVC Aircraft
Near-OOD
0.409
100
15.6±4.2
2.3
0.904
Pokémon
Distant-OOD
0.411
18
14.4±4.3
2.2
0.959
Retinal OCT
Distant-OOD
0.578
8
18.8±5.4
2.5
0.810
Appendix
Table 9 : Trait extraction statistics across the OOD spectrum (core datasets; MVTec is reported per-category in Table 10 ). Traits/Image and MeanNormIDF show no systematic decline with MK-MMD, confirming that descriptive volume and class-informativeness are preserved even when discriminative capacity collapses.
MVTec Category
MK-MMD
#Cls
Traits/Image
Avg. Len
MeanNormIDF
Grid
0.566
6
15.3±3.5
2.1
0.972
Wood
0.587
6
15.8±3.6
2.4
0.912
Metal Nut
0.603
5
17.2±2.9
2.2
0.910
Hazelnut
0.610
5
14.2±2.9
2.1
0.927
Tile
0.615
6
15.0±3.1
2.2
0.932
Toothbrush
0.620
2
15.0±2.0
2.3
0.889
Appendix
Table 10 : Per-category trait extraction statistics on MVTec AD (15 independent 1-shot datasets), ordered by MK-MMD. MeanNormIDF remains consistently high (0.874–0.972) across all categories despite large domain gap (MK-MMD ≥0.566 ).
Dataset/Category
C
Random
Corr. → trait
Incorr. → wrong
× chance
Pokémon
Pokémon (881 images)
18
5.6%
81.2%
72.3%
14.6 ×
MVTec AD (per-category)
Transistor
5
20.0%
90.9%
86.3%
4.5 ×
Tile
6
16.7%
76.5%
71.1%
4.6 ×
Wood
6
16.7%
76.5%
68.2%
4.6 ×
Appendix
Table 11 : Unified trait-class alignment results (Qwen2.5-VL, 1-shot). The random baseline 1/C is the expected correct → trait-driven rate under class-independent trait citation. MVTec numbers are averaged over 3 support-set runs.
Type
Top-3 canonical traits
Fire
orange body, red fur, flame-tipped tail
Steel
metallic limbs, metallic surface, gray body
Ice
ice crystal structure, blue accents, ice cube head
Bug
antennae-like appendages, spiky appearance, black legs
Appendix
Table 12 : Representative per-class top-3 trait signatures on the Pokémon benchmark.
Class
Top-3 canonical traits
crack
darker brown crack, visible separation, cracked shell
Table 13 : Per-class top-3 trait signatures on the MVTec-Hazelnut benchmark (semantic-mode IVL, 5-shot). The defect-name token appears verbatim in the top-3 traits for every anomaly class.
Figure 7 : Per-trait grounding on a normal MVTec Hazelnut sample. Left to right: original image; IVL query attention for “clean surface” (none trait); IVL query attention for “brown crack” (crack trait); SFT attention for none; SFT attention for crack.
Configuration
Accuracy (%)
Δ Accuracy (%)
Hybrid (Default)
38.9
—
Semantic
32.6
−6.2
Low-Level
27.8
−11.1
Appendix
Table 14 : Ablation: Trait Prompt Variants
Configuration
Accuracy (%)
Δ Accuracy (%)
Two-Stage (Default)
38.9
—
Global Only
39.6
+0.7
Local Only
37.5
−1.4
No Filtering
36.8
−2.1
Strict Filtering ( k2=50 )
33.3
−5.6
Appendix
Table 15 : Ablation: Filtering Strategy
Configuration
Accuracy (%)
Δ Accuracy (%)
HDBSCAN (Default)
38.9
—
K-means
41.0
+2.1
Similarity-based
32.6
−6.2
Appendix
Table 16 : Ablation: Clustering Method
Configuration
Dimension
Accuracy (%)
Δ Accuracy (%)
Sentence Transformer (Default)
768
38.9
—
Qwen2.5-VL Encoder
3584
37.5
−1.4
Appendix
Table 17 : Ablation: Clustering Embedding Method
Configuration
Salience
Accuracy (%)
Δ Accuracy (%)
Salience Filtering (Default)
On
38.9
—
w/o Salience Filtering
Off
36.3
−2.6
Appendix
Table 18 : Ablation: Salience Filtering
Configuration
Visual Grounding
Accuracy (%)
Δ Accuracy (%)
With Grounding (Default)
On
38.9
—
w/o Grounding
Off
34.7
−4.2
Appendix
Table 19 : Ablation: Localized Visual Grounding
Figure 8 : Human and Model Performance on Normal vs. Special Case Pokémon
Figure 9 : Confusion matrix of IVL on the MVTec AD Transistor dataset. Nearly all defect types, including none , cut_lead , damaged_case , and misplaced , are misclassified as bent_lead . This collapse reflects a systematic failure mode. IVL interprets the presence of curved metal leads as definitive evidence for the bent_lead defect, even though normal samples and several other defect types naturally exhibit curved leads.
Figure 10 : Example images from two FractalDB classes. The visual patterns are generated by different IFS parameter families but are difficult for humans to distinguish, limiting trait-based reasoning.
Figure 11 : Representative samples of all five defect types in the MVTec AD Transistor category: none , bent_lead , cut_lead , damaged_case , and misplaced .
Figure 12 : Reasoning breakdown on a normal MVTec AD sample ( none ). IVL flags the object as bent_lead based on Visual Observation 1 ("The metal leads appear to be bent"). While the leads are indeed curved, they lack the unnatural distortion of the actual defect (refer to Fig. 11 for a comparison). The failure stems from the lack of a comparative mechanism: IVL detects the "bent" primitive but, without comparing against a reference image, cannot discern that this curvature is within standard tolerance.
IVL
SFT
V-RFT
ICL-1
ICL-8
Offline (one-time adaptation)
Wall-clock (s)
603
309
412
—
—
# GPUs
1
6
6
—
—
GPU-seconds
603
1,854
2,472
0
0
Peak VRAM (GB)
32
199
199
—
—
Online (per-image inference)
Appendix
Table 20 : Unified efficiency comparison across adaptation methods. Offline cost is a one-time expense amortized over all subsequent queries; online cost is per test image. IVL achieves the lowest total GPU compute for adaptation while requiring only a single consumer-grade GPU.
C
#Test
Offline s/class
Ratio
Online s/image
Ratio
Inference s/image
Ratio
20
100
68.60
1.000
16.85
1.000
9.056
1.000
40
200
63.65
0.928
15.80
0.938
8.902
0.983
60
300
61.23
0.893
15.45
0.917
8.711
0.962
80
400
58.84
0.858
15.34
0.910
8.404
0.928
100
500
61.60
0.898
15.17
0.900
9.052
1.000
Appendix
Table 21 : Scalability of IVL with number of classes C on FGVC Aircraft (Qwen2.5-VL-7B, 1-shot). “Offline s/class” is the per-class adaptation wall-clock time (stages S1–S4). “Online s/image” is total per-image inference time (S5 + retrieval). “Inference s/image” is the VLM generation time within S5 only. Ratio columns are relative to the C=20 baseline.
Standalone
MVTec AD (per-category)
Method
Pokémon
Ret. OCT
WM811k
Bottle
Cable
Capsule
Carpet
Grid
Hazelnut
Leather
Metal Nut
Pill
Screw
Tile
T.brush
Trans.
Wood
Zipper
MVT. Avg
CLIP ViT-B/32
CLIP Vanilla
42.54
12.29
13.49
49.4
9.9
16.7
16.2
11.1
15.2
30.5
21.8
12.6
13.6
31.5
72.5
12.6
13.7
11.2
22.6
CuPL [ 30 ]
36.11
12.68
9.64
34.2
6.4
20.6
16.2
16.7
35.2
22.9
34.6
5.0
14.9
58.6
72.5
9.5
35.6
11.2
26.3
Tip-Adapter [ 45 ]
36.11
32.93
36.75
46.8
34.0
26.4
43.2
44.4
48.6
36.4
41.8
27.7
25.3
45.9
70.0
12.6
57.5
12.6
38.1
CoOp [ 49 ]
27.78
32.3
23.1
48.1
14.2
26.2
34.2
13.9
40.9
59.3
19.1
12.0
16.2
81.1
65.0
11.6
65.8
9.8
34.5
Appendix
Table 22 : Non-generative vision-language baselines on distant-OOD benchmarks (1-shot for Tip-Adpater and CoOp, 0-shot for other methods).
Hyperparameter
Value
Description
Trait Filtering ( k1 , k2 )
Coarse ratio
0.8
Fraction of traits retained at global stage
k1 (coarse min)
50
Minimum number of globally retained traits
k1 (coarse max)
1000
Maximum number of globally retained traits
k2 (fine filter)
250
Number of traits retained after local refinement
HDBSCAN Clustering
Appendix
Table 23 : IVL pipeline default hyperparameters.
Figure 13 : Extraction Prompt.
Figure 14 : VLM Salience Prompt.
Figure 15 : VLM Classification Prompt.
Figure 16 : ICL Prompt. For each dataset, we use a simple format string to make the model aware of which dataset it is inferring on.
Figure 17 : More qualitative results on the MVTec AD dataset.
Figure 18 : More qualitative results on the MVTec AD dataset.
Figure 19 : More qualitative results on the Retinal OCT, WM-811K dataset.
Medical Vision-Language Models (VLMs) exhibit strong zero-shot performance, yet their effectiveness still declines on out-of-distribution (OOD) data due to domain shifts and class bias inherited from large-scale pretraining. Existing few-shot adaptation methods typically introduce additional trainable components, which can be unstable in extremely low-data regimes (e.g., 1-shot), and lack robustness on different medical data. We present TCLA, a purely training-free few-shot adaptation method for Medical VLMs, which is fast and model-agnostic. TCLA corrects inference logits based on a small set of support samples, boosting pretrained VLMs performance by improving inter-class deconfusion and reducing domain shift. Extensive experiments on nine datasets across multiple medical imaging modalities including X-ray, Ultrasound, MRI, CT, Histopathology, demonstrate that TCLA consistently improves OOD performance of Medical VLMs and, in most of cases, outperforms existing training-based adaptation methods.
Tianyou Jiang, Ziyu Zhou
University of Bern, Bern, Switzerland · Shanghai Jiao Tong University, Shanghai, China
Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel domains that exhibit substantial distribution shifts from their pretraining data. Existing domain adaptation methods rely on finetuning standard VLM components; however, depending on which components are updated, these approaches either limit the model's ability to learn domain-specific representations or cause catastrophic forgetting of previously acquired capabilities. We introduce Vision Contextualized Probing (VisCoP), a parameter-efficient adaptation framework that augments the VLM vision encoder with a compact set of learnable visual probes. By learning domain-specific visual representations through these probes while requiring only minimal updates to pretrained model components, VisCoP effectively adapts to new domains without sacrificing existing knowledge. We evaluate VisCoP across three challenging adaptation settings: cross-view (exocentric to egocentric), cross-modal (RGB to depth), and cross-task (human understanding to robot control). Across all scenarios, VisCoP consistently outperforms existing domain adaptation strategies, achieving superior target-domain performance while preserving the pretrained VLM's capabilities on the source domain. These results demonstrate that lightweight visual probing provides an effective and robust solution for adapting VLMs under substantial distribution shifts. Code, models, and evaluation protocols are available at https://github.com/dominickrei/VisCoP.
Dominick Reilly, Manish Kumar Govind, Le Xue +1
University of North Carolina at Charlotte · Elorian AI
We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting, where a small number of continuous prompt tokens are optimized while the pretrained backbone remains frozen. We identify two key design choices. First, placing prompt tokens at the cross-modal boundary between visual and text tokens outperforms other placements (10.0 vs. 8.4 mAP). Second, initializing prompts from the empty space token outperforms semantic and random initialization. With these choices, one to three learned tokens (7,168 parameters on average) match the best LoRA configuration on Roboflow20-VL (14.2 mAP, 10-shot) while training over 20,000x fewer parameters. Soft prompting remains harder to optimize, exhibiting higher variance across random seeds. Unlike LoRA, however, it causes no forgetting: the LoRA rank matching our accuracy reduces NaturalBench VQA accuracy by 35% relative, rising to 56% at the largest rank, whereas soft prompting leaves pretrained performance unchanged. The learned tokens behave like prompts rather than weights. They transfer to a newer model without retraining (+0.8 mAP on Qwen3.5-9B) and can be verbalized into readable prompts competitive with prompt-search methods (matching DetPO and outperforming GEPA). The approach also extends beyond detection. On RoboCasa manipulation tasks, the frozen π0.5 vision-language-action policy benefits from soft prompting, matching the LoRA baseline on two of three tasks when tokens are placed at the gradient bottleneck. These results suggest modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask.