Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs
Organizations: National Tsing Hua University, Hsinchu, Taiwan · National Taiwan University, Taipei, Taiwan
Abstract
Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining distribution, yet fail on specialized domains where the required discriminative features were never learned, a regime we term distant out-of-distribution (OOD). Standard adaptation methods cannot overcome this representational absence because they operate within the encoder's existing feature space. However, VLMs retain a robust descriptive capacity even when discrimination collapses: a model that cannot classify a medical scan can still articulate its visual patterns. Exploiting this asymmetry, we introduce Inductive Visual Logic (IVL), a training-free framework that constructs classification knowledge from the model's surviving descriptive ability. IVL extracts visual traits from few-shot support images through dual-mode prompting, combining semantic descriptions with primitive visual observations, and organizes them into per-class trait dictionaries. At inference, hierarchical filtering identifies spatially grounded trait evidence for classification. Across multiple distant-OOD benchmarks, IVL achieves the highest aggregate accuracy under two VLM backbones while producing interpretable, trait-traceable predictions.
Figures & tables
| Dataset | Zero-Shot | MajorVote | ICL | CuPL [ 30 ] | Menon et al . [ 23 ] | V-RFT [ 18 ] | SFT+LoRA | SFT | Tip-Adapter [ 45 ] | CoOp [ 49 ] | IVL (Ours) |
| Qwen2.5-VL-7B | |||||||||||
| Pokémon | 47.47 | 48.23 | 22.00 | 15.28 | 25.77 | 48.77 | 45.59 | 45.52 | 9.72 | 20.01 | 50.30 |
| Retinal OCT | 14.36 | 16.39 | 13.75 | 12.50 | 11.46 | 14.00 | 13.69 | 14.77 | 12.50 | 13.39 | 20.83 |
| WM811k | 13.13 | 13.73 | 12.41 | 12.77 | 12.29 | 9.52 | 9.16 | 12.41 | 11.08 | 9.30 | 16.51 |
| 8-shot | |||||||||||
| Pokémon | – | – | 34.24 | – | – | 49.09 | 50.09 | 49.94 | 6.94 | 29.17 | 52.20 |
| Bottle | Cable | Capsule | Carpet | Grid | Hazelnut | Leather | Metal Nut | Pill | Screw | Tile | Toothbrush | Transistor | Wood | Zipper | Mean | |
| Semantic | 39.2 | 14.7 | 35.7 | 36.9 | 34.3 | 52.7 | 28.0 | 26.7 | 5.9 | 32.3 | 57.7 | 51.7 | 15.4 | 40.2 | 17.0 | 32.6 |
| Low-level | 40.5 | 30.3 | 35.7 | 50.8 | 36.6 | 46.7 | 39.5 | 31.2 | 10.5 | 30.3 | 52.6 | 72.5 | 14.4 | 47.9 | 29.3 | 37.9 |
| Dual | 46.8 | 45.4 | 44.4 | 60.4 | 43.1 | 72.4 | 48.3 | 40.9 | 22.0 | 33.1 | 65.8 | 65.0 | 11.6 | 65.8 | 26.6 | 46.1 |
| +6.3 | +15.1 | +8.7 | +9.6 | +6.5 | +19.7 | +8.8 | +9.7 | +11.5 | +0.8 | +8.1 | +17.9 | +7.2 |
| Component | Configuration | Acc. (%) | |
| Extraction ( Sec. 3.5 ) | Semantic only | 32.6 | |
| Low-level only | 27.8 | ||
| Filtering ( Sec. 3.6 ) | Global only | 39.6 | |
| Local only | 37.5 | ||
| No filtering | 36.8 | ||
| Quality ( Sec. 3.5 ) | w/o salience | 36.3 |
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Symbol | Description |
| Problem Setting | Generative VLM (composed of and ) | |
| Vision encoder | ||
| Language decoder | ||
| Pretraining distribution | ||
| Target distribution | ||
| Few-shot support set |
| Dataset | MK-MMD | NMI | 0-shot (%) | (%) | Regime |
| Reference | |||||
| ImageNet | — | 0.881 | — | — | (reference) |
| Near-OOD (low MK-MMD, high NMI) | |||||
| Pets-37 | 0.174 | 0.651 | 80.9 | +6.3 | Near-OOD |
| Cars-196 | 0.289 | 0.750 | 56.2 | +27.3 | Near-OOD |
| Flowers-102 | 0.310 | 0.975 | 72.4 | +23.9 | Near-OOD |
| Zero-Shot | SFT | CoOp | IVL (Ours) | |
| 1-shot | 18.26 | 28.14 | 2.93 | 31.34 |
| 8-shot | — | 29.13 | 5.63 | 33.18 |
| Dataset | Backbone | ZS | SFT | IVL |
| Pets-37 | Qwen2.5-VL | 80.9 | 87.2 | 79.8 |
| ImageNet-A | Qwen2.5-VL | 79.4 | 82.0 | 78.9 |
| Datasets unchanged | |
| 0.30 | 22/22 |
| 0.36 | 22/22 |
| 0.42 (Otsu) | 22/22 |
| 0.48 | 22/22 |
| 0.54 | 22/22 |
| 0.60 | 22/22 |
| Dataset | Category | MK-MMD | #Cls | Traits/Image | Avg. Len | MeanNormIDF |
| Oxford Pets | Near-OOD | 0.174 | 37 | 2.2 | 0.890 | |
| Stanford Cars | Near-OOD | 0.289 | 196 | 2.2 | 0.897 | |
| Flowers-102 | Near-OOD | 0.310 | 102 | 2.3 | 0.919 | |
| FGVC Aircraft | Near-OOD | 0.409 | 100 | 2.3 | 0.904 | |
| Pokémon | Distant-OOD | 0.411 | 18 | 2.2 | 0.959 | |
| Retinal OCT | Distant-OOD | 0.578 | 8 | 2.5 | 0.810 |
| MVTec Category | MK-MMD | #Cls | Traits/Image | Avg. Len | MeanNormIDF |
| Grid | 0.566 | 6 | 2.1 | 0.972 | |
| Wood | 0.587 | 6 | 2.4 | 0.912 | |
| Metal Nut | 0.603 | 5 | 2.2 | 0.910 | |
| Hazelnut | 0.610 | 5 | 2.1 | 0.927 | |
| Tile | 0.615 | 6 | 2.2 | 0.932 | |
| Toothbrush | 0.620 | 2 | 2.3 | 0.889 |
| Dataset/Category | Random | Corr. trait | Incorr. wrong | chance | |
| Pokémon | |||||
| Pokémon (881 images) | 18 | 5.6% | 81.2% | 72.3% | 14.6 |
| MVTec AD (per-category) | |||||
| Transistor | 5 | 20.0% | 90.9% | 86.3% | 4.5 |
| Tile | 6 | 16.7% | 76.5% | 71.1% | 4.6 |
| Wood | 6 | 16.7% | 76.5% | 68.2% | 4.6 |
| Type | Top-3 canonical traits |
| Fire | orange body, red fur, flame-tipped tail |
| Steel | metallic limbs, metallic surface, gray body |
| Ice | ice crystal structure, blue accents, ice cube head |
| Bug | antennae-like appendages, spiky appearance, black legs |
| Class | Top-3 canonical traits |
| crack | darker brown crack, visible separation, cracked shell |
| cut | sharp cut edges, irregular cut edge, visible separation lines |
| hole | smooth contrasting texture, nut hole detail, darker brown hole |
| text-like white spots, white shell contrast, white spot circle | |
| none | clean surface, smooth ridged surface, consistent shell structure |
| Configuration | Accuracy (%) | Accuracy (%) |
| Hybrid (Default) | 38.9 | — |
| Semantic | 32.6 | |
| Low-Level | 27.8 |
| Configuration | Accuracy (%) | Accuracy (%) |
| Two-Stage (Default) | 38.9 | — |
| Global Only | 39.6 | |
| Local Only | 37.5 | |
| No Filtering | 36.8 | |
| Strict Filtering ( ) | 33.3 |
| Configuration | Accuracy (%) | Accuracy (%) |
| HDBSCAN (Default) | 38.9 | — |
| K-means | 41.0 | |
| Similarity-based | 32.6 |
| Configuration | Dimension | Accuracy (%) | Accuracy (%) |
| Sentence Transformer (Default) | 768 | 38.9 | — |
| Qwen2.5-VL Encoder | 3584 | 37.5 |
| Configuration | Salience | Accuracy (%) | Accuracy (%) |
| Salience Filtering (Default) | On | 38.9 | — |
| w/o Salience Filtering | Off | 36.3 |
| Configuration | Visual Grounding | Accuracy (%) | Accuracy (%) |
| With Grounding (Default) | On | 38.9 | — |
| w/o Grounding | Off | 34.7 |
| IVL | SFT | V-RFT | ICL-1 | ICL-8 | |
| Offline (one-time adaptation) | |||||
| Wall-clock (s) | 603 | 309 | 412 | — | — |
| # GPUs | 1 | 6 | 6 | — | — |
| GPU-seconds | 603 | 1,854 | 2,472 | 0 | 0 |
| Peak VRAM (GB) | 32 | 199 | 199 | — | — |
| Online (per-image inference) | |||||
| #Test | Offline s/class | Ratio | Online s/image | Ratio | Inference s/image | Ratio | |
| 20 | 100 | 68.60 | 1.000 | 16.85 | 1.000 | 9.056 | 1.000 |
| 40 | 200 | 63.65 | 0.928 | 15.80 | 0.938 | 8.902 | 0.983 |
| 60 | 300 | 61.23 | 0.893 | 15.45 | 0.917 | 8.711 | 0.962 |
| 80 | 400 | 58.84 | 0.858 | 15.34 | 0.910 | 8.404 | 0.928 |
| 100 | 500 | 61.60 | 0.898 | 15.17 | 0.900 | 9.052 | 1.000 |
| Standalone | MVTec AD (per-category) | ||||||||||||||||||
| Method | Pokémon | Ret. OCT | WM811k | Bottle | Cable | Capsule | Carpet | Grid | Hazelnut | Leather | Metal Nut | Pill | Screw | Tile | T.brush | Trans. | Wood | Zipper | MVT. Avg |
| CLIP ViT-B/32 | |||||||||||||||||||
| CLIP Vanilla | 42.54 | 12.29 | 13.49 | 49.4 | 9.9 | 16.7 | 16.2 | 11.1 | 15.2 | 30.5 | 21.8 | 12.6 | 13.6 | 31.5 | 72.5 | 12.6 | 13.7 | 11.2 | 22.6 |
| CuPL [ 30 ] | 36.11 | 12.68 | 9.64 | 34.2 | 6.4 | 20.6 | 16.2 | 16.7 | 35.2 | 22.9 | 34.6 | 5.0 | 14.9 | 58.6 | 72.5 | 9.5 | 35.6 | 11.2 | 26.3 |
| Tip-Adapter [ 45 ] | 36.11 | 32.93 | 36.75 | 46.8 | 34.0 | 26.4 | 43.2 | 44.4 | 48.6 | 36.4 | 41.8 | 27.7 | 25.3 | 45.9 | 70.0 | 12.6 | 57.5 | 12.6 | 38.1 |
| CoOp [ 49 ] | 27.78 | 32.3 | 23.1 | 48.1 | 14.2 | 26.2 | 34.2 | 13.9 | 40.9 | 59.3 | 19.1 | 12.0 | 16.2 | 81.1 | 65.0 | 11.6 | 65.8 | 9.8 | 34.5 |
| Hyperparameter | Value | Description |
| Trait Filtering ( , ) | ||
| Coarse ratio | 0.8 | Fraction of traits retained at global stage |
| (coarse min) | 50 | Minimum number of globally retained traits |
| (coarse max) | 1000 | Maximum number of globally retained traits |
| (fine filter) | 250 | Number of traits retained after local refinement |
| HDBSCAN Clustering | ||