Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.
Figures & tables
Figure 1: Verbal comparison versus visual-difference conditioning. (a,c) Two reasoning routes using the same query–reference pair. (b) Example predictions and post-hoc attention maps, shown for interpretation only and not used as model inputs.
Figure 2: VD-DeepStack overview. Blue paths provide appearance context to both images; orange paths inject difference vectors E only into query tokens, weighted by spatial scores S . Injection occurs during prefill without adding tokens.
Model
MVTec AD
VisA
MVTec 3D
MPDD
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Commercial MLLMs
Gemini-2.5-pro
79.4/34.4
98.4 /46.0
79.0/48.8
65.7/10.5
97.2 /17.1
62.7/21.3
75.7/10.9
90.1 /18.3
83.2/19.9
62.1/16.3
92.2/23.2
60.8/26.7
GPT-4o
68.9/8.1
74.1/12.9
82.4/16.9
57.6/3.7
69.4/5.8
61.1/9.2
67.8/7.0
74.0/11.5
83.4/14.8
63.1/13.6
84.3/20.4
64.3/22.3
GPT-5
78.7/41.8
95.9/64.2
79.4/53.0
65.8/18.5
94.5/31.0
63.3/30.7
75.0/27.2
85.3/41.1
84.1/43.2
68.4/21.2
91.9/34.5
67.0/29.2
Qwen3.8-Flash
76.0/58.0
95.1/71.5
72.2/73.4
78.9/32.0
86.3/43.2
76.2/53.0
78.3 /29.9
82.2/40.1
90.4/46.7
63.0/37.9
79.9/49.7
62.2/52.5
Table 1: Detection/localization (%) on four benchmarks. Bold and underline mark the best and second-best values, respectively, for detection and localization separately, including VD-DeepStack and excluding full-shot reference methods. Ties share the same rank. VD-DeepStack results include RL refinement. IAD-R1 ∗ is our reproduction trained and evaluated using MMR-AD data. Dashes denote unavailable results.
Variant
VisA
MVTec 3D
Acc.
Rec.
Prec.
Acc.
Rec.
Prec.
Baseline
78.0/38.3
66.5/42.0
91.5/81.6
70.4/38.7
67.7/45.4
93.0/72.5
w/o Evi
78.9/43.4
69.6/46.9
90.4/84.0
67.7/43.2
64.7/47.2
97.3 /75.1
E-tokens
78.3/41.0
67.3/44.7
91.7/82.2
71.9/44.8
69.0/48.6
94.9/75.7
S-only
79.7/44.3
69.3/47.5
92.8 / 86.3
72.4/47.0
68.7/50.6
95.9/78.6
w/o Ctx
81.8 /45.2
76.3/49.6
91.1/81.7
74.2/48.6
70.7/52.6
96.2/78.9
Table 2: Ablation of visual-difference conditioning. Entries report detection/localization (%) at localization IoU ≥ 0.1. Ctx and Evi denote the visual-context and difference-evidence paths. All variants except VD-DeepStack-7B are trained without RL. Bold marks the best result among the variants without RL.
Model
BrainMRI
HeadCT
Average
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Qwen3.8-Flash
71.4 ± 4.7
88.1 ± 2.6
71.9 ± 4.2
50.5 ± 2.2
86.4 ± 2.1
50.3 ± 1.3
60.9
87.3
61.1
Qwen3.8-27B
50.4 ± 7.3
39.2 ± 10.5
65.8 ± 9.1
53.6 ± 2.2
13.8 ± 4.7
68.1 ± 8.6
52.0
26.5
67.0
IAD-R1 ∗
78.7 ± 5.6
92.4 ± 1.0
77.6 ± 5.4
73.5 ± 4.0
97.0 ± 2.5
66.2 ± 3.6
76.1
94.7
71.9
VD-DeepStack-7B
83.9 ± 2.4
94.3 ± 0.8
82.2 ± 2.9
77.1 ± 2.0
95.8 ± 1.6
69.8 ± 2.1
80.5
95.1
76.0
VD-DeepStack-8B
87.7 ± 1.2
98.5 ± 0.7
84.2 ± 1.6
77.0 ± 3.8
99.8 ± 0.4
68.8 ± 3.7
82.3
99.1
76.5
Table 3: Cross-domain anomaly detection performance on medical datasets (%). Results are means over five 1-shot runs with different normal reference images; standard deviations are shown where available. Average denotes the unweighted mean across the two datasets.
Model
MVTec AD
MVTec 3D
VisA
MPDD
Overall
Abn.
Norm.
Abn.
Norm.
Abn.
Norm.
Abn.
Norm.
IAD-R1 ∗
69.9
75.9
51.2
81.9
49.8
90.1
58.2
73.9
68.3
Qwen3.8-Flash
74.7
57.5
58.1
61.4
62.4
72.2
66.5
57.7
65.6
VD-DeepStack-7B
68.1
86.1
55.6
86.6
56.2
87.8
60.0
74.4
69.4
Table 4: Quality of generated CoT reasoning scored by GLM-5.3-Flash (0–100; higher is better). The judge compares CoT text with ground-truth annotations, excluding final answers. Abn./Norm. denote abnormal/normal samples. Overall is the mean score over all evaluated samples. Bold indicates the best score in each column.
Figure 3: Qualitative comparison with IAD-R1 ∗ . Attention is extracted from the token position immediately preceding the final answer, after reasoning, to query-image tokens and averaged over the last four LLM layers. Attention differences are computed as VD-DeepStack minus IAD-R1 ∗ ; red and blue indicate increased and decreased attention, respectively.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Level
Native visual block
DINO blocks
Decoder block
1
13
10–13
2
2
19
14–17
5
3
25
18–20
9
4
31
21–23
13
Appendix
Table 5: Feature extraction and decoder injection configuration for VD-DeepStack-7B. All block indices are zero-based.
Placement
Blocks
Det. Acc.
Loc. Acc.
Wide span
5/11/16/22
78.54
43.74
First half †
2/5/9/13
79.65
44.47
First quarter
0/2/4/6
79.30
44.71
Appendix
Table 6: Sensitivity to decoder injection layers (zero-based indices).
Figure 4: Frozen visual encoders under a common one-shot matching protocol. Bubble area indicates vision-backbone parameter count; Qwen labels identify the parent vision–language models.
Model
MVTec AD
VisA
MVTec 3D
MPDD
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Qwen3.8-Flash
18.9
29.4
30.4
6.4
11.0
13.0
7.9
13.0
14.4
10.9
17.5
18.4
Qwen3-VL-32B
29.6
45.7
41.9
11.0
16.4
22.8
16.1
25.4
28.7
26.5
38.1
32.2
Qwen3.8-27B
41.7
53.7
57.3
13.7
21.6
26.6
18.7
28.2
32.4
25.8
36.7
32.0
IAD-R1*
34.9
44.3
55.2
24.2
29.3
53.5
26.0
34.6
52.0
23.3
27.5
40.6
VD-DeepStack-7B
42.6
53.4
61.7
26.1
32.7
52.0
35.0
42.3
60.9
30.2
36.6
48.3
Appendix
Table 7: Localization results (%) at IoU ≥ 0.5.
Sample
Score
Criterion
Abnormal ( y=1 )
90–100
Correct anomaly judgment, with defect type and location consistent with GT.
75–89
Correct type with a vague or slightly inaccurate location, or correct location with a synonymous or near-equivalent type.
55–74
Correct broad region but incorrect type, or only one of multiple GT defects identified.
30–54
Anomaly stated, but both the described type and location differ from GT.
10–29
Normal judgment or missed anomaly, or irrelevant content.
0–9
Empty CoT or wholly fabricated content.
Appendix
Table 8: Scoring anchors for generated CoTs. GT denotes the ground-truth annotation.
Figure 5: An anomalous black bracket from MPDD, showing GT annotations and the two models’ responses and predicted boxes. Red keywords disagree with GT; blue keywords convey similar or matching content.
Figure 6: A normal candle sample from VisA. IAD-R1 ∗ reports a broken part, while VD-DeepStack-7B predicts no anomaly. Red keywords disagree with GT; blue keywords convey similar or matching content.
Figure 7: An abnormal BrainMRI query paired with a normal reference. Blue highlights mark the response identifying an anomaly, and ochre highlights mark the incorrect normal judgment. GT supports the image-level decision only; the highlighted regional descriptions and coordinates are unverified model outputs.
As a classic vision task, anomaly detection has been widely applied in industrial inspection and medical imaging. In this task, data scarcity is often a frequently-faced issue. To solve it, the few-shot anomaly detection (FSAD) scheme is attracting increasing attention. In recent years, beyond traditional visual paradigm, Vision-Language Model (VLM) has been extensively explored to boost this field. However, in currently-existing VLM-based FSAD schemes, almost all perform anomaly inference only by pairwise feature matching, ignoring structural dependencies and global consistency. To further redound to FSAD via VLM, we propose a Heterogeneous Hypergraph Vision-Language Reasoning (H2VLR) framework. It reformulates the FSAD as a high-order inference problem of visual-semantic relations, by jointly modeling visual regions and semantic concepts in a unified hypergraph. Experimental comparisons verify the effectiveness and advantages of H2VLR. It could often achieve state-of-the-art (SOTA) performance on representative industrial and medical benchmarks. Our code will be released upon acceptance.
Jianghong Huang, Luping Ji, Weiwei Duan +1
School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, China
Zero-shot anomaly detection aims to identify defects in unseen categories without target-specific training. Existing methods usually apply the same feature transformation to all samples, treating normal and anomalous data uniformly despite their fundamentally asymmetric distributions, compact normals versus diverse anomalies. We instead exploit this natural asymmetry by proposing AVA-DINO, an anomaly-aware vision-language adaptation framework with dual specialized branches for normal and anomalous patterns that adapt frozen DINOv3 visual features. During training on auxiliary data, the two branches are learned jointly with a text-guided routing mechanism and explicit routing regularization that encourages branch specialization. At test time, only the input image and fixed, predefined language descriptions are used to dynamically combine the two branches, enabling an asymmetric activation. This design prevents degenerate uniform routing and allows context-specific feature transformations. Experiments across nine industrial and medical benchmarks demonstrate state-of-the-art performance, achieving 93.5% image-AUROC on MVTec-AD and strong cross-domain generalization to medical imaging without domain-specific fine-tuning. https://github.com/aqeeelmirza/AVA-DINO
Muhammad Aqeel, Maham Nazir, Uzair Khan +2
Dept. of Engineering for Innovation Medicine, University of Verona, Italy · School of Computer Science and Engineering, Beihang University, China · Dept. of Engineering for Innovation +1
Few-shot anomaly detection (FSAD) has made significant strides, yet existing methods still face critical challenges: (i) dependence on task- or dataset-specific training/fine-tuning, (ii) reliance on language supervision or carefully hand-crafted prompts, and (iii) limited robustness across domains. In this paper, we introduce HyperFSAD, a novel FSAD framework that is training-free, language-free, and robust across domains, offering a powerful solution to these challenges. Built upon DINOv3 and a hypergraph-based inference mechanism, our approach performs inference without any task-specific optimization or text prompts, while remaining competitive. Specifically, we replace sensitive nearest-neighbor / top-n matching with \textbf{Sparse Hyper Matching}: \textit{sparsemax} first selects the most relevant support patches, which are then aggregated into a \textit{hyperedge} as compact normal evidence to suppress background noise and distractors. We further introduce \textbf{Dual-Branch Image Scoring}, which fuses \emph{spatial anomaly evidence} from the patch-grid anomaly map with \emph{global semantic deviation} captured by support-aware CLS matching, yielding a robust image-level anomaly score in a strictly visual manner. Notably, all components of HyperFSAD are purely visual, eliminating the need for labor-intensive hand-crafted text prompts. Under the stringent training-free and language-free setting, HyperFSAD achieves state-of-the-art performance across six datasets spanning four industrial datasets (MVTecAD, VisA, MPDD, BTAD) and two medical datasets (RESC, BraTS).
Guohuan Xie, Xin He, Dingying Fan +2
Nankai University · Tianjin University of Technology · Tsinghua University