Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.
Figures & tables
Figure 1: Verbal comparison versus visual-difference conditioning. (a,c) Two reasoning routes using the same query–reference pair. (b) Example predictions and post-hoc attention maps, shown for interpretation only and not used as model inputs.
Figure 2: VD-DeepStack overview. Blue paths provide appearance context to both images; orange paths inject difference vectors E only into query tokens, weighted by spatial scores S . Injection occurs during prefill without adding tokens.
Model
MVTec AD
VisA
MVTec 3D
MPDD
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Commercial MLLMs
Gemini-2.5-pro
79.4/34.4
98.4 /46.0
79.0/48.8
65.7/10.5
97.2 /17.1
62.7/21.3
75.7/10.9
90.1 /18.3
83.2/19.9
62.1/16.3
92.2/23.2
60.8/26.7
GPT-4o
68.9/8.1
74.1/12.9
82.4/16.9
57.6/3.7
69.4/5.8
61.1/9.2
67.8/7.0
74.0/11.5
83.4/14.8
63.1/13.6
84.3/20.4
64.3/22.3
GPT-5
78.7/41.8
95.9/64.2
79.4/53.0
65.8/18.5
94.5/31.0
63.3/30.7
75.0/27.2
85.3/41.1
84.1/43.2
68.4/21.2
91.9/34.5
67.0/29.2
Qwen3.8-Flash
76.0/58.0
95.1/71.5
72.2/73.4
78.9/32.0
86.3/43.2
76.2/53.0
78.3 /29.9
82.2/40.1
90.4/46.7
63.0/37.9
79.9/49.7
62.2/52.5
Table 1: Detection/localization (%) on four benchmarks. Bold and underline mark the best and second-best values, respectively, for detection and localization separately, including VD-DeepStack and excluding full-shot reference methods. Ties share the same rank. VD-DeepStack results include RL refinement. IAD-R1 ∗ is our reproduction trained and evaluated using MMR-AD data. Dashes denote unavailable results.
Variant
VisA
MVTec 3D
Acc.
Rec.
Prec.
Acc.
Rec.
Prec.
Baseline
78.0/38.3
66.5/42.0
91.5/81.6
70.4/38.7
67.7/45.4
93.0/72.5
w/o Evi
78.9/43.4
69.6/46.9
90.4/84.0
67.7/43.2
64.7/47.2
97.3 /75.1
E-tokens
78.3/41.0
67.3/44.7
91.7/82.2
71.9/44.8
69.0/48.6
94.9/75.7
S-only
79.7/44.3
69.3/47.5
92.8 / 86.3
72.4/47.0
68.7/50.6
95.9/78.6
w/o Ctx
81.8 /45.2
76.3/49.6
91.1/81.7
74.2/48.6
70.7/52.6
96.2/78.9
Table 2: Ablation of visual-difference conditioning. Entries report detection/localization (%) at localization IoU ≥ 0.1. Ctx and Evi denote the visual-context and difference-evidence paths. All variants except VD-DeepStack-7B are trained without RL. Bold marks the best result among the variants without RL.
Model
BrainMRI
HeadCT
Average
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Qwen3.8-Flash
71.4 ± 4.7
88.1 ± 2.6
71.9 ± 4.2
50.5 ± 2.2
86.4 ± 2.1
50.3 ± 1.3
60.9
87.3
61.1
Qwen3.8-27B
50.4 ± 7.3
39.2 ± 10.5
65.8 ± 9.1
53.6 ± 2.2
13.8 ± 4.7
68.1 ± 8.6
52.0
26.5
67.0
IAD-R1 ∗
78.7 ± 5.6
92.4 ± 1.0
77.6 ± 5.4
73.5 ± 4.0
97.0 ± 2.5
66.2 ± 3.6
76.1
94.7
71.9
VD-DeepStack-7B
83.9 ± 2.4
94.3 ± 0.8
82.2 ± 2.9
77.1 ± 2.0
95.8 ± 1.6
69.8 ± 2.1
80.5
95.1
76.0
VD-DeepStack-8B
87.7 ± 1.2
98.5 ± 0.7
84.2 ± 1.6
77.0 ± 3.8
99.8 ± 0.4
68.8 ± 3.7
82.3
99.1
76.5
Table 3: Cross-domain anomaly detection performance on medical datasets (%). Results are means over five 1-shot runs with different normal reference images; standard deviations are shown where available. Average denotes the unweighted mean across the two datasets.
Model
MVTec AD
MVTec 3D
VisA
MPDD
Overall
Abn.
Norm.
Abn.
Norm.
Abn.
Norm.
Abn.
Norm.
IAD-R1 ∗
69.9
75.9
51.2
81.9
49.8
90.1
58.2
73.9
68.3
Qwen3.8-Flash
74.7
57.5
58.1
61.4
62.4
72.2
66.5
57.7
65.6
VD-DeepStack-7B
68.1
86.1
55.6
86.6
56.2
87.8
60.0
74.4
69.4
Table 4: Quality of generated CoT reasoning scored by GLM-5.3-Flash (0–100; higher is better). The judge compares CoT text with ground-truth annotations, excluding final answers. Abn./Norm. denote abnormal/normal samples. Overall is the mean score over all evaluated samples. Bold indicates the best score in each column.
Figure 3: Qualitative comparison with IAD-R1 ∗ . Attention is extracted from the token position immediately preceding the final answer, after reasoning, to query-image tokens and averaged over the last four LLM layers. Attention differences are computed as VD-DeepStack minus IAD-R1 ∗ ; red and blue indicate increased and decreased attention, respectively.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Level
Native visual block
DINO blocks
Decoder block
1
13
10–13
2
2
19
14–17
5
3
25
18–20
9
4
31
21–23
13
Appendix
Table 5: Feature extraction and decoder injection configuration for VD-DeepStack-7B. All block indices are zero-based.
Placement
Blocks
Det. Acc.
Loc. Acc.
Wide span
5/11/16/22
78.54
43.74
First half †
2/5/9/13
79.65
44.47
First quarter
0/2/4/6
79.30
44.71
Appendix
Table 6: Sensitivity to decoder injection layers (zero-based indices).
Figure 4: Frozen visual encoders under a common one-shot matching protocol. Bubble area indicates vision-backbone parameter count; Qwen labels identify the parent vision–language models.
Model
MVTec AD
VisA
MVTec 3D
MPDD
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Acc.
Recall
Prec.
Qwen3.8-Flash
18.9
29.4
30.4
6.4
11.0
13.0
7.9
13.0
14.4
10.9
17.5
18.4
Qwen3-VL-32B
29.6
45.7
41.9
11.0
16.4
22.8
16.1
25.4
28.7
26.5
38.1
32.2
Qwen3.8-27B
41.7
53.7
57.3
13.7
21.6
26.6
18.7
28.2
32.4
25.8
36.7
32.0
IAD-R1*
34.9
44.3
55.2
24.2
29.3
53.5
26.0
34.6
52.0
23.3
27.5
40.6
VD-DeepStack-7B
42.6
53.4
61.7
26.1
32.7
52.0
35.0
42.3
60.9
30.2
36.6
48.3
Appendix
Table 7: Localization results (%) at IoU ≥ 0.5.
Sample
Score
Criterion
Abnormal ( y=1 )
90–100
Correct anomaly judgment, with defect type and location consistent with GT.
75–89
Correct type with a vague or slightly inaccurate location, or correct location with a synonymous or near-equivalent type.
55–74
Correct broad region but incorrect type, or only one of multiple GT defects identified.
30–54
Anomaly stated, but both the described type and location differ from GT.
10–29
Normal judgment or missed anomaly, or irrelevant content.
0–9
Empty CoT or wholly fabricated content.
Appendix
Table 8: Scoring anchors for generated CoTs. GT denotes the ground-truth annotation.
Figure 5: An anomalous black bracket from MPDD, showing GT annotations and the two models’ responses and predicted boxes. Red keywords disagree with GT; blue keywords convey similar or matching content.
Figure 6: A normal candle sample from VisA. IAD-R1 ∗ reports a broken part, while VD-DeepStack-7B predicts no anomaly. Red keywords disagree with GT; blue keywords convey similar or matching content.
Figure 7: An abnormal BrainMRI query paired with a normal reference. Blue highlights mark the response identifying an anomaly, and ochre highlights mark the incorrect normal judgment. GT supports the image-level decision only; the highlighted regional descriptions and coordinates are unverified model outputs.
Dept. of Engineering for Innovation Medicine, University of Verona, Italy · School of Computer Science and Engineering, Beihang University, China · Dept. of Engineering for Innovation +1