cs.AISep 28, 2026

VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection

Authors: Mengyang Zhao, Zhuolin He, Haiyang Yu, Yuxuan Liang, Yifang Xu, Yuchuan Wu, Xiaolei Chen, Zhengtao Yao, +4 more

Organizations: Fudan University · University of Southern California · Fuzhou University · Tongji University

Abstract

Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. H2VLR: Heterogeneous Hypergraph Vision-Language Reasoning for Few-Shot Anomaly Detection

    Apr 16, 2026Jianghong Huang, Luping Ji, Weiwei Duan +1Semantic Region

  2. Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection

    May 12, 2026Muhammad Aqeel, Maham Nazir, Uzair Khan +2Vision-Language Model AdaptationDino

  3. Hypergraph-Enhanced Training-Free and Language-Free Few-Shot Anomaly Detection

    May 11, 2026Guohuan Xie, Xin He, Dingying Fan +2Multimodal Anomaly DetectionUnsupervised Detection