Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?
Organizations: Hunan University · University of Aberdeen · South China Normal University · South China University of Technology
Abstract
A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU). We show that in-domain training does not close this gap. On MMAD, a widely adopted MM-IAU benchmark, trained specialists reach at most 75.5% accuracy in defect localization, against 92.3% for human experts, and even detect anomalies less accurately than their untrained base model. Meanwhile, different MLLMs offer complementary strengths but share this weakness in fine-grained perception, so combining them alone cannot remove it. We therefore propose SiGMA, a spatially grounded multi-agent framework that divides labor between heterogeneous MLLM agents and a dedicated visual defect expert. A multimodal searcher supplies industrial knowledge and normal references, the defect expert turns query-reference comparison into calibrated anomaly evidence, and a label-free reliability controller weighs each source by task-wise competence and query-level evidence quality. SiGMA reaches 85.2% average accuracy on MMAD, 4.0% above the strongest trained specialist and Gemini-2.5-Pro and within 1.5% of human experts. Even with three agents of at most 9B parameters, it reaches 84.4%, and new MLLMs join without retraining.
Figures & tables
| Anomaly | Defect | Object | ||||||
| Method | Detection | Classification | Localization | Description | Analysis | Classification | Analysis | Average |
| Human (Expert) | 95.2 | 75.0 | 92.3 | 83.3 | 94.2 | 86.1 | 80.4 | 86.7 |
| Large-scale MLLMs | ||||||||
| InternVL2-76B ( Chen et al., 2024 ) | 68.3 | 54.2 | 56.7 | 66.3 | 80.5 | 86.4 | 82.9 | 70.8 |
| GPT-4o ( Hurst et al., 2024 ) | 68.6 | 65.8 | 55.6 | 73.2 | 83.4 | 95.0 | 82.8 | 74.9 |
| GPT-5-mini ( Singh et al., 2025 ) | 64.1 | 67.4 | 69.1 | 79.0 | 86.7 | 94.0 | 83.4 | 77.7 |
| Configuration | AD | DL | Others | Avg. |
| (a) Best single (Qwen3-VL-4B) | 74.1 | 63.6 | 77.1 | 74.8 |
| (b) Multimodal searcher | 74.3 | 64.8 | 82.6 | 78.9 |
| (c) Heterogeneous agents (MV) | 74.7 | 65.4 | 85.8 | 81.3 |
| (d) Defect expert (MV) | 80.5 | 73.6 | 85.8 | 83.3 |
| (e) Reliability controller | 82.6 | 79.6 | 86.8 | 85.2 |
| (f) SiGMA without defect expert | 75.1 | 65.8 | 86.7 | 82.0 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Symbol | Value | Role |
| Base agents | – | 1 | Retrieved normal references per agent input |
| Defect expert | – | Size of the category-specific normal bank | |
| Noise-robust calibration | 0.99 | Quantile of normal residuals defining | |
| 3 | Expansion coefficient defining | ||
| 0.99 | Quantile of normal area fractions defining | ||
| 0.01 | Margin of the spatial-coherence threshold |
| Readers | Similar reference accuracy (%) | Random reference accuracy (%) |
| 1 | 84.7 | 84.7 |
| 3 | 85.2 | 85.2 |
| 4 | 85.1 | 85.0 |
| 5 | 85.2 | 85.1 |
| 10 | 85.2 | 85.0 |
| Anomaly | Defect | Object | |||||||
| Configuration | Det. | Cls. | Loc. | Desc. | Anal. | Cls. | Anal. | Avg. | |
| (a) | Best single MLLM (Qwen3-VL-4B) | 74.1 | 59.7 | 63.6 | 70.1 | 80.0 | 92.4 | 83.4 | 74.8 |
| (b) | + Multimodal searcher | 74.3 | 72.0 | 64.8 | 78.8 | 82.1 | 94.8 | 85.5 | 78.9 |
| (c) | + Heterogeneous agents (majority vote) | 74.7 | 77.6 | 65.4 | 82.4 | 85.6 | 96.3 | 87.1 | 81.3 |
| (d) | + Defect expert (majority vote) | 80.5 | 77.6 | 73.6 | 82.4 | 85.6 | 96.3 | 87.1 | 83.3 |
| (e) | + Reliability controller (SiGMA) | 82.6 | 79.2 | 79.6 | 83.2 | 86.4 | 97.3 | 88.1 | 85.2 |
| Use of expert evidence | AD | DL | FPR |
| No defect expert | 75.1 | 65.8 | 17.8 |
| Raw overlay on agent inputs | 69.8 | 75.0 | 30.6 |
| Raw map through readers | 73.5 | 76.8 | 24.2 |
| Calibrated map through readers | 80.9 | 79.6 | 12.8 |
| + Coherence gate (SiGMA) | 82.6 | 79.6 | 9.4 |
| Method | Accuracy (%) | In-domain training/ optimizing | Inference and token cost | Latency (s/sample) |
| Echo | 77.3 | Retrieval and expert-guided MLLM reasoning | 1.51 | |
| Reason-IAD | 79.4 | 10 latent-reasoning iterations | — | |
| JUDO | 81.2 | Single-pass inference with a specialized MLLM | 1.23 | |
| AgentIAD | 82.9 | Multi-round reasoning with adaptive tool use | 2.90 | |
| SiGMA | 85.2 | Parallel agent inference and CPU-based reliability aggregation; 25.04–228.44 tokens/question | 1.88 |
| Configuration | MLLM calls | Latency (s) | Peak mem. (GiB) | Avg. acc. (%) |
| Single 7B MLLM | 1 | 1.2 | 15.2 | 81.2 |
| SiGMA, three base agents | 3 + up to 3 | 1.5 | 85 | 84.4 |
| SiGMA, eight base agents | 8 + up to 3 | 1.88 | 105 | 85.2 |