Visual anomaly detectors identify deviations from known-normal data, but their anomaly signals may mix evidence of actual anomalies with benign visual variation. We investigate whether automated interpretability can augment visual anomaly detectors by identifying and intervening on different components of this signal. We decompose PatchCore nearest-normal residuals into sparse features using Sparse Autoencoders (SAEs), and provide high-activation and contrastive non-active examples to a Multimodal LLM, which describes each feature and labels it as anomaly, distractor, or uncertain. These labels guide interventions in the SAE hidden representation, where distractor features are suppressed and anomaly features amplified. The edited representation is then used to reconstruct patch embeddings, which are rescored with PatchCore. Across 40 categories from four benchmarks, applying both interventions jointly improves macro-average image-level AUROC from 0.8724 to 0.8857 on source data and from 0.8066 to 0.8210 under synthetic corruptions. On three additional RobustAD categories with real acquisition shifts, the same interventions improve AUROC from 0.8745 to 0.9056 on source data and from 0.6069 to 0.6599 under real acquisition shifts. Finally, individual feature interventions across all 43 categories show that the MLLM labels are aligned in aggregate with how features differently affect normal and anomalous images.
Figures & tables
Figure 1: Overview of the pipeline. (1) Given a frozen vision backbone and a PatchCore memory bank, we compute anomaly residuals relative to the nearest normal patch embeddings and decompose them using a tied Sparse Autoencoder (SAE). (2) A multimodal LLM is prompted with top-activating and contrastive non-active examples to generate SAE feature descriptions and assign an anomaly, distractor, or uncertain label. (3) We amplify anomaly features and suppress distractor ones. Then we reconstruct the edited patch embeddings and rescore them with PatchCore.
Method
Intervention
Evaluation
VisA
MVTec AD
MVTec LOCO
MVTec AD 2
All
(12)
(15)
(5)
(8)
(40)
Original PatchCore
Source
0.8953
0.9859
0.8430
0.6440
0.8724
Corrupted
0.7868
0.9311
0.7831
0.6177
0.8066
PatchCore + extra normals
Source
0.8955
0.9866
0.8482
0.6521
0.8751
Corrupted
0.7862
0.9316
0.7874
0.6236
0.8083
DINOv3 linear probe
Source
0.8979
0.8544
0.6199
0.5998
0.7872
Table 1: Image-level AUROC on source and synthetically corrupted data, macro-averaged across categories. Best results are in bold. Random interventions are averaged over 10 seeds.
Method
Intervention
Evaluation
PCB
MetalParts
PiledBags
All (3)
Original PatchCore
Source
0.8926
0.7423
0.9887
0.8745
Shifted
0.3943
0.6415
0.7849
0.6069
PatchCore + extra normals
Source
0.9202
0.7386
0.9889
0.8826
Shifted
0.4154
0.6413
0.7828
0.6131
DINOv3 linear probe
Source
0.9998
0.7495
0.8057
0.8517
Shifted
0.5165
0.7106
0.7177
0.6482
Table 2: Image-level AUROC on RobustAD source and acquisition shifts, macro-averaged across categories. Best results are in bold. Random interventions are averaged over 10 seeds.
Figure 2: Effects of suppressing and amplifying SAE features across 43 categories. Points show normalized mean score changes per feature, and diamonds show category-balanced means by label.
Figure 3: Examples of anomaly and distractor SAE features. Italic text gives the MLLM-generated feature description. The top row shows a test image and three top-activating calibration examples shown to the MLLM. The bottom row shows, for the test image, the SAE feature map and PatchCore anomaly maps before intervention, after suppression, and after amplification. The three PatchCore anomaly maps for each example share a min-max color scale.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
VisA
MVTec AD
MVTec LOCO
MVTec AD 2
All
(12)
(15)
(5)
(8)
(40)
Training MSE
21.11
30.41
26.86
42.86
29.67
Test MSE
25.00
52.09
31.80
57.43
42.50
Training cosine
0.550
0.633
0.568
0.646
0.602
Test cosine
0.504
0.513
0.550
0.500
0.512
Recovered AUROC (%)
97.23
97.55
92.92
91.15
95.95
Appendix
Table 3: SAE reconstruction fidelity evaluated on the test split, macro-averaged across categories. Recovered AUROC is reported as a percentage of the original PatchCore AUROC.
Metric
PCB
MetalParts
PiledBags
All (3)
Training MSE
13.55
32.28
28.36
24.73
Test MSE
13.49
32.82
33.10
26.47
Training cosine
0.599
0.553
0.679
0.610
Test cosine
0.586
0.538
0.645
0.590
Recovered AUROC (%)
110.37
84.72
98.74
98.74
Dead features / 64
0.00
0.00
4.00
1.33
Appendix
Table 4: SAE reconstruction fidelity evaluated on the RobustAD source test split. Recovered AUROC is reported as a percentage of the original PatchCore AUROC.
Metric
Tied, bias-free
Untied, with biases
Test reconstruction MSE
42.50
43.35
Recovered AUROC (%)
95.95
95.82
Dead features / 64
0.05
11.43
Removals increasing patch distance (%)
0.00
3.34
Appendix
Table 5: Comparison of tied and untied SAEs across the 40 RQ1 categories.
Width
TopK
Source
Corrupted
m
k
Suppression
Amplification
Joint
Suppression
Amplification
Joint
Original PatchCore
0.8592
0.7836
32
2
0.8816
0.9418
0.9465
0.7998
0.8544
0.8571
32
4
0.8924
0.9263
0.9373
0.8027
0.8545
0.8543
32
8
0.8903
0.9123
0.8903
0.8023
0.8414
0.8023
64
2
0.8858
0.8821
0.8965
0.7969
0.8097
0.8126
Appendix
Table 6: Width and TopK sensitivity on VisA capsules and MVTec AD pill , macro-averaged across the two categories.
Source
Corrupted
Intervention
Error retained
Error omitted
Error retained
Error omitted
None
0.8724
0.8368
0.8066
0.7760
Suppression
0.8808
0.8516
0.8126
0.8102
Amplification
0.8819
0.8524
0.8152
0.7964
Joint
0.8857
0.8551
0.8210
0.8136
Appendix
Table 7: Effect of retaining the SAE reconstruction error across the 40 RQ1 categories, macro-averaged across categories. For no intervention, the error-retained condition corresponds to original PatchCore, while the error-omitted condition corresponds to the SAE-only reconstruction.
Intervention
Source
Corrupted
None
0.8368
0.7760
Suppression
0.7807
0.7274
Amplification
0.8273
0.7715
Joint
0.7899
0.7348
Appendix
Table 8: AUROC for random interventions without SAE reconstruction error across the 40 RQ1 categories, macro-averaged across categories. None corresponds to the unedited SAE-only reconstruction. Random interventions use a single feature assignment.
Source
Corrupted
Multiplier
Amplification
Joint
Amplification
Joint
1×
0.8724
0.8808
0.8066
0.8126
1.5×
0.8761
0.8831
0.8155
0.8199
2×
0.8785
0.8836
0.8210
0.8245
3×
0.8801
0.8789
0.8246
0.8259
4×
0.8745
0.8735
0.8236
0.8237
Appendix
Table 9: Effect of the anomaly-feature amplification multiplier on amplification-only and joint intervention across the 40 RQ1 categories, macro-averaged across categories. Best results in each column are in bold.
Source
Corrupted
MLLM
Suppression
Amplification
Joint
Suppression
Amplification
Joint
Original PatchCore
0.8653
0.7951
GPT-5.6 Sol
0.8779
0.8808
0.8849
0.8024
0.8157
0.8243
Gemma 4 31B-IT
0.8786
0.8853
0.8899
0.8006
0.8165
0.8207
Qwen3.5-27B
0.8773
0.8708
0.8791
0.8043
0.8055
0.8149
Appendix
Table 10: Effect of the MLLM used for feature interpretation on a fixed subset of ten categories, macro-averaged across categories.
Detector
Intervention
Source
Corrupted
AUROC
Δ
AUROC
Δ
PatchCore
Original
0.8969
–
0.8428
–
Suppression
0.9105
+1.36
0.8464
+0.37
Amplification
0.9166
+1.96
0.8714
+2.86
Joint
0.9255
+2.86
0.8716
+2.88
FRE
Original
0.8475
–
0.7488
–
Appendix
Table 11: Anomaly detector ablation on five categories. Deltas are AUROC percentage points relative to the corresponding original detector. Best intervention results for each detector and evaluation setting are in bold.
Source
Corrupted
Backbone
Original
Suppression Δ
Joint Δ
Original
Suppression Δ
Joint Δ
DINOv3-L/16
0.8969
+1.36
+2.86
0.8428
+0.37
+2.88
SwinV2-L
0.8073
+1.41
+1.78
0.7177
+1.02
+1.22
WideResNet-50-2
0.9213
-0.19
-0.19
0.8297
-0.07
-0.07
Appendix
Table 12: Vision backbone ablation over five categories. Deltas are AUROC percentage points relative to the corresponding original detector. Largest intervention gains are in bold.
Dept. of Engineering for Innovation Medicine, University of Verona, Italy · School of Computer Science and Engineering, Beihang University, China · Dept. of Engineering for Innovation +1