Current explanation methods for contrastive vision-language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce Mask-guided Adaptive Counterfactual Explanations (MACE), a targeted visual counterfactual method designed specifically for CLIP zero-shot classification. MACE constructs an editable region from either source attribution or source-target attribution differences and expands the mask only when needed to reach a specified target class. A latent diffusion inpainting model then modifies the selected region, while a frozen CLIP model provides modification guidance and anchors the remaining image content to the original input. We evaluate MACE on ImageNet, Food-101, Oxford Pets, and CUB-200. The source-mask variant achieves the highest target top-1 success rate across all four datasets, while the difference-mask variant produces the smallest pixel level and perceptual changes and the best realism scores. Both variants improve proximity and realism over a Stable Diffusion-only baseline using the same generative backbone. These results show that adaptive mask-guided editing produces effective CLIP counterfactuals. They further reveal a tradeoff between counterfactual validity and source-image preservation.
Figures & tables
Figure 1 : Overview of MACE. Given a source image and a target class, CLIP attribution identifies regions that support the source prediction or favor the source over the target. Starting with the top 25% of attribution values ( M0,r0=0.25 ), MACE progressively expands the resulting editable mask in 5% steps until a candidate reaches the target class and then uses CLIP-guided latent diffusion to inpaint only the selected region while preserving the remaining image content.
Dataset
Method/Variant
Validity ↑
ℓ1norm↓
ℓ2norm↓
LPIPS ↓
SSIM ↑
FID ↓
KID ↓
ImageNet
MACE s
0.9700
0.0742
0.1603
0.2526
0.6948
49.7978
0.0035
MACE d
0.7805
0.0407
0.1059
0.1517
0.7940
44.8410
0.0007
Stable Diffusion
0.7195
0.3346
0.4066
0.8397
0.1301
84.1666
0.0051
Food-101
MACE s
0.9990
0.0554
0.1249
0.2177
0.7322
23.5741
0.0018
MACE d
0.8744
0.0402
0.1003
0.1674
0.7914
22.3334
0.0015
Stable Diffusion
0.9701
0.3437
0.4218
0.8219
0.1728
151.4007
0.0308
Table 1: Quantitative comparison on ImageNet, Food-101, Oxford Pets, and CUB-200. We report validity as the target top-1 success rate, followed by proximity and realism metrics. Stable Diffusion-only is the baseline. Lower values are better for FID, KID, and distance metrics, whereas higher values are better for validity and SSIM. Bold and underlined values indicate the best and second-best results, respectively, within each dataset and metric column.
Figure 2 : Qualitative comparison across ImageNet, Food-101, Oxford Pets, and CUB-200. Each row contains one success case on the left and one failure case on the right. For each success case, we show the source image and target class, the editable masks and counterfactuals produced by MACE s and MACE d , and the Stable Diffusion-only (SD) result. All methods reach the target class in these examples. For each failure case, we show the source image followed by the outputs of all three methods. None reaches the target class, except in the Food-101 example, where only MACE s succeeds.
Figure 3 : Validity ablations across all four datasets. Solid and dashed lines show the source and difference masks, respectively. Panels (a) and (b) vary the CLIP guidance scale and fixed mask fraction, respectively, while panel (c) reports validity when each mask is adaptively expanded up to the indicated maximum fraction.
Figure 4 : Proximity ablations measured by normalized ℓ1 distance, where lower values indicate smaller image changes. Solid and dashed lines show the source and difference masks, respectively. Panels (a) and (b) vary the CLIP guidance scale and fixed mask fraction, while (c) adaptively expands each mask up to the indicated maximum fraction.
School of Cyber Science and Technology, Sun Yat-Sen University, Guangdong Province, China · Institute of Perception, PC Laboratory, Guangdong Province, China · University of Chinese Academy of Sciences, Beijing, China +3