VisionLogic: Discovering and Grounding Decision-Relevant Visual Concepts
Authors: Chuqin Geng, Yuhe Jiang, Li Zhang, Zhaoyue Wang, Haolin Ye, Mark Zhang, Jingkai Xu, Xujie Si
Organizations: School of Computer Science, McGill University, Montreal, Canada · Department of Computer Science, University of Toronto, Toronto, Canada
Concept-based explanations help users understand vision models through recognizable visual patterns. However, existing methods often rely on correlational signals without directly validating which image cues support prediction-relevant internal features. To this end, we introduce VisionLogic, a post-hoc framework that grounds these features in visual concepts through intervention-based validation. VisionLogic first identifies compact sets of features whose contributions reproduce the model's original prediction. It then represents their activation states as predicates using class-specific thresholds. An iterative refinement procedure grounds these predicates in visual regions through ablation tests. A region is accepted when its removal deactivates the corresponding predicate, linking the feature's numerical role to visual evidence in the input. The same predicates allow us to examine how features are activated, selected, and reused across images and classes. Across CNNs and vision transformers on ImageNet-1k, we find that only a few features are selected to explain each prediction, and frequently active features are not always selected. In a large-scale human evaluation with 465 participants, VisionLogic significantly improves participants' understanding of model behavior over established methods ACE and CRAFT. Code is available at https://github.com/allengeng123/VisionLogic.
Figures & tables
Figure 1: Intervention-based visual grounding. Neuron-specific attribution maps guide the search for candidate regions. We ablate these regions by replacing their content with a blurred version. If the predicate remains active, we iteratively refine the proposals by including weaker heatmap responses. The first union of boxes whose ablation deactivates the target predicate is accepted. Foreground segmentation highlights the visual content within the accepted boxes.
Figure 2: Activation and selection frequencies on source images in four models. Each hexbin counts class-specific predicates, with a shared logarithmic color scale. The dashed line indicates equal frequencies. Selection implies activation on source images by threshold construction. The large gaps below the line show that many frequently active features are selected much less often.
Figure 3: Cross-class reuse of signed features in four models. For each feature selected in the source images, we count the classes in which it is selected in at least 1% of source examples. The vertical axis is shared and logarithmic; horizontal ranges differ across models. Predicates for shared features use class-specific thresholds, and feature indices are specific to each model.
Husky vs. Wolf
Otter vs. Beaver
Kit Fox vs. Red Fox
Session n ∘
1
2
3
Utility
1
2
3
Utility
1
2
3
Utility
Baseline
65.7
68.6
70.3
1.00
84.4
90.3
92.2
1.00
84.1
89.0
84.1
1.00
Control
55.3
63.6
70.0
0.92
85.1
88.3
92.9
1.00
80.8
79.2
79.2
0.93
ACE
60.4
71.1
74.6
1.01
80.4
85.7
90.5
0.96
80.6
83.2
76.2
0.93
CRAFT
55.5
60.8
65.3
0.89
86.3
90.9
90.9
1.00
76.8
81.8
76.8
0.92
VisionLogic
74.8
90.0
91.0
1.25
96.8
98.4
99.2
1.10
84.1
84.5
82.9
0.98
Table 1: Utility scores in three application scenarios. The Utility benchmark measures how well explanations help users identify general rules that transfer to unseen instances. During training, participants are shown images along with the explanations and model predictions, and are asked to infer the underlying decision rules. At test time, the benchmark evaluates participants’ accuracy in predicting the model’s output on novel images. Higher Utility scores indicate that the explanations provide more useful information for understanding the model’s behavior on new samples. For each scenario, the first and second best results are in bold and underlined respectively. VisionLogic achieves higher Utility scores than ACE and CRAFT across all three scenarios, with statistically significant improvements over both methods in the first two.
Figure 5: Participant accuracy distributions for each explanation method. VisionLogic achieves higher mean accuracy than the ACE and CRAFT baselines across all three scenarios.
Figure 6: Human-recognizable visual concepts grounded through predicates. In each image, the predicate appears above the frame, and the colored overlay highlights the associated visual content.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Attribution method
Predicate success
Image coverage
Median occupancy
ViT-B/16
Grad-CAM
242/276 ( 87.7% )
98/99 ( 99.0% )
23.0%
Score-CAM
254/276 ( 92.0% )
97/99 ( 98.0% )
33.1%
Grad-CAM → Score-CAM
274/276 ( 99.3% )
99/99 ( 100.0% )
24.0%
ResNet-50
Grad-CAM
95/129 ( 73.6% )
79/100 ( 79.0% )
41.2%
Score-CAM
109/129 ( 84.5% )
86/100 ( 86.0% )
57.9%
Grad-CAM → Score-CAM
116/129 ( 89.9% )
90/100 ( 90.0% )
46.3%
Appendix
Table 2: Comparison of attribution methods for visual grounding. All methods use the same 100-image sample per model and predicate-deactivation test. Image coverage is measured over images containing at least one eligible predicate. Median occupancy is computed over successful predicate tests. The combined search applies Grad-CAM first and Score-CAM only when the Grad-CAM search fails.
Model
Guided success
Baseline success
Image coverage
Guided occupancy
Baseline occupancy
ViT-B/16
274/276 (99.3%)
256/276 (92.8%)
99/99 (100%)
24.0% [11.8, 50.0]
27.9% [16.8, 56.4]
ResNet-50
116/129 (89.9%)
105/129 (81.4%)
90/100 (90.0%)
46.3% [24.8, 70.8]
60.7% [37.0, 78.7]
Appendix
Table 3: Guided grounding versus area-matched search on 100 images per model. The baseline mirrors both stages of the guided search with approximately area-matched proposals placed at random locations. Image coverage refers to the guided search. Occupancy is summarized over successful predicate tests as the median percentage of image area covered; brackets give the interquartile range.
Model
Dim.
Source images
Vocab.
Source
Held-out
ViT-B/16
768
1,206,951
95
7/3
7/3
ResNet-50
2,048
1,169,769
7
2/1
2/1
ConvNeXt-Base
1,024
1,216,462
81
9/4
9/4
Swin-T
768
1,153,631
133
15/7
15/7
Appendix
Table 4: Image populations and predicate counts. Dim. gives the number of features supplied to the classification head, and Vocab. gives the median number of source predicates per predicted class. Source and Held-out report median active/selected predicate counts per image. Models follow the order used in the main figures.
Figure 7: Within-class concentration of predicate activations and selections on source images. Predicates are ranked separately by activation and selection frequency, and curves show the cumulative share of occurrences. Colors identify classes nearest the 10th, 50th, and 90th percentiles of selection concentration, measured by 1−N80/∣Vc∣ . Solid lines show selection and dashed lines activation; both axes are normalized to [0,1] .
Figure 8: Source predicate coverage and threshold retention on held-out images. Eligible shows the percentage of all selected occurrences with a matching class-specific source predicate. Retained shows the percentage of eligible occurrences that satisfy the corresponding fixed threshold. Values are printed above the bars, showing how well the source predicates cover selected features and retain activation on unseen images.
Figure 9: The online questionnaire begins with a consent form.
Figure 10: The online questionnaire displaying an example of the training session with the original image, the explanation, and the model prediction.
Figure 11: The online questionnaire displaying the testing session with a reservoir containing all examples during the training phase.
Figure 12: Complete statistical test results for scenario 1, page 1.
Figure 13: Complete statistical test results for scenario 1, page 2.
Figure 14: Complete statistical test results for scenario 2, page 1.
Figure 15: Complete statistical test results for scenario 2, page 2.
Figure 16: Complete statistical test results for scenario 3, page 1.
Figure 17: Complete statistical test results for scenario 3, page 2. The overall test does not reject at the 0.05 level.