VisionLogic: Discovering and Grounding Decision-Relevant Visual Concepts
Authors: Chuqin Geng, Yuhe Jiang, Li Zhang, Zhaoyue Wang, Haolin Ye, Mark Zhang, Jingkai Xu, Xujie Si
Organizations: School of Computer Science, McGill University, Montreal, Canada · Department of Computer Science, University of Toronto, Toronto, Canada
Concept-based explanations help users understand vision models through recognizable visual patterns. However, existing methods often rely on correlational signals without directly validating which image cues support prediction-relevant internal features. To this end, we introduce VisionLogic, a post-hoc framework that grounds these features in visual concepts through intervention-based validation. VisionLogic first identifies compact sets of features whose contributions reproduce the model's original prediction. It then represents their activation states as predicates using class-specific thresholds. An iterative refinement procedure grounds these predicates in visual regions through ablation tests. A region is accepted when its removal deactivates the corresponding predicate, linking the feature's numerical role to visual evidence in the input. The same predicates allow us to examine how features are activated, selected, and reused across images and classes. Across CNNs and vision transformers on ImageNet-1k, we find that only a few features are selected to explain each prediction, and frequently active features are not always selected. In a large-scale human evaluation with 465 participants, VisionLogic significantly improves participants' understanding of model behavior over established methods ACE and CRAFT. Code is available at https://github.com/allengeng123/VisionLogic.
Figures & tables
Figure 1: Intervention-based visual grounding. Neuron-specific attribution maps guide the search for candidate regions. We ablate these regions by replacing their content with a blurred version. If the predicate remains active, we iteratively refine the proposals by including weaker heatmap responses. The first union of boxes whose ablation deactivates the target predicate is accepted. Foreground segmentation highlights the visual content within the accepted boxes.
Figure 2: Activation and selection frequencies on source images in four models. Each hexbin counts class-specific predicates, with a shared logarithmic color scale. The dashed line indicates equal frequencies. Selection implies activation on source images by threshold construction. The large gaps below the line show that many frequently active features are selected much less often.
Figure 3: Cross-class reuse of signed features in four models. For each feature selected in the source images, we count the classes in which it is selected in at least 1% of source examples. The vertical axis is shared and logarithmic; horizontal ranges differ across models. Predicates for shared features use class-specific thresholds, and feature indices are specific to each model.
Husky vs. Wolf
Otter vs. Beaver
Kit Fox vs. Red Fox
Session n ∘
1
2
3
Utility
1
2
3
Utility
1
2
3
Utility
Baseline
65.7
68.6
70.3
1.00
84.4
90.3
92.2
1.00
84.1
89.0
84.1
1.00
Control
55.3
63.6
70.0
0.92
85.1
88.3
92.9
1.00
80.8
79.2
79.2
0.93
ACE
60.4
71.1
74.6
1.01
80.4
85.7
90.5
0.96
80.6
83.2
76.2
0.93
CRAFT
55.5
60.8
65.3
0.89
86.3
90.9
90.9
1.00
76.8
81.8
76.8
0.92
VisionLogic
74.8
90.0
91.0
1.25
96.8
98.4
99.2
1.10
84.1
84.5
82.9
0.98
Table 1: Utility scores in three application scenarios. The Utility benchmark measures how well explanations help users identify general rules that transfer to unseen instances. During training, participants are shown images along with the explanations and model predictions, and are asked to infer the underlying decision rules. At test time, the benchmark evaluates participants’ accuracy in predicting the model’s output on novel images. Higher Utility scores indicate that the explanations provide more useful information for understanding the model’s behavior on new samples. For each scenario, the first and second best results are in bold and underlined respectively. VisionLogic achieves higher Utility scores than ACE and CRAFT across all three scenarios, with statistically significant improvements over both methods in the first two.
Figure 5: Participant accuracy distributions for each explanation method. VisionLogic achieves higher mean accuracy than the ACE and CRAFT baselines across all three scenarios.
Figure 6: Human-recognizable visual concepts grounded through predicates. In each image, the predicate appears above the frame, and the colored overlay highlights the associated visual content.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Attribution method
Predicate success
Image coverage
Median occupancy
ViT-B/16
Grad-CAM
242/276 ( 87.7% )
98/99 ( 99.0% )
23.0%
Score-CAM
254/276 ( 92.0% )
97/99 ( 98.0% )
33.1%
Grad-CAM → Score-CAM
274/276 ( 99.3% )
99/99 ( 100.0% )
24.0%
ResNet-50
Grad-CAM
95/129 ( 73.6% )
79/100 ( 79.0% )
41.2%
Score-CAM
109/129 ( 84.5% )
86/100 ( 86.0% )
57.9%
Grad-CAM → Score-CAM
116/129 ( 89.9% )
90/100 ( 90.0% )
46.3%
Appendix
Table 2: Comparison of attribution methods for visual grounding. All methods use the same 100-image sample per model and predicate-deactivation test. Image coverage is measured over images containing at least one eligible predicate. Median occupancy is computed over successful predicate tests. The combined search applies Grad-CAM first and Score-CAM only when the Grad-CAM search fails.
Model
Guided success
Baseline success
Image coverage
Guided occupancy
Baseline occupancy
ViT-B/16
274/276 (99.3%)
256/276 (92.8%)
99/99 (100%)
24.0% [11.8, 50.0]
27.9% [16.8, 56.4]
ResNet-50
116/129 (89.9%)
105/129 (81.4%)
90/100 (90.0%)
46.3% [24.8, 70.8]
60.7% [37.0, 78.7]
Appendix
Table 3: Guided grounding versus area-matched search on 100 images per model. The baseline mirrors both stages of the guided search with approximately area-matched proposals placed at random locations. Image coverage refers to the guided search. Occupancy is summarized over successful predicate tests as the median percentage of image area covered; brackets give the interquartile range.
Model
Dim.
Source images
Vocab.
Source
Held-out
ViT-B/16
768
1,206,951
95
7/3
7/3
ResNet-50
2,048
1,169,769
7
2/1
2/1
ConvNeXt-Base
1,024
1,216,462
81
9/4
9/4
Swin-T
768
1,153,631
133
15/7
15/7
Appendix
Table 4: Image populations and predicate counts. Dim. gives the number of features supplied to the classification head, and Vocab. gives the median number of source predicates per predicted class. Source and Held-out report median active/selected predicate counts per image. Models follow the order used in the main figures.
Figure 7: Within-class concentration of predicate activations and selections on source images. Predicates are ranked separately by activation and selection frequency, and curves show the cumulative share of occurrences. Colors identify classes nearest the 10th, 50th, and 90th percentiles of selection concentration, measured by 1−N80/∣Vc∣ . Solid lines show selection and dashed lines activation; both axes are normalized to [0,1] .
Figure 8: Source predicate coverage and threshold retention on held-out images. Eligible shows the percentage of all selected occurrences with a matching class-specific source predicate. Retained shows the percentage of eligible occurrences that satisfy the corresponding fixed threshold. Values are printed above the bars, showing how well the source predicates cover selected features and retain activation on unseen images.
Figure 9: The online questionnaire begins with a consent form.
Figure 10: The online questionnaire displaying an example of the training session with the original image, the explanation, and the model prediction.
Figure 11: The online questionnaire displaying the testing session with a reservoir containing all examples during the training phase.
Figure 12: Complete statistical test results for scenario 1, page 1.
Figure 13: Complete statistical test results for scenario 1, page 2.
Figure 14: Complete statistical test results for scenario 2, page 1.
Figure 15: Complete statistical test results for scenario 2, page 2.
Figure 16: Complete statistical test results for scenario 3, page 1.
Figure 17: Complete statistical test results for scenario 3, page 2. The overall test does not reject at the 0.05 level.
Concept-based explanations offer a promising approach for explaining the predictions of deep neural networks in terms of high-level, human-understandable concepts. However, existing methods either do not establish a causal connection between the concepts and model predictions or are limited in expressivity and only able to infer causal explanations involving single concepts. At the same time, the parallel line of work on formal abductive and contrastive explanations computes the minimal set of input features causally relevant for model outcomes but only considers low-level features such as pixels. Merging these two threads, in this work, we propose the notion of concept-based abductive and contrastive explanations that capture the minimal sets of high-level concepts causally relevant for model outcomes. We then present a family of algorithms that enumerate all minimal explanations while using concept erasure procedures to establish causal relationships. By appropriately aggregating such explanations, we are not only able to understand model predictions on individual images but also on collections of images where the model exhibits a user-specified, common behavior. We evaluate our approach on multiple models, datasets, and behaviors, and demonstrate its effectiveness in computing helpful, user-friendly explanations.
Ronaldo Canizales, Divya Gopinath, Corina Păsăreanu +1
Colorado State University, Fort Collins CO, USA · KBR Inc., NASA Ames, Moffett Field CA, USA · Carnegie Mellon University, Pittsburgh PA, USA
The growing demand for transparency in automated decision-making has propelled eXplainable Artificial Intelligence (XAI) to the forefront of machine learning research. In computer vision, however, existing explanation methods often prioritize end-user accessibility at the expense of formal guarantees, leaving a critical gap between practical utility and theoretical rigor. In this paper, we address this gap by introducing OPTIMUS, a novel framework for generating concept-based visual explanations for deep classification models. OPTIMUS explanations take the form of visual heatmaps that not only remain interpretable to end users, but are grounded in the well-established theory of prime implicants, providing formal guarantees that have been largely absent from existing saliency-based methods. Specifically, OPTIMUS explanations satisfy two desirable properties: sufficiency, ensuring that the highlighted concepts provably guarantee the classifier's prediction, and minimality, ensuring that no strict subset of those concepts retains this guarantee. Together, these properties yield explanations that are both logically tight and visually coherent. We validate our approach on a visual classification benchmark, demonstrating that OPTIMUS heatmaps naturally and faithfully surface the decision-relevant concepts underlying model predictions.
Arthur Hoarau, Chenrui Zhu, Vu Linh Nguyen
Université de Lorraine, CentraleSupélec Loria, CNRS, Metz, France · Université de technologie de Compiègne UMR CNRS 7253 Heudiasyc, France
When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-language models (VLMs) such as Gemini-3-Pro and GPT-5 only respond with text, which can be difficult for users to verify. We present SketchVLM, a training-free, model-agnostic framework that enables VLMs to produce non-destructive, editable SVG overlays on the input image to visually explain their answers. Across seven benchmarks spanning visual reasoning (maze navigation, ball-drop trajectory prediction, and object counting) and drawing (part labeling, connecting-the-dots, and drawing shapes around objects), SketchVLM improves visual reasoning task accuracy by up to +28.5 percentage points and annotation quality by up to 1.48x relative to image-editing and fine-tuned sketching baselines, while also producing annotations that are more faithful to the model's stated answer. We find that single-turn generation already achieves strong accuracy and annotation quality, and multi-turn generation opens up further opportunities for human-AI collaboration. An interactive demo and code are at https://sketchvlm.github.io/.