cs.CVOct 5, 2026

Spatial Supervision Without Attribution Optimization: Improving Post-Hoc Class Activation Maps via Box-Guided Evidence Routing

Authors: Wenhao Liang, Liangwei Nathan Zheng, Lin Yue, Wei Emma Zhang, Mingyu Guo, Weitong Chen

Organizations: Adelaide University, Adelaide, Australia

Abstract

Post-hoc class activation maps (CAMs) are a standard tool for inspecting the evidence behind an image classifier's predictions, yet nothing in ordinary training encourages these maps to be spatially appropriate. We study whether inexpensive spatial supervision can improve a classifier's own predicted-class Grad-CAM without ever optimizing an attribution map. Box-Guided Evidence Routing (BGER) trains a lightweight gate on the final feature map under box or mask supervision and routes classification through the gated features, while Grad-CAM is computed separately at the pre-gate representation, so the evaluated map never enters the training objective. With a BCE routing loss, BGER raises MaxBoxAccV2 from 0.5840.584 to 0.7150.715 on CUB-200-2011 and from 0.7570.757 to 0.8320.832 on Stanford Dogs at comparable accuracy. Matched controls attribute most of the ResNet-50 gain to the spatial supervision reshaping the backbone rather than to routing itself: when classification bypasses the gate, most of the improvement remains, and detaching gradients through the gate leaves the ResNet-50 result nearly unchanged. The same detachment preserves most of the gain in two DenseNet-121 chest X-ray settings but removes the apparent gain on Swin-T, and directly supervising the CAM reaches stronger localization at a larger accuracy cost. Overall, spatial supervision can improve separately evaluated post-hoc CAMs, but both the mechanism and the size of the benefit depend on the architecture and the evaluation setting.

Figures & tables

Appendix figures & tables35 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations

    Aug 12, 2026AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini +2Transformer InterpretabilityFeature Attribution

  2. Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations

    Jul 26, 2026Xianghao Jiao, Ruoyu Chen, Wei Wang +6Feature AttributionImage Classification