AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
Authors: Mohamed Eltahir, Fardows Adam, Duaa M. Tahir, Lama Alamoudi, Sana Ammar, Atheer A. Alboloshi, Jory Albluey, Tanveer Hussain, +1 more
Organizations: King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia · Department of Computer Science, Edge Hill University, Ormskirk, England
When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free, task-agnostic, black-box visual rationale constructed from the output head. The image is cut into K row and K column bands, each shown alone to the frozen model along with the query in the format of a yes/no relevance question. The outer product of the row and column ``yes'' posteriors gives the query-conditioned spatial map. Crucially, by defining a fixed read-out R (e.g., expectation, maximum) on top of AnswerMap, we can derive continuous outputs like location natively. This bypasses the reliance on discrete text tokens for continuous-output tasks and guarantees an image-dependent answer by construction. However, a rationale can be confabulated, so we validate AnswerMap across four models and three query distributions with two tests: (a) agreement with the model's own generated point and (b) deletion of the map's region. The map lands where the model points (AUC 0.85 against 0.38 for attention), and deleting its region flips 53% of correct answers (against 19% for attention's). Beyond establishing faithfulness, we demonstrate the map's task-agnostic utility through three distinct read-outs: its maximum flags hallucinated objects without generation, its expectation localizes correctly when the model's own pointing fails, and its top-mass region, fed back as a crop, fixes half of the model's wrong answers. AnswerMap thus offers a new lens on VLM interpretability and, through its read-outs, a new output interface for visual tasks beyond text tokens.
Figures & tables
Figure 1: Three ways to explain one frozen VLM’s answer. A text rationale is in a mismatched modality. An internal read-out is in the right modality but originates too early and needs the weights. AnswerMap reads the output head one band at a time, the map is visual and black-box, and a fixed read-out computes the answer from it.
Figure 2: AnswerMap (multigrid) from GPT-6-sol through its API, one query each.
Rationale
Spatial
Any answer
Cell score
Black box
Passes per map
Words (text generation)
✗
✓
none
✓
1
Attention
✓
✓
routing
✗
1
Gradient relevancy (e.g., T-MM)
✓
✓
routing
✗
1
Token activation (e.g., TAM)
✓
✗
activation
✗
1
Occlusion
✓
✓
necessity
✓
K2
AnswerMap
✓
✓
sufficiency
✓
2K
Table 1: Positioning. Any answer: the rationale exists for non-word answers such as coordinates. Cell score: what a cell’s value measures (where attention routed, how strongly a word activates, whether the cell is necessary when removed from the full image, or whether a band shown alone suffices). Black box: needs only the model’s output token probabilities. Passes: forward passes for one K×K map.
Figure 3: The AnswerMap operator. K row and K column band questions fill a K×K map through the outer product, and a read-out computes the answer from the map.
Figure 4: Multigrid product. Coprime grids are probed separately, upsampled to a shared grid, and combined, so a region survives only if every band width endorses it.
Figure 5: The two-test faithfulness protocol. Left, agreement, score the full map at the model’s own generated point. Right, deletion, remove the map’s top-mass region against a matched random region and re-ask.
RefCOCOg
RefCOCO+
CAVE
Method
NSS
AUC
NSS
AUC
NSS
AUC
AnswerMap ( K=8 )
1.36
0.824
1.44
0.817
1.73
0.828
AnswerMap ( {3,5} , equal cost)
1.56
0.853
1.52
0.840
1.70
0.845
Attention (best layer)
−0.22
0.380
−0.18
0.432
−0.21
0.409
Attention (last layer)
−0.15
0.464
−0.11
0.490
−0.17
0.456
Attention rollout
−0.16
0.457
−0.15
0.475
−0.13
0.543
Table 2: Test 1. Does the map land where the model points? NSS and AUC at the model’s own generated point (chance 0 and 0.5). RefCOCOg n=800, RefCOCO+ n=800, CAVE n=334. TAM is answer-conditioned and runs in its own regime (Section 4.1 ). Controls: the probe run on a blank image, and on the image of another example, with the model’s generated point held fixed.
Method
TextVQA
TextVQA (blur)
GQA
GQA (non-binary)
AnswerMap (ours)
53.4
54.0
22.4
26.2
Relevancy T-MM
36.8
36.5
17.1
15.3
Attention (best layer)
18.8
19.3
3.0
2.2
Attention (last layer)
18.3
19.6
4.2
4.9
Attention rollout
9.5
9.5
3.4
5.3
Token activation TAM (own regime)
33.2
32.7
8.8
13.9
Table 3: Test 2. Does deleting the map’s region break the answer? Answer change rate (%) on questions answered correctly with the full image. TextVQA n=400, GQA n=800. Right columns: blur instead of grey, and non-binary questions only, those whose answer is not yes or no. TAM is answer-conditioned and runs in its own regime.
Test 1, agreement
Test 2, deletion
Probe
Best baseline
Model
NSS
AUC
NSS
AUC
Probe
Occlusion
Random
Qwen3-VL-4B
1.36
0.824
occlusion
1.04
0.649
53.4
48.8
8.2
Qwen3-VL-8B
1.49
0.823
attention
0.33
0.634
51.1
50.3
9.4
Qwen3-VL-30B-A3B
1.39
0.813
occlusion
1.03
0.671
45.8
45.0
7.4
InternVL3.5-8B
1.03
0.768
occlusion
1.03
0.699
50.4
50.0
9.6
Table 4: Both tests at 8B, 30B, and on a second family. Test 1: the probe and the strongest baseline run on that model, attention at the model’s own swept best layer. Test 2 on TextVQA: the probe, occlusion, and a uniform random region of the same size, all measured on the same model and examples.
Figure 6: AnswerMap on GPT-6-sol through its API, Query: "a truck number 14 on a snow bank." Logprobs only. Top: an 8×8 grid, 16 calls. Bottom: the multigrid product, 58 calls, narrows the map onto the truck the query names.
Qwen3-VL-4B
Lingshu-7B
Read-out
CAVE
RefCOCOg
RefCOCOg
Generated point (the interface)
77.8
94.5
46.7
Probe expectation ( AnswerMap )
71.3
67.5
57.7
Attention (best layer)
57.2
40.0
40.0
Image-centre prior
58.1
40.5
39.3
Random point
24.3
24.8
24.3
Table 5: Does the map localize where the model’s own pointing fails? Point-in-region accuracy (%) against ground-truth boxes, parse failures scored as misses after getting retried 3 times. Qwen3-VL-4B on CAVE and RefCOCOg, and the medical Lingshu-7B on RefCOCOg.
Yes claim
Map region
Random
Grounded
41.0
1.4
Hallucinated
69.5
11.2
Table 6: Does deletion tell a grounded yes from a hallucinated one? Flip rate (%) of yes claims on POPE when the map’s region or a random region is deleted.
Yes claim
Map region
Random
Grounded
41.0
1.4
Hallucinated
69.5
11.2
Table 6: Does deletion tell a grounded yes from a hallucinated one? Flip rate (%) of yes claims on POPE when the map’s region or a random region is deleted.
Added crop
Accuracy
None (image alone)
91.6
Own map region
92.4
Random region
89.0
Table 7: Does a crop of the map’s region help the model? Accuracy (%) on TextVQA, n=500, Qwen3-VL-4B.
Axis
Configuration
Queries
NSS
Refine one grid
K=2
4
0.75
K=4
8
1.25
K=6
12
1.46
K=8 (default)
16
1.14
K=17
34
1.11
K=40
80
0.22
Table 8: How should a query budget be spent? NSS at the model’s own point on a fixed 296-example RefCOCOg split, Qwen3-VL-4B.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Map expectation against the model’s generated point on RefCOCOg. The probe tracks the diagonal, best-layer attention is a flat band regardless of where the model points.
Figure 8: Faithfulness against query budget. Refining one grid (left) peaks narrowly and collapses, composing coprime grids (right) is flat-topped. Grey markers, the fusion and independence controls of Table 8 .
Method
Queries
Access
4B
30B
Probe map (ours)
16
logits only
0.41
1.25
Occlusion
65
logits only
1.80
5.90
Attention (best layer)
1
internals, eager
0.07
0.38
Attention rollout
1
internals, eager
0.59
1.15
Relevancy T-MM
1
internals, eager, gradients
0.18
1.35
Appendix
Table 9: Seconds per explanation map on one A100, 20 samples at image side 512. Bold: the faster of the two black-box methods. T-MM additionally carries gradient memory, up to 47% extra at 4B.
Figure 9: Qualitative examples of the probe map. Each column shows one example, with the input image above its corresponding probe map.
Figure 10: Agreement of each decoder layer’s attention map with the model’s point on Qwen3-VL-4B. A mid-stack band peaks at layer 15, the last layer is negative.
Model
Image used
Image unused
Random (unused)
Qwen3-VL-4B
57.4
22.0
2.4
Qwen3-VL-8B
53.8
31.1
2.2
Qwen3-VL-30B-A3B
48.4
22.2
0.0
InternVL3.5-8B
54.5
20.6
5.9
GQA non-binary, 4B
30.0
18.7
0.0
Appendix
Table 10: The map matters where the image mattered. Probe-region flip rate (%) on questions the image was needed for against questions it was not, labeled by whole-image deletion. TextVQA unless noted.
Method
NSS
AUC
Dist. ↓
Pearson r
Probe map ( K=8 )
1.36
0.824
0.130
0.665
Probe map ( {3,5} )
1.56
0.853
0.122
0.711
Attention (best layer, ℓ=15 )
−0.22
0.380
0.193
0.224
Attention (last layer)
−0.15
0.464
0.193
−0.036
Attention rollout
−0.16
0.457
0.292
0.120
Relevancy T-MM (all layers)
0.33
0.656
0.159
0.706
Appendix
Table 11: Full agreement detail on RefCOCOg, all four metrics. Bold: best per column among the methods.
Configuration
Low tercile
High tercile
TextVQA, 4B, blank
44.3
64.2
TextVQA, 4B, blur
49.2
61.8
GQA, 4B
16.1
28.4
TextVQA, 30B
26.9
58.5
TextVQA, InternVL3.5-8B
48.3
60.7
Appendix
Table 12: Map-maximum grading. Probe-region flip rate (%) on the lowest and highest tercile of the map maximum, per configuration.
Probe image long side
512 px
Answering and deletion long side
1024 px
Self-conditioning long side
1280 px (the crop inherits it)
Grid
K=8 default, multigrid product {3,5} or {2,3,5}
Multigrid shared grid
64×64 , nearest-neighbour upsampling
Band presentation
one concatenated strip image per question
Deletion region
top-mass cells, 12% of the grid (8 of 64)
Appendix
Table 13: Per-experiment configuration. All values are fixed across models and benchmarks unless a sweep is the experiment.
Understanding how vision-language models (VLMs) interpret data visualizations remains an open problem, and is increasingly important as these models are used for analytical tasks where reliable reasoning is essential. We introduce a lightweight, diagnostic saliency map method tailored for text generation over images using transformer models, the current state-of-the-art models in visualization interpretation. Our approach aggregates the language model's attention over the visual tokens across all heads and layers, then maps this attention back onto the vision encoder's patch grid to localise it over the image, producing a direct correspondence between each generated answer token and the image regions it attended to. This yields fast, gradient-free saliency maps that expose how VLMs allocate focus across visual elements during answer generation, enabling inspection of whether model attention aligns with semantically relevant components. We evaluate our approach using a deletion metric which validates the causal faithfulness of our saliency maps to the model's behavior.
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterfactual image pairs isolate changes in the visual evidence needed to answer. Across four models, direct answers, hidden-state readouts, and state interventions show that the individual judgments can be made without explicit reasoning and that intervening on the corresponding states can affect the answer. During reasoning, the Composite answer becomes decodable from hidden states and usable from shortened traces, often before the model stops on its own. We train a small detector to predict this readiness and stop reasoning at that point. On MMStar and RealWorldQA, this reduces mean reasoning tokens by 79.1% and 74.5%, while average accuracy rises by 3.13 and 3.30 percentage points, respectively. These findings connect the internal development of answer readiness to a practical rule for allocating reasoning computation.
Rong Yu Xu, Prayag Tiwari, Shaolei Zhang
Shenzhen College of International Education · Halmstad University, Sweden · Renmin University of China
Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on localized patch-based embeddings that are insufficient to extract semantics in multi-step reasoning. We propose "Decompose, Look, and Reason" (DLR), a reinforced latent reasoning framework that dynamically decomposes queries into textual premises, extracts premise-conditioned continuous visual latents, and deduces answers through grounded rationales. We introduce a three-stage training pipeline and propose a novel Spherical Gaussian Latent Policy, to enable effective exploration in the latent space. Extensive experiments on vision-centric benchmarks show that DLR consistently outperforms strong baselines, including text-only, interleaved multimodal CoT, and latent reasoning methods, while providing superior stepwise interpretability.
Mengdan Zhu, Senhao Cheng, Liang Zhao
Emory University · University of Michigan, Ann Arbor