AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
Authors: Mohamed Eltahir, Fardows Adam, Duaa M. Tahir, Lama Alamoudi, Sana Ammar, Atheer A. Alboloshi, Jory Albluey, Tanveer Hussain, +1 more
Organizations: King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia · Department of Computer Science, Edge Hill University, Ormskirk, England
When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free, task-agnostic, black-box visual rationale constructed from the output head. The image is cut into K row and K column bands, each shown alone to the frozen model along with the query in the format of a yes/no relevance question. The outer product of the row and column ``yes'' posteriors gives the query-conditioned spatial map. Crucially, by defining a fixed read-out R (e.g., expectation, maximum) on top of AnswerMap, we can derive continuous outputs like location natively. This bypasses the reliance on discrete text tokens for continuous-output tasks and guarantees an image-dependent answer by construction. However, a rationale can be confabulated, so we validate AnswerMap across four models and three query distributions with two tests: (a) agreement with the model's own generated point and (b) deletion of the map's region. The map lands where the model points (AUC 0.85 against 0.38 for attention), and deleting its region flips 53% of correct answers (against 19% for attention's). Beyond establishing faithfulness, we demonstrate the map's task-agnostic utility through three distinct read-outs: its maximum flags hallucinated objects without generation, its expectation localizes correctly when the model's own pointing fails, and its top-mass region, fed back as a crop, fixes half of the model's wrong answers. AnswerMap thus offers a new lens on VLM interpretability and, through its read-outs, a new output interface for visual tasks beyond text tokens.
Figures & tables
Figure 1: Three ways to explain one frozen VLM’s answer. A text rationale is in a mismatched modality. An internal read-out is in the right modality but originates too early and needs the weights. AnswerMap reads the output head one band at a time, the map is visual and black-box, and a fixed read-out computes the answer from it.
Figure 2: AnswerMap (multigrid) from GPT-6-sol through its API, one query each.
Rationale
Spatial
Any answer
Cell score
Black box
Passes per map
Words (text generation)
✗
✓
none
✓
1
Attention
✓
✓
routing
✗
1
Gradient relevancy (e.g., T-MM)
✓
✓
routing
✗
1
Token activation (e.g., TAM)
✓
✗
activation
✗
1
Occlusion
✓
✓
necessity
✓
K2
AnswerMap
✓
✓
sufficiency
✓
2K
Table 1: Positioning. Any answer: the rationale exists for non-word answers such as coordinates. Cell score: what a cell’s value measures (where attention routed, how strongly a word activates, whether the cell is necessary when removed from the full image, or whether a band shown alone suffices). Black box: needs only the model’s output token probabilities. Passes: forward passes for one K×K map.
Figure 3: The AnswerMap operator. K row and K column band questions fill a K×K map through the outer product, and a read-out computes the answer from the map.
Figure 4: Multigrid product. Coprime grids are probed separately, upsampled to a shared grid, and combined, so a region survives only if every band width endorses it.
Figure 5: The two-test faithfulness protocol. Left, agreement, score the full map at the model’s own generated point. Right, deletion, remove the map’s top-mass region against a matched random region and re-ask.
RefCOCOg
RefCOCO+
CAVE
Method
NSS
AUC
NSS
AUC
NSS
AUC
AnswerMap ( K=8 )
1.36
0.824
1.44
0.817
1.73
0.828
AnswerMap ( {3,5} , equal cost)
1.56
0.853
1.52
0.840
1.70
0.845
Attention (best layer)
−0.22
0.380
−0.18
0.432
−0.21
0.409
Attention (last layer)
−0.15
0.464
−0.11
0.490
−0.17
0.456
Attention rollout
−0.16
0.457
−0.15
0.475
−0.13
0.543
Table 2: Test 1. Does the map land where the model points? NSS and AUC at the model’s own generated point (chance 0 and 0.5). RefCOCOg n=800, RefCOCO+ n=800, CAVE n=334. TAM is answer-conditioned and runs in its own regime (Section 4.1 ). Controls: the probe run on a blank image, and on the image of another example, with the model’s generated point held fixed.
Method
TextVQA
TextVQA (blur)
GQA
GQA (non-binary)
AnswerMap (ours)
53.4
54.0
22.4
26.2
Relevancy T-MM
36.8
36.5
17.1
15.3
Attention (best layer)
18.8
19.3
3.0
2.2
Attention (last layer)
18.3
19.6
4.2
4.9
Attention rollout
9.5
9.5
3.4
5.3
Token activation TAM (own regime)
33.2
32.7
8.8
13.9
Table 3: Test 2. Does deleting the map’s region break the answer? Answer change rate (%) on questions answered correctly with the full image. TextVQA n=400, GQA n=800. Right columns: blur instead of grey, and non-binary questions only, those whose answer is not yes or no. TAM is answer-conditioned and runs in its own regime.
Test 1, agreement
Test 2, deletion
Probe
Best baseline
Model
NSS
AUC
NSS
AUC
Probe
Occlusion
Random
Qwen3-VL-4B
1.36
0.824
occlusion
1.04
0.649
53.4
48.8
8.2
Qwen3-VL-8B
1.49
0.823
attention
0.33
0.634
51.1
50.3
9.4
Qwen3-VL-30B-A3B
1.39
0.813
occlusion
1.03
0.671
45.8
45.0
7.4
InternVL3.5-8B
1.03
0.768
occlusion
1.03
0.699
50.4
50.0
9.6
Table 4: Both tests at 8B, 30B, and on a second family. Test 1: the probe and the strongest baseline run on that model, attention at the model’s own swept best layer. Test 2 on TextVQA: the probe, occlusion, and a uniform random region of the same size, all measured on the same model and examples.
Figure 6: AnswerMap on GPT-6-sol through its API, Query: "a truck number 14 on a snow bank." Logprobs only. Top: an 8×8 grid, 16 calls. Bottom: the multigrid product, 58 calls, narrows the map onto the truck the query names.
Qwen3-VL-4B
Lingshu-7B
Read-out
CAVE
RefCOCOg
RefCOCOg
Generated point (the interface)
77.8
94.5
46.7
Probe expectation ( AnswerMap )
71.3
67.5
57.7
Attention (best layer)
57.2
40.0
40.0
Image-centre prior
58.1
40.5
39.3
Random point
24.3
24.8
24.3
Table 5: Does the map localize where the model’s own pointing fails? Point-in-region accuracy (%) against ground-truth boxes, parse failures scored as misses after getting retried 3 times. Qwen3-VL-4B on CAVE and RefCOCOg, and the medical Lingshu-7B on RefCOCOg.
Yes claim
Map region
Random
Grounded
41.0
1.4
Hallucinated
69.5
11.2
Table 6: Does deletion tell a grounded yes from a hallucinated one? Flip rate (%) of yes claims on POPE when the map’s region or a random region is deleted.
Yes claim
Map region
Random
Grounded
41.0
1.4
Hallucinated
69.5
11.2
Table 6: Does deletion tell a grounded yes from a hallucinated one? Flip rate (%) of yes claims on POPE when the map’s region or a random region is deleted.
Added crop
Accuracy
None (image alone)
91.6
Own map region
92.4
Random region
89.0
Table 7: Does a crop of the map’s region help the model? Accuracy (%) on TextVQA, n=500, Qwen3-VL-4B.
Axis
Configuration
Queries
NSS
Refine one grid
K=2
4
0.75
K=4
8
1.25
K=6
12
1.46
K=8 (default)
16
1.14
K=17
34
1.11
K=40
80
0.22
Table 8: How should a query budget be spent? NSS at the model’s own point on a fixed 296-example RefCOCOg split, Qwen3-VL-4B.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Map expectation against the model’s generated point on RefCOCOg. The probe tracks the diagonal, best-layer attention is a flat band regardless of where the model points.
Figure 8: Faithfulness against query budget. Refining one grid (left) peaks narrowly and collapses, composing coprime grids (right) is flat-topped. Grey markers, the fusion and independence controls of Table 8 .
Method
Queries
Access
4B
30B
Probe map (ours)
16
logits only
0.41
1.25
Occlusion
65
logits only
1.80
5.90
Attention (best layer)
1
internals, eager
0.07
0.38
Attention rollout
1
internals, eager
0.59
1.15
Relevancy T-MM
1
internals, eager, gradients
0.18
1.35
Appendix
Table 9: Seconds per explanation map on one A100, 20 samples at image side 512. Bold: the faster of the two black-box methods. T-MM additionally carries gradient memory, up to 47% extra at 4B.
Figure 9: Qualitative examples of the probe map. Each column shows one example, with the input image above its corresponding probe map.
Figure 10: Agreement of each decoder layer’s attention map with the model’s point on Qwen3-VL-4B. A mid-stack band peaks at layer 15, the last layer is negative.
Model
Image used
Image unused
Random (unused)
Qwen3-VL-4B
57.4
22.0
2.4
Qwen3-VL-8B
53.8
31.1
2.2
Qwen3-VL-30B-A3B
48.4
22.2
0.0
InternVL3.5-8B
54.5
20.6
5.9
GQA non-binary, 4B
30.0
18.7
0.0
Appendix
Table 10: The map matters where the image mattered. Probe-region flip rate (%) on questions the image was needed for against questions it was not, labeled by whole-image deletion. TextVQA unless noted.
Method
NSS
AUC
Dist. ↓
Pearson r
Probe map ( K=8 )
1.36
0.824
0.130
0.665
Probe map ( {3,5} )
1.56
0.853
0.122
0.711
Attention (best layer, ℓ=15 )
−0.22
0.380
0.193
0.224
Attention (last layer)
−0.15
0.464
0.193
−0.036
Attention rollout
−0.16
0.457
0.292
0.120
Relevancy T-MM (all layers)
0.33
0.656
0.159
0.706
Appendix
Table 11: Full agreement detail on RefCOCOg, all four metrics. Bold: best per column among the methods.
Configuration
Low tercile
High tercile
TextVQA, 4B, blank
44.3
64.2
TextVQA, 4B, blur
49.2
61.8
GQA, 4B
16.1
28.4
TextVQA, 30B
26.9
58.5
TextVQA, InternVL3.5-8B
48.3
60.7
Appendix
Table 12: Map-maximum grading. Probe-region flip rate (%) on the lowest and highest tercile of the map maximum, per configuration.
Probe image long side
512 px
Answering and deletion long side
1024 px
Self-conditioning long side
1280 px (the crop inherits it)
Grid
K=8 default, multigrid product {3,5} or {2,3,5}
Multigrid shared grid
64×64 , nearest-neighbour upsampling
Band presentation
one concatenated strip image per question
Deletion region
top-mass cells, 12% of the grid (8 of 64)
Appendix
Table 13: Per-experiment configuration. All values are fixed across models and benchmarks unless a sweep is the experiment.