Composition, Not Conversation: VLMs Lose the Scene, Not the Thread
Authors: L. D. M. S. Sai Teja, Ufaq Khan, N. Siva Gopala Krishna, Satyajit Tourani, Ashshak Sharifdeen, Fida Mohammad Thoker, Bernard Ghanem, Muhammad Haris Khan
Vision-language models (VLMs) increasingly reason over visual evidence that is cropped, segmented, retrieved, or revealed over time. Yet most VQA benchmarks present the complete image and question at once. We ask what models lose when the same information is fragmented. We introduce Layered-VQA, with 93 scenes and 300 questions. Each image is decomposed into ordered RGBA layers that exactly recompose the original scene, and each question is annotated with supporting, minimal-sufficient, and distractor layers. We evaluate eleven open-weight VLMs from 3B to 32B parameters and two proprietary models with a scale of 187,200 conversations, graded by 1.74M open-model cross-judgments. We find three consistent failures. Loss in Composition: fragmenting the question has a small effect, but fragmenting the scene substantially reduces accuracy; recomposing the same layers largely restores performance. Oracle Inversion: even oracle-selected sufficient evidence can perform worse than the complete scene. Loss in Grounding: as more evidence is required, grounding degrades much faster than answer accuracy. Together, these results show that having the right visual evidence is not enough. How that evidence is composed and presented determines whether models can use and ground it. The right evidence is not enough: VLMs need the scene it came from.
Figures & tables
Figure 1: Vision-language models lose the scene, not the thread. Left: Scene fragmentation hurts accuracy, while recomposition recovers it. Middle: Oracle-selected minimal evidence can perform worse than the complete scene. Right: As evidence demand grows, grounding falls much faster than answer accuracy.
Figure 2: Layered-VQA construction. Scenes are decomposed into RGBA layers, manually annotated, expanded into question shards and evidence mappings with GPT-5.6-Luna, and verified before finalization.
Figure 3: Controlled delivery conditions. The sample is presented under different question, visual, alignment, and evidence schedules.
Figure 4: Self-preference after cross-judging. Boxed cells correspond to models judging their own outputs.
base
question in text
question in layers
alignment
order
retained %
Model
full
concat
sharded
recap
snowball
concat
sharded
recap
snowball
aligned
misaligned
ev-first
ev-last
shuffled
Text
Visual
Layered
Qwen2.5-VL-3B
.408 ± .02
.334 ± .01
.142 ± .01
.439 ± .01
.213 ± .05
.250 ± .02
.150 ± .02
.373 ± .01
.224 ± .03
.201 ± .03
.170 ± .01
.097 ± .02
.193 ± .03
.333 ± .03
69
61
41
Ministral-3-8B
.463 ± .03
.433 ± .02
.199 ± .04
.402 ± .01
.200 ± .03
.247 ± .02
.180 ± .01
.523 ± .03
.272 ± .01
.143 ± .02
.176 ± .01
.113 ± .02
.298 ± .02
.454 ± .03
67
66
39
Qwen2.5-VL-7B
.556 ± .02
.496 ± .02
.439 ± .03
.672 ± .01
.497 ± .01
.340 ± .02
.208 ± .03
.649 ± .03
.280 ± .01
.309 ± .01
.321 ± .03
.114 ± .01
.329 ± .01
.478 ± .04
95
66
48
Qwen2.5-VL-32B
.559 ± .01
.549 ± .03
.533 ± .00
.711 ± .04
.630 ± .03
.426 ± .00
.327 ± .02
.708 ± .03
.433 ± .00
.337 ± .02
.367 ± .02
.274 ± .04
.432 ± .02
.557 ± .02
108
85
63
Pixtral-12B
.572 ± .01
.556 ± .01
.393 ± .02
.577 ± .02
.471 ± .04
.360 ± .01
.233 ± .02
.579 ± .01
.392 ± .03
.271 ± .01
.299 ± .00
.154 ± .01
.271 ± .02
.502 ± .01
87
68
43
Table 1: What it costs to deliver the evidence across turns. Consensus final-turn accuracy.
answer accuracy
evidence selection
error modes
joint
Model
rgb
all-lyr
oracle
Precision
Recall
F1-score
Ex Match
Comp.
Exc.
Dist.
Inv.
Acc.
Q2.5-VL-3B
.408 ± .02
.371 ± .03
.303 ± .01
.561 ± .01
.522 ± .01
.503 ± .01
.101 ± .01
.279 ± .00
.387 ± .02
.520 ± .02
.137 ± .01
.048 ± .01
Min-3-8B
.463 ± .03
.420 ± .02
.267 ± .02
.617 ± .01
.565 ± .01
.562 ± .01
.174 ± .00
.295 ± .01
.318 ± .01
.409 ± .02
.159 ± .02
.080 ± .01
Q2.5-VL-7B
.556 ± .02
.463 ± .02
.372 ± .01
.469 ± .01
.364 ± .02
.385 ± .02
.105 ± .01
.154 ± .02
.395 ± .02
.327 ± .02
.251 ± .01
.059 ± .01
Q2.5-VL-32B
.559 ± .01
.629 ± .01
.474 ± .02
.519 ± .01
.437 ± .02
.447 ± .01
.126 ± .02
.205 ± .03
.455 ± .01
.314 ± .02
.413 ± .03
.083 ± .02
Pixtral-12B
.572 ± .01
.478 ± .02
.396 ± .03
.649 ± .03
.494 ± .01
.531 ± .02
.168 ± .02
.230 ± .01
.324 ± .03
.383 ± .02
.092 ± .02
.087 ± .02
Table 2: What models can do when evidence is handed to them.
Figure 5: Layer citation over time.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Question type
n
full
concat-q
sharded-q
concat-v
sharded-v
recap-v
all-lyr
oracle
Direct ask
83
.663
.632
.513
.538
.376
.693
.650
.573
Attribute binding
65
.601
.565
.434
.507
.398
.682
.604
.549
Occlusion
17
.594
.590
.412
.346
.248
.624
.560
.443
Counting
44
.642
.539
.492
.298
.235
.641
.604
.342
Spatial
54
.524
.513
.527
.280
.252
.658
.496
.305
Multi-hop
37
.550
.495
.293
.372
.282
.615
.546
.394
Appendix
Table 3: Accuracy by question type and delivery condition.
Figure 6: Response length and accuracy across successive answer attempts. The dashed line is the single-turn baseline for the same questions. Attempts get terser, and accuracy peaks mid-conversation.
Figure 7: Aptitude against unreliability. Hollow markers deliver the evidence in one turn, filled markers shard it. An arrow joins the two states of one model.
first answer attempt
commits
Model
earliest
early
midway
late
latest
share
0–20%
20–40%
40–60%
60–80%
80–100%
Q2.5-VL-3B
.206
.220
.232
.261
.203
51%
Q3-VL-4B
.417
.561
.545
.514
.500
63%
Q2.5-VL-7B
.332
.393
.398
.390
.333
53%
Q3-VL-8B
.502
.626
.657
.646
.574
45%
Appendix
Table 4: When models first attempt an answer.
evidence in one turn
evidence across turns
Model
shortest
short
median
long
longest
shortest
short
median
long
longest
Q2.5-VL-3B
.311
.365
.349
.369
.253
.270
.249
.227
.236
.218
Q3-VL-4B
.649
.706
.715
.647
.451
.592
.569
.564
.534
.474
Q2.5-VL-7B
.478
.565
.554
.403
.336
.412
.412
.415
.381
.327
Q3-VL-8B
.696
.676
.732
.694
.521
.649
.657
.623
.622
.557
IVL3.5-8B
.524
.468
.512
.546
.407
.458
.431
.385
.426
.448
Appendix
Table 5: Verbosity: Accuracy by evidence length and delivery mode. Models are grouped by the length of the evidence they receive, either in one turn or spread across multiple turns.
ungrounded share
accuracy
Model
Visual
Layered
Order
as scored
grounded only
Q2.5-VL-3B
.603
.556
.606
.223
.090
Q3-VL-4B
.397
.369
.274
.482
.312
Q2.5-VL-7B
.454
.404
.376
.338
.196
Q3-VL-8B
.326
.291
.221
.573
.411
IVL3.5-8B
.478
.260
.377
.403
.243
Appendix
Table 6: Ungrounded answers remain common across delivery settings. Left : reports the share of responses produced before the required evidence was available; Right : compares standard accuracy with accuracy restricted to grounded answers.
Layers
Questions
full
Δ concat-v
Δ oracle
2
11
.606
−.050
−.003
3
100
.597
−.149
−.128
4
110
.611
−.193
−.155
5
51
.594
−.224
−.156
6
28
.616
−.269
−.229
Appendix
Table 7: Loss by scene depth. Both losses grow with the number of layers, and neither depends on a particular granularity.
Condition
Relational
Single-object
Extra loss
95% CI
concat-v
−.259
−.126
−.133
[−.189,−.074]
oracle
−.228
−.081
−.147
[−.202,−.089]
Appendix
Table 8: Relational questions lose more under visual decomposition. Extra loss is the difference between relational and single-object degradation.
Figure 8: The loss is concentrated, and it scales with depth. Left: questions ranked by loss against the running total of that loss. Right: both losses by the number of layers a scene decomposes into, with 95% scene-level bootstrap intervals.
Figure 9: Restoring a fragmented stream overshoots the baseline. Left: the cost of fragmenting each stream and the recovery when it is restored, against full . Right: the two overshoots per model. Points on the diagonal mean the final turn is worth the same either way.
Figure 10: The scene and its four layers. Left to right: the composed scene, then layer 0 (background, not needed), layer 1 (the worker, the reference landmark), layer 2 (the bottle, the answer), and layer 3 (the man taking a selfie, the relational bridge). Every layer render is used exactly as shown.
Shard
Text
Supported by
1
can you help me identify a particular thing in this image?
—
2
The answer should name an object.
2
3
The target object is positioned to the right of the man.
2, 3
4
The referenced man is taking a selfie.
3
5
The target object is directly in front of the worker.
1, 2
6
The referenced worker is wearing a white helmet and an orange-red
1
Appendix
Table 9: The six question shards, and the layers each one depends on. A shard with no supporting layer is asking about the question rather than the image.
Condition
Turns
What the model receives
The question arrives in parts, the scene stays whole
full
1
Q with I
concat-q
1
q1q2q3q4q5q6 with I
sharded-q
6
q1 with I ; then q2 , q3 , q4 , q5 , q6 , one per turn
recap-q
7
As sharded-q , then a final turn repeating Q with I
snowball-q
6
q1 with I ; then q1q2 ; then q1q2q3 ; and so on
Appendix
Table 10: Sixteen delivery schedules for one question.
2 1 + bg
BG background L1 barn
DIRECT ASK 2 layers 5 shards What natural landscape is visible rising behind the barns in the distance? Answer A mountain range. 1 can you help me figure out what is being identified in this image? 2 The answer should name a natural landscape. BG 3 The target is the natural landscape that is visible and rising. 4 It is behind the barns. 5 The barns are in the distance.
3 2 + bg
BG background L1 bird L2 bird
SPATIAL 3 layers 6 shards Where is the brown hawk positioned relative to the white bird with the yellow-and-black beak on the rocky shoreline? Answer The brown hawk is to the left of the white bird. 1 can you help me figure out a spatial property of a particular thing? 2 The answer should report where the thing is located. 3 The target thing is the brown hawk. L2 4 Determine the brown hawk’s position relative to a comparison entity. L1 5 Use the white bird with the yellow-and-black beak as the comparison entity. 6 The white bird with the yellow-and-black beak is on the rocky shoreline. BG
4 3 + bg
BG background L1 dog L2 dog L3 dog
COUNTING 4 layers 6 shards How many dogs in the fenced dirt area are facing generally toward the camera or frontward, and how many are turned backward or away? Answer Two dogs are facing frontward, and one dog is facing backward. 1 can you help me figure out the counts for two groups of subjects? L1 L2 L3 2 Each requested result is a count. 3 The subjects to count are dogs. 4 The dogs are in the fenced dirt area. BG 5 Count the dogs that are facing generally toward the camera or frontward. 6 Include a count of the dogs that are turned backward or away.
5 4 + bg
BG background L1 symbol L2 woman L3 toy L4 man
MULTI-HOP REFERENCE 5 layers 5 shards What object is being held by the person standing to the right of the man who is pointing toward her? Answer A teddy bear. 1 can you help me identify a particular thing in this image? 2 The answer should name an object. L3 3 The target object is being held by a person. L2 4 The person holding the target object is standing to the right of a man. L4 5 The referenced man is pointing toward the person holding the target object.
COUNTING 6 layers 5 shards How many lions are running across the grassy field in this scene? Answer There are five lions visible in the field… 1 can you help me figure out a quantity for particular animals? L1 L2 L3 L4 L5 2 The requested quantity is a count. 3 The animals to count are lions. 4 Count lions that are running across the grassy field. BG 5 Restrict the count to the lions in this scene.
Appendix
Table 11: Scene depth across the benchmark. One scene per depth, from a single object over a background up to five. Depth is a property of the scene and shard count is a property of the question.
Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly reflect reliance on visual evidence. Motivated by a surprising observation that removing a substantial fraction of image tokens only degrades model performance very slightly on a widely used hallucination benchmark, we systematically investigate this mismatch in a set of open-source VLMs. Our analysis spans multiple levels of granularity, spanning global visual degradation, localized occlusion, question reformulation, answer-space expansion, and decision-level analyses beyond standard accuracy. We further complement these behavioral results with a layer-wise analysis of vision-token geometry. Throughout the experiments, we find that although VLMs do incorporate visual input, their predictions are less sensitive to the loss of fine-grained visual evidence that standard accuracy should have suggested. Even when the final prediction remains unchanged, the model's internal support for the correct answer may already be weakened. We further complement a representation-level analysis, which shows increasing similarity among visual tokens in deeper layers, providing a possible explanation for our findings. Together, these results suggest that current benchmarks are not sufficient to reliably evaluate fine-grained visual grounding in VLMs.
Zixuan Lan, Luzhe Sun, Matthew R. Walter +1
University of Chicago · Toyota Technological Institute at Chicago · Stony Brook University
Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separately yet still fail on the original multimodal question, a distinction that overall answer accuracy cannot reveal. To study this, we perform a question-level empirical analysis across multiple VLMs and visual domains. We define an exactly scorable task state (i.e., the visual information sufficient to solve a question) and use it to test whether the same model can extract the required state, solve the question from the ground-truth state, and answer the original multimodal question. We find that composition failures, where extraction and solving both succeed but direct answering fails, account for 17.7% to 75.6% of direct-answering errors across multiple VLMs and datasets. To address this failure mode, we introduce a simple yet effective method, termed State Realization Tuning (SRT). SRT fine-tunes LoRA adapters attached to the language-model layers while keeping the pretrained VLM weights frozen. It trains the model to output the ground-truth task state before the final answer in a single autoregressive response. SRT improves over standard supervised fine-tuning by 1.7 to 14.1 percentage points and repairs 92.5% to 98.1% of diagnosed composition failures. A single LoRA adapter trained with SRT also improves performance across substantially different task-state structures. Our work shows that having both visual extraction and problem-solving capabilities does not guarantee correct multimodal answering. Requiring the model to first output the visual information needed to solve the question can help bridge this gap.
Ziheng Wang, Mingxuan Xie, Yilin Liu +3
Sun Yat-sen University · Zhejiang University · The Hong Kong University of Science and Technology +2
Counterfactual analysis is widely used to study evidence use in vision-language models, but its diagnostic value is limited on well-posed tasks: when several cues independently support the same answer, removing one may not change the prediction. We propose monocular metric object-size estimation as an ill-posed diagnostic setting for evidence selection: because physical size cannot be determined from a single uncalibrated image, models must rely on imperfect cues category priors, target appearance, local context, apparent image size, and scene geometry. We assemble Metric VQA (10,813 dimension queries from Objectron and 331 tape-measured in-the-wild scenes) and evaluate 12 open-weight VLMs (3--397,B parameters) with counterfactual analysis decomposing six visual and language evidence channels. Even the largest VLMs tested (Qwen3-VL-235B, Qwen3.5-397B, InternVL3.5-241B) trail a text-only frontier LLM on the in-the-wild split. The diagnostic analysis shows: target identity is the most load-bearing cue, target pixels and local context help only some models, apparent size shifts predictions without a directional readout, and global scene geometry is largely unused. We analyze LoRA fine-tuning as an actionable intervention specific to metric estimation: while the task is learnable, the models do not learn to leverage scene geometry.