Composition, Not Conversation: VLMs Lose the Scene, Not the Thread
Organizations: NIT Silchar · MBZUAI · BML Munjal University · KAUST
Abstract
Vision-language models (VLMs) increasingly reason over visual evidence that is cropped, segmented, retrieved, or revealed over time. Yet most VQA benchmarks present the complete image and question at once. We ask what models lose when the same information is fragmented. We introduce Layered-VQA, with 93 scenes and 300 questions. Each image is decomposed into ordered RGBA layers that exactly recompose the original scene, and each question is annotated with supporting, minimal-sufficient, and distractor layers. We evaluate eleven open-weight VLMs from 3B to 32B parameters and two proprietary models with a scale of 187,200 conversations, graded by 1.74M open-model cross-judgments. We find three consistent failures. Loss in Composition: fragmenting the question has a small effect, but fragmenting the scene substantially reduces accuracy; recomposing the same layers largely restores performance. Oracle Inversion: even oracle-selected sufficient evidence can perform worse than the complete scene. Loss in Grounding: as more evidence is required, grounding degrades much faster than answer accuracy. Together, these results show that having the right visual evidence is not enough. How that evidence is composed and presented determines whether models can use and ground it. The right evidence is not enough: VLMs need the scene it came from.
Figures & tables
| base | question in text | question in layers | alignment | order | retained % | ||||||||||||
| Model | full | concat | sharded | recap | snowball | concat | sharded | recap | snowball | aligned | misaligned | ev-first | ev-last | shuffled | Text | Visual | Layered |
| Qwen2.5-VL-3B | .408 .02 | .334 .01 | .142 .01 | .439 .01 | .213 .05 | .250 .02 | .150 .02 | .373 .01 | .224 .03 | .201 .03 | .170 .01 | .097 .02 | .193 .03 | .333 .03 | 69 | 61 | 41 |
| Ministral-3-8B | .463 .03 | .433 .02 | .199 .04 | .402 .01 | .200 .03 | .247 .02 | .180 .01 | .523 .03 | .272 .01 | .143 .02 | .176 .01 | .113 .02 | .298 .02 | .454 .03 | 67 | 66 | 39 |
| Qwen2.5-VL-7B | .556 .02 | .496 .02 | .439 .03 | .672 .01 | .497 .01 | .340 .02 | .208 .03 | .649 .03 | .280 .01 | .309 .01 | .321 .03 | .114 .01 | .329 .01 | .478 .04 | 95 | 66 | 48 |
| Qwen2.5-VL-32B | .559 .01 | .549 .03 | .533 .00 | .711 .04 | .630 .03 | .426 .00 | .327 .02 | .708 .03 | .433 .00 | .337 .02 | .367 .02 | .274 .04 | .432 .02 | .557 .02 | 108 | 85 | 63 |
| Pixtral-12B | .572 .01 | .556 .01 | .393 .02 | .577 .02 | .471 .04 | .360 .01 | .233 .02 | .579 .01 | .392 .03 | .271 .01 | .299 .00 | .154 .01 | .271 .02 | .502 .01 | 87 | 68 | 43 |
| answer accuracy | evidence selection | error modes | joint | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | rgb | all-lyr | oracle | Precision | Recall | F1-score | Ex Match | Comp. | Exc. | Dist. | Inv. | Acc. |
| Q2.5-VL-3B | .408 .02 | .371 .03 | .303 .01 | .561 .01 | .522 .01 | .503 .01 | .101 .01 | .279 .00 | .387 .02 | .520 .02 | .137 .01 | .048 .01 |
| Min-3-8B | .463 .03 | .420 .02 | .267 .02 | .617 .01 | .565 .01 | .562 .01 | .174 .00 | .295 .01 | .318 .01 | .409 .02 | .159 .02 | .080 .01 |
| Q2.5-VL-7B | .556 .02 | .463 .02 | .372 .01 | .469 .01 | .364 .02 | .385 .02 | .105 .01 | .154 .02 | .395 .02 | .327 .02 | .251 .01 | .059 .01 |
| Q2.5-VL-32B | .559 .01 | .629 .01 | .474 .02 | .519 .01 | .437 .02 | .447 .01 | .126 .02 | .205 .03 | .455 .01 | .314 .02 | .413 .03 | .083 .02 |
| Pixtral-12B | .572 .01 | .478 .02 | .396 .03 | .649 .03 | .494 .01 | .531 .02 | .168 .02 | .230 .01 | .324 .03 | .383 .02 | .092 .02 | .087 .02 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Question type | full | concat-q | sharded-q | concat-v | sharded-v | recap-v | all-lyr | oracle | |
|---|---|---|---|---|---|---|---|---|---|
| Direct ask | 83 | .663 | .632 | .513 | .538 | .376 | .693 | .650 | .573 |
| Attribute binding | 65 | .601 | .565 | .434 | .507 | .398 | .682 | .604 | .549 |
| Occlusion | 17 | .594 | .590 | .412 | .346 | .248 | .624 | .560 | .443 |
| Counting | 44 | .642 | .539 | .492 | .298 | .235 | .641 | .604 | .342 |
| Spatial | 54 | .524 | .513 | .527 | .280 | .252 | .658 | .496 | .305 |
| Multi-hop | 37 | .550 | .495 | .293 | .372 | .282 | .615 | .546 | .394 |
| first answer attempt | commits | |||||
| Model | earliest | early | midway | late | latest | share |
| 0–20% | 20–40% | 40–60% | 60–80% | 80–100% | ||
| Q2.5-VL-3B | .206 | .220 | .232 | .261 | .203 | 51% |
| Q3-VL-4B | .417 | .561 | .545 | .514 | .500 | 63% |
| Q2.5-VL-7B | .332 | .393 | .398 | .390 | .333 | 53% |
| Q3-VL-8B | .502 | .626 | .657 | .646 | .574 | 45% |
| evidence in one turn | evidence across turns | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | shortest | short | median | long | longest | shortest | short | median | long | longest |
| Q2.5-VL-3B | .311 | .365 | .349 | .369 | .253 | .270 | .249 | .227 | .236 | .218 |
| Q3-VL-4B | .649 | .706 | .715 | .647 | .451 | .592 | .569 | .564 | .534 | .474 |
| Q2.5-VL-7B | .478 | .565 | .554 | .403 | .336 | .412 | .412 | .415 | .381 | .327 |
| Q3-VL-8B | .696 | .676 | .732 | .694 | .521 | .649 | .657 | .623 | .622 | .557 |
| IVL3.5-8B | .524 | .468 | .512 | .546 | .407 | .458 | .431 | .385 | .426 | .448 |
| ungrounded share | accuracy | ||||
|---|---|---|---|---|---|
| Model | Visual | Layered | Order | as scored | grounded only |
| Q2.5-VL-3B | .603 | .556 | .606 | .223 | .090 |
| Q3-VL-4B | .397 | .369 | .274 | .482 | .312 |
| Q2.5-VL-7B | .454 | .404 | .376 | .338 | .196 |
| Q3-VL-8B | .326 | .291 | .221 | .573 | .411 |
| IVL3.5-8B | .478 | .260 | .377 | .403 | .243 |
| Layers | Questions | full | concat-v | oracle |
|---|---|---|---|---|
| 2 | 11 | .606 | ||
| 3 | 100 | .597 | ||
| 4 | 110 | .611 | ||
| 5 | 51 | .594 | ||
| 6 | 28 | .616 |
| Condition | Relational | Single-object | Extra loss | CI |
|---|---|---|---|---|
| concat-v | ||||
| oracle |
| Shard | Text | Supported by |
|---|---|---|
| 1 | can you help me identify a particular thing in this image? | — |
| 2 | The answer should name an object. | 2 |
| 3 | The target object is positioned to the right of the man. | 2, 3 |
| 4 | The referenced man is taking a selfie. | 3 |
| 5 | The target object is directly in front of the worker. | 1, 2 |
| 6 | The referenced worker is wearing a white helmet and an orange-red | 1 |
| Condition | Turns | What the model receives |
| The question arrives in parts, the scene stays whole | ||
| full | 1 | with |
| concat-q | 1 | with |
| sharded-q | 6 | with ; then , , , , , one per turn |
| recap-q | 7 | As sharded-q , then a final turn repeating with |
| snowball-q | 6 | with ; then ; then ; and so on |
| 2 1 + bg | BG background L1 barn | DIRECT ASK 2 layers 5 shards What natural landscape is visible rising behind the barns in the distance? Answer A mountain range. 1 can you help me figure out what is being identified in this image? 2 The answer should name a natural landscape. BG 3 The target is the natural landscape that is visible and rising. 4 It is behind the barns. 5 The barns are in the distance. |
| 3 2 + bg | BG background L1 bird L2 bird | SPATIAL 3 layers 6 shards Where is the brown hawk positioned relative to the white bird with the yellow-and-black beak on the rocky shoreline? Answer The brown hawk is to the left of the white bird. 1 can you help me figure out a spatial property of a particular thing? 2 The answer should report where the thing is located. 3 The target thing is the brown hawk. L2 4 Determine the brown hawk’s position relative to a comparison entity. L1 5 Use the white bird with the yellow-and-black beak as the comparison entity. 6 The white bird with the yellow-and-black beak is on the rocky shoreline. BG |
| 4 3 + bg | BG background L1 dog L2 dog L3 dog | COUNTING 4 layers 6 shards How many dogs in the fenced dirt area are facing generally toward the camera or frontward, and how many are turned backward or away? Answer Two dogs are facing frontward, and one dog is facing backward. 1 can you help me figure out the counts for two groups of subjects? L1 L2 L3 2 Each requested result is a count. 3 The subjects to count are dogs. 4 The dogs are in the fenced dirt area. BG 5 Count the dogs that are facing generally toward the camera or frontward. 6 Include a count of the dogs that are turned backward or away. |
| 5 4 + bg | BG background L1 symbol L2 woman L3 toy L4 man | MULTI-HOP REFERENCE 5 layers 5 shards What object is being held by the person standing to the right of the man who is pointing toward her? Answer A teddy bear. 1 can you help me identify a particular thing in this image? 2 The answer should name an object. L3 3 The target object is being held by a person. L2 4 The person holding the target object is standing to the right of a man. L4 5 The referenced man is pointing toward the person holding the target object. |
| 6 5 + bg | BG background L1 lion L2 lion L3 lion L4 lion L5 lion | COUNTING 6 layers 5 shards How many lions are running across the grassy field in this scene? Answer There are five lions visible in the field… 1 can you help me figure out a quantity for particular animals? L1 L2 L3 L4 L5 2 The requested quantity is a count. 3 The animals to count are lions. 4 Count lions that are running across the grassy field. BG 5 Restrict the count to the lions in this scene. |