Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterfactual image pairs isolate changes in the visual evidence needed to answer. Across four models, direct answers, hidden-state readouts, and state interventions show that the individual judgments can be made without explicit reasoning and that intervening on the corresponding states can affect the answer. During reasoning, the Composite answer becomes decodable from hidden states and usable from shortened traces, often before the model stops on its own. We train a small detector to predict this readiness and stop reasoning at that point. On MMStar and RealWorldQA, this reduces mean reasoning tokens by 79.1% and 74.5%, while average accuracy rises by 3.13 and 3.30 percentage points, respectively. These findings connect the internal development of answer readiness to a practical rule for allocating reasoning computation.
Figures & tables
Figure 1: Overview of our study. Controlled visual tasks test individual judgments and their combination in Composite scenes. Linear readouts and counterfactual state interventions examine what visual information is represented and affects the answer. During native reasoning, readouts and answers from shortened traces track when the Composite answer becomes available. A hidden-state readiness detector uses this principle to stop reasoning early on MMStar and RealWorldQA.
Figure 2: Every column shows a counterfactual pair in which the task-relevant variable and the correct answer change together while the remaining scene content is preserved.
Task
Qwen3.5-4B
Qwen3.5-9B
Gemma-4-E4B
Gemma-4-12B
Guess
Feature Binding
98.0
99.8
98.6
99.9
25.0
Numerosity
79.6
85.2
72.4
86.5
50.0
Spatial Relations
99.8
100.0
97.5
89.5
25.0
Amodal Completion
76.8
80.5
81.4
72.9
25.0
Avg
88.6
91.4
87.5
87.2
31.3
Table 1: All four component judgments can be performed above the guess level without explicit reasoning. Answer accuracy (%) on 1,000 images per task and model; guess is uniform random selection over the answer choices (two for Numerosity, four for the others).
Sample Accuracy
Pair-Consistent Accuracy
Budget hit
Model
Non-thinking
Thinking
Δ
Non-thinking
Thinking
Δ
Count (%)
Qwen3.5-4B
58.5 [53.0,64.0]
80.0 [74.0,86.0]
↑21.5
26.0 [18.0,35.0]
67.0 [58.0,76.0]
↑41.0
38 (19.0%)
Qwen3.5-9B
66.5 [61.5,71.5]
94.0 [90.5,97.0]
↑27.5
36.0 [27.0,45.0]
89.0 [83.0,95.0]
↑53.0
11 (5.5%)
Gemma-4-E4B
52.0 [49.0,55.0]
73.5 [67.0,80.0]
↑21.5
7.0 [2.0,12.0]
57.0 [47.0,67.0]
↑50.0
0 (0.0%)
Gemma-4-12B
57.5 [54.0,61.0]
54.5 [47.0,62.0]
↓3.0
15.0 [8.0,22.0]
34.0 [25.0,43.0]
↑19.0
84 (42.0%)
Avg
58.6
75.5
↑16.9
21.0
61.8
↑40.8
133 (16.6%)
Table 2: Native thinking improves counterfactual consistency. Values are accuracy (%) with pair-bootstrap 95% CIs; Δ is the reasoning-induced change in percentage points. Budget hit counts thinking-mode runs that reach the generation limit, out of 200 per model; the summary row pools all 800 runs.
Accuracy (%)
Reasoning Tokens
Benchmark
Model
Full
Stop
Δ
Full
Stop
Reduction
MMStar
Qwen3.5-4B
64.87
64.93
↑0.07
2,041
255
87.5%
Qwen3.5-9B
66.87
69.87
↑3.00
1,951
255
86.9%
Gemma-4-E4B
58.07
58.67
↑0.60
753
585
22.3%
Gemma-4-12B
56.93
65.80
↑8.87
1,854
286
84.6%
Avg
61.69
64.82
↑3.13
1,650
345
79.1%
Table 3: Adaptive stopping cuts reasoning cost on both benchmarks while raising average accuracy. All held-out cases per model (1,500 MMStar and 765 RealWorldQA); accuracy in %, reasoning tokens as means, Δ in percentage points, reduction as the token decrease.
Total output tokens
Stopping (%)
Model
Full
Stop
Reduction
Any check
First check
MMStar
Qwen3.5-4B
2,159
326
84.9%
95.1%
94.7%
Qwen3.5-9B
2,066
315
84.8%
94.7%
94.3%
Gemma-4-E4B
944
724
23.3%
24.7%
15.9%
Gemma-4-12B
1,974
351
82.2%
76.3%
68.1%
Table 4: Mean total output tokens and detector-triggered stopping (% of all questions).
Missing answer
Full status
Net correct gain
Model
Full
Stop
Unclosed
Closed
Unclosed
Closed
Total
MMStar
Qwen3.5-4B
297
24
267
1,233
96
−95
1
Qwen3.5-9B
276
44
243
1,257
121
−76
45
Gemma-4-E4B
104
87
32
1,468
20
−11
9
Gemma-4-12B
318
56
207
1,293
100
33
133
Table 5: Answer completion and changes in correct-answer counts. Missing includes traces without a parsed answer. Open/closed refer to full thinking; direct answers count as closed. Gains are stopping minus full thinking.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Analysis Split
Dataset
Images
Choices
Train
Validation
Test
Split Unit
Feature Binding
1,000
4
800
–
200
Image
Numerosity
1,000
2
800
–
200
Image
Spatial Relations
1,000
4
800
–
200
Image
Amodal Completion
1,000
4
800
–
200
Pair
Composite
1,000
2
640
160
200
Pair
Appendix
Table 6: Synthetic dataset sizes and analysis splits. Counts are images. Primitive splits are used for linear probing, and the Composite split is used for reasoning-trajectory analysis before trace filtering. Behavioral evaluation uses all 1,000 images per primitive and the 200-image Composite test set. Paired images remain in the same split.
Setting
Specification
Inputs
No-thinking image and question
Visual features
Mean-pooled image-token states at each tested depth
Question-only control
Mean-pooled question span, including candidates, with no image
Standardization
Per-dimension training mean and standard deviation
Label classes
4 for Feature Binding, 2 for Numerosity, 4 for Spatial Relations, and 6 for Amodal Completion
Appendix
Table 7: Primitive linear-readout settings. Dataset splits are given in Table 6 .
Task
Qwen3.5-4B
Qwen3.5-9B
Gemma-4-E4B
Gemma-4-12B
Feature Binding
24
25
22
25
Numerosity
22
23
18
24
Spatial Relations
25
25
22
25
Amodal Completion
48
62
72
75
Appendix
Table 8: Retained counterfactual pairs for primitive interventions.
Checkpoint
State position
Pre
Last prompt token before reasoning
Early
Reasoning token ⌊T/3⌋
Middle
Reasoning token ⌊2T/3⌋
Late
Last reasoning token, T
Answer
Token immediately preceding the first final-choice token
Appendix
Table 9: Reasoning-trajectory checkpoints. T counts reasoning tokens, and positions are one-based.
Source fold
Held-out fold
Dataset
Evaluation
Training
Validation
Test
MMStar
A → B
600
150
750
B → A
600
150
750
RealWorldQA
A → B
306
77
382
B → A
305
77
383
Appendix
Table 10: Early-stopping evaluation splits. Counts are questions. Each source fold is divided into detector training and threshold validation, and the other fold is held out for testing.
Accuracy (%)
Reasoning Tokens
Model
Train → Test
Full
Stop
Δ
Full
Stop
Qwen3.5-4B
A → B
64.13
64.13
0.00
2,101
256
B → A
65.60
65.73
↑0.13
1,982
254
Qwen3.5-9B
A → B
66.93
69.73
↑2.80
1,979
256
B → A
66.80
70.00
↑3.20
1,922
254
Gemma-4-E4B
A → B
57.47
57.47
0.00
767
437
Appendix
Table 11: Per-direction held-out MMStar results. Each row evaluates 750 held-out cases. Accuracy values are percentages and reasoning-token counts are means. Δ denotes the change from full thinking to adaptive stopping in percentage points.
Accuracy (%)
Reasoning Tokens
Model
Train → Test
Full
Stop
Δ
Full
Stop
Qwen3.5-4B
A → B
71.73
74.87
↑3.14
1,692
247
B → A
73.89
75.98
↑2.09
1,700
322
Qwen3.5-9B
A → B
73.30
78.27
↑4.97
1,521
252
B → A
72.58
77.55
↑4.96
1,692
244
Gemma-4-E4B
A → B
57.33
56.81
↓0.52
448
347
Appendix
Table 12: Per-direction held-out RealWorldQA results. A → B evaluates 382 questions, and B → A evaluates 383. Accuracy values are percentages and reasoning-token counts are means. Δ denotes the change from full thinking to adaptive stopping in percentage points.