From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models
Organizations: Shenzhen College of International Education · Halmstad University, Sweden · Renmin University of China
Abstract
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterfactual image pairs isolate changes in the visual evidence needed to answer. Across four models, direct answers, hidden-state readouts, and state interventions show that the individual judgments can be made without explicit reasoning and that intervening on the corresponding states can affect the answer. During reasoning, the Composite answer becomes decodable from hidden states and usable from shortened traces, often before the model stops on its own. We train a small detector to predict this readiness and stop reasoning at that point. On MMStar and RealWorldQA, this reduces mean reasoning tokens by 79.1% and 74.5%, while average accuracy rises by 3.13 and 3.30 percentage points, respectively. These findings connect the internal development of answer readiness to a practical rule for allocating reasoning computation.
Figures & tables
| Task | Qwen3.5-4B | Qwen3.5-9B | Gemma-4-E4B | Gemma-4-12B | Guess |
| Feature Binding | 98.0 | 99.8 | 98.6 | 99.9 | 25.0 |
| Numerosity | 79.6 | 85.2 | 72.4 | 86.5 | 50.0 |
| Spatial Relations | 99.8 | 100.0 | 97.5 | 89.5 | 25.0 |
| Amodal Completion | 76.8 | 80.5 | 81.4 | 72.9 | 25.0 |
| Avg | 88.6 | 91.4 | 87.5 | 87.2 | 31.3 |
| Sample Accuracy | Pair-Consistent Accuracy | Budget hit | |||||
| Model | Non-thinking | Thinking | Non-thinking | Thinking | Count (%) | ||
| Qwen3.5-4B | 58.5 [53.0,64.0] | 80.0 [74.0,86.0] | 26.0 [18.0,35.0] | 67.0 [58.0,76.0] | 38 (19.0%) | ||
| Qwen3.5-9B | 66.5 [61.5,71.5] | 94.0 [90.5,97.0] | 36.0 [27.0,45.0] | 89.0 [83.0,95.0] | 11 (5.5%) | ||
| Gemma-4-E4B | 52.0 [49.0,55.0] | 73.5 [67.0,80.0] | 7.0 [2.0,12.0] | 57.0 [47.0,67.0] | 0 (0.0%) | ||
| Gemma-4-12B | 57.5 [54.0,61.0] | 54.5 [47.0,62.0] | 15.0 [8.0,22.0] | 34.0 [25.0,43.0] | 84 (42.0%) | ||
| Avg | 58.6 | 75.5 | 21.0 | 61.8 | 133 (16.6%) | ||
| Accuracy (%) | Reasoning Tokens | ||||||
| Benchmark | Model | Full | Stop | Full | Stop | Reduction | |
| MMStar | Qwen3.5-4B | 64.87 | 64.93 | 2,041 | 255 | ||
| Qwen3.5-9B | 66.87 | 69.87 | 1,951 | 255 | |||
| Gemma-4-E4B | 58.07 | 58.67 | 753 | 585 | |||
| Gemma-4-12B | 56.93 | 65.80 | 1,854 | 286 | |||
| Avg | 61.69 | 64.82 | 1,650 | 345 | |||
| Total output tokens | Stopping (%) | ||||
| Model | Full | Stop | Reduction | Any check | First check |
| MMStar | |||||
| Qwen3.5-4B | 2,159 | 326 | 84.9% | 95.1% | 94.7% |
| Qwen3.5-9B | 2,066 | 315 | 84.8% | 94.7% | 94.3% |
| Gemma-4-E4B | 944 | 724 | 23.3% | 24.7% | 15.9% |
| Gemma-4-12B | 1,974 | 351 | 82.2% | 76.3% | 68.1% |
| Missing answer | Full status | Net correct gain | |||||
| Model | Full | Stop | Unclosed | Closed | Unclosed | Closed | Total |
| MMStar | |||||||
| Qwen3.5-4B | 297 | 24 | 267 | 1,233 | 96 | 1 | |
| Qwen3.5-9B | 276 | 44 | 243 | 1,257 | 121 | 45 | |
| Gemma-4-E4B | 104 | 87 | 32 | 1,468 | 20 | 9 | |
| Gemma-4-12B | 318 | 56 | 207 | 1,293 | 100 | 33 | 133 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Analysis Split | ||||||
| Dataset | Images | Choices | Train | Validation | Test | Split Unit |
| Feature Binding | 1,000 | 4 | 800 | – | 200 | Image |
| Numerosity | 1,000 | 2 | 800 | – | 200 | Image |
| Spatial Relations | 1,000 | 4 | 800 | – | 200 | Image |
| Amodal Completion | 1,000 | 4 | 800 | – | 200 | Pair |
| Composite | 1,000 | 2 | 640 | 160 | 200 | Pair |
| Setting | Specification |
| Inputs | No-thinking image and question |
| Visual features | Mean-pooled image-token states at each tested depth |
| Question-only control | Mean-pooled question span, including candidates, with no image |
| Standardization | Per-dimension training mean and standard deviation |
| Label classes | 4 for Feature Binding, 2 for Numerosity, 4 for Spatial Relations, and 6 for Amodal Completion |
| Task | Qwen3.5-4B | Qwen3.5-9B | Gemma-4-E4B | Gemma-4-12B |
| Feature Binding | 24 | 25 | 22 | 25 |
| Numerosity | 22 | 23 | 18 | 24 |
| Spatial Relations | 25 | 25 | 22 | 25 |
| Amodal Completion | 48 | 62 | 72 | 75 |
| Checkpoint | State position |
| Pre | Last prompt token before reasoning |
| Early | Reasoning token |
| Middle | Reasoning token |
| Late | Last reasoning token, |
| Answer | Token immediately preceding the first final-choice token |
| Source fold | Held-out fold | |||
| Dataset | Evaluation | Training | Validation | Test |
| MMStar | A B | 600 | 150 | 750 |
| B A | 600 | 150 | 750 | |
| RealWorldQA | A B | 306 | 77 | 382 |
| B A | 305 | 77 | 383 | |
| Accuracy (%) | Reasoning Tokens | |||||
| Model | Train Test | Full | Stop | Full | Stop | |
| Qwen3.5-4B | A B | 64.13 | 64.13 | 2,101 | 256 | |
| B A | 65.60 | 65.73 | 1,982 | 254 | ||
| Qwen3.5-9B | A B | 66.93 | 69.73 | 1,979 | 256 | |
| B A | 66.80 | 70.00 | 1,922 | 254 | ||
| Gemma-4-E4B | A B | 57.47 | 57.47 | 767 | 437 | |
| Accuracy (%) | Reasoning Tokens | |||||
| Model | Train Test | Full | Stop | Full | Stop | |
| Qwen3.5-4B | A B | 71.73 | 74.87 | 1,692 | 247 | |
| B A | 73.89 | 75.98 | 1,700 | 322 | ||
| Qwen3.5-9B | A B | 73.30 | 78.27 | 1,521 | 252 | |
| B A | 72.58 | 77.55 | 1,692 | 244 | ||
| Gemma-4-E4B | A B | 57.33 | 56.81 | 448 | 347 | |