cs.CVSep 28, 2026

From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models

Authors: Rong Yu Xu, Prayag Tiwari, Shaolei Zhang

Organizations: Shenzhen College of International Education · Halmstad University, Sweden · Renmin University of China

Abstract

Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterfactual image pairs isolate changes in the visual evidence needed to answer. Across four models, direct answers, hidden-state readouts, and state interventions show that the individual judgments can be made without explicit reasoning and that intervening on the corresponding states can affect the answer. During reasoning, the Composite answer becomes decodable from hidden states and usable from shortened traces, often before the model stops on its own. We train a small detector to predict this readiness and stop reasoning at that point. On MMStar and RealWorldQA, this reduces mean reasoning tokens by 79.1% and 74.5%, while average accuracy rises by 3.13 and 3.30 percentage points, respectively. These findings connect the internal development of answer readiness to a practical rule for allocating reasoning computation.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models

    May 19, 2026Juncheng Wu, Hardy Chen, Haoqin Tu +6Recent Vision-Language ModelsVisual Reasoning

  2. Seeing and Solving Are Not Enough for Vision-Language Models

    Sep 27, 2026Ziheng Wang, Mingxuan Xie, Yilin Liu +3Multimodal QueryState-Tracking

  3. Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models

    Apr 16, 2026Danae Sánchez Villegas, Samuel Lewis-Lim, Nikolaos Aletras +1Large Reasoning ModelsVisual Evidence