cs.CVSep 28, 2026

WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning

Authors: Yuheng Zha, Yilei Wang, Qiyue Gao, Junrong Chen, Yujia Wu, Zhengfeng Lai, Zhengzhong Liu, Eric P. Xing

Organizations: UC San Diego · Institute of Foundation Models, MBZUAI · Carnegie Mellon University

Abstract

Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states. Our two-stage training first teaches the model to generate the next visual state and then to use that state for reasoning. We programmatically construct spatial reasoning tasks with verifiable intermediate visual states. These tasks allow us to evaluate how well the model generates visual states and how much it relies on them to answer the question. On 2D and 3D mental rotation tasks, WM-VLM consistently outperforms the supervised fine-tuned backbone, with gains of up to 39.25 percentage points. Ablations suggest that these gains depend on the generated visual states, as removing or corrupting them sharply reduces performance. Together, these results suggest that internal world models offer a promising path toward VLMs that reason in both language and visual space.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning

    Apr 29, 2026Wanyue Zhang, Wenxiang Wu, Wang Xu +6Spatial ReasoningVideo World Models

  2. GeoWorld-VLM: Geometry from World Models for Vision-Language Models

    May 15, 2026Renjie Gu, Kaichen Zhou, Yan Luo +1Spatial SupervisionVideo World Models

  3. Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

    Jun 4, 2026Chenming Zhu, Jingli Lin, Yilin Long +4Spatial ReasoningImagination