cs.CVOct 8, 2026

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

Authors: Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang, Canyu Chen, Jie Hao, Xing Fan, +4 more

Organizations: Northwestern University · UNC Chapel Hill · Carnegie Mellon University · Amazon

Abstract

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.

Figures & tables

Appendix figures & tables76 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning

    May 25, 2026Longteng Guo, Yifan Wang, Pengkang Huo +4Visual ReasoningVLM Reasoning

  2. WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

    Jun 4, 2026Yida Yin, Harish Krishnakumar, Chung Peng Lee +9Multimodal Large Language ModelsMultimodal Reasoning

  3. Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

    Jul 28, 2026Jiaang Li, Chengzu Li, Zhaochong An +4VLM EvaluationMultimodal Large Language Models