cs.CVSep 30, 2026

MindWorldBench: Evaluating Mental-State-to-Behavior Reasoning in Image-to-Video Generation

Authors: Ruiqi Li, Xuanyi Liu, Sijia Li, Haofeng Wang, Yuxin Liu, Feng Xie, Songchao Tan, Shiqi Wang, +4 more

Organizations: Peking University Beijing, China · University of Science and Technology Beijing Beijing, China · City University of Hong Kong Hong Kong, China · Nanyang Technological University Singapore, Singapore

Abstract

Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instructions. We introduce MindWorldBench to evaluate mental-state-conditioned video generation. We formalize this as mental-state-to-behavior reasoning, where models generate actions from a world state and latent variables without explicit action prompts. MindWorldBench utilizes Zero-Action Prompting and a counterfactual design with 744 prompts to isolate the causal effects of mental states. An automated pipeline evaluates video quality, commonsense plausibility, and mental-state consistency. Evaluations of 11 models show that despite visual fidelity and physical reasoning, models fail to align behaviors with latent mental states. We identify a failure mode, termed Omniscient Bias, where models default to the objective world state rather than human's subjective belief. These results demonstrate a disconnect between visual generation and cognitive reasoning, suggesting a need for explicit mental-state modeling in video generation systems. Project website: https://richard2049-lee.github.io/MindWorldBench/

Figures & tables

Explore similar work

CardsList
  1. WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

    May 11, 2026Keming Wu, Yijing Cui, Wenhan Xue +11Long Video GenerationVideo Generation

  2. Thinking in Video: Can Video Generators Really Reason About the Real World?

    Jul 20, 2026Yongheng Zhang, Guang Yang, Ruihan Hou +12Generative Video ModelsVideo Generation

  3. From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

    Sep 12, 2026Meng Luo, Yicheng Liu, Jiahao Wang +5Generative Video ModelsVideo Understanding