Organizations: Peking University Beijing, China · University of Science and Technology Beijing Beijing, China · City University of Hong Kong Hong Kong, China · Nanyang Technological University Singapore, Singapore
Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instructions. We introduce MindWorldBench to evaluate mental-state-conditioned video generation. We formalize this as mental-state-to-behavior reasoning, where models generate actions from a world state and latent variables without explicit action prompts. MindWorldBench utilizes Zero-Action Prompting and a counterfactual design with 744 prompts to isolate the causal effects of mental states. An automated pipeline evaluates video quality, commonsense plausibility, and mental-state consistency. Evaluations of 11 models show that despite visual fidelity and physical reasoning, models fail to align behaviors with latent mental states. We identify a failure mode, termed Omniscient Bias, where models default to the objective world state rather than human's subjective belief. These results demonstrate a disconnect between visual generation and cognitive reasoning, suggesting a need for explicit mental-state modeling in video generation systems. Project website: https://richard2049-lee.github.io/MindWorldBench/
Figures & tables
Figure 1 . Mental-State-to-Behavior Reasoning and the MindWorldBench Taxonomy. (Left) Evolution of Generation Paradigms: Unlike traditional Text-to-Video (explicit instruction alignment) or Video Reasoning (objective physical rules), our paradigm requires models to autonomously deduce behaviors from latent cognitive states. As illustrated, the model must prioritize the man’s subjective belief over objective reality to generate the correct reasoning-driven action. (Right) MindWorldBench Taxonomy: Grounded in the BDI-P framework, we systematically categorize mental variables into three primary dimensions— Perception , Belief , and Desire —to comprehensively evaluate Theory of Mind capabilities in video generation.
Figure 2 . Causal Graph of Mental-State-to-Behavior Generation. Shaded gray nodes ( Wt,V ) are observed visual states. Blue nodes ( M={P,B,D},I,A ) are latent cognitive variables. Models implicitly infer the latent intention I and action A to generate the video V .
Figure 3 . Representative examples from MindWorldBench. The benchmark spans the three mental dimensions: Belief , Desire , and Perception , covering diverse subcategories. Each example shows the generation input (image plus mental-state prompt); the displayed question is evaluation-only and is not provided to the generation model.
Models
Params
Visual Quality ↑
Commonsense Plausibility ↑
Intention Accuracy ↑
World-State Maintenance ↑
Kling ( Team et al., 2025a )
-
4.1
4.2
47.2
66.3
Veo3.1 ( Gallegos and Iljic, 2025 )
-
3.8
4.0
59.5
61.4
Seedance1.5Pro ( Seedance et al., 2025 )
-
4.3
4.3
29.3
52.9
Gen4 ( Runway, 2025 )
-
4.0
3.9
18.6
49.4
Hailuo2.3 ( MiniMax, 2025 )
-
3.8
3.8
13.0
32.5
CogVideoX ( Yang et al., 2024 )
5 B
3.9
4.0
27.7
54.2
Table 1 . Overall performance comparison on MindWorldBench.
Figure 4 . Performance breakdown across the three major mental dimensions: Belief, Perception, and Desire. Different models exhibit distinct strengths across dimensions. A brief plain-text description of the full-width figure.
Figure 5 . Diagnostic Analysis of Reasoning Failures. (a) Generated action type distribution. (b) Counterfactual paired outcomes. (c) Decomposition of action and world-state results into Both Correct, Action Correct, State Correct, and Both Incorrect.
Figure 6 . Action Completion (x-axis) vs. Conditional Correctness (y-axis), illustrating the trade-off between action generation and reasoning accuracy.
School of Computing, National University of Singapore, Singapore, Singapore · Kling Team, Kuaishou Technology, Beijing, China · Department of Computer Science and Technology, Nanjing University, Nanjing, China +1