World Model Evaluation

Momentum

29 papers in the last four weeks, up 123% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 148

All topics
CardsList
  1. RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

    Jun 22, 2026Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen +10VLM EvaluationRobot Manipulation Benchmarks

  2. Current World Models Lack a Persistent State Core

    Jun 18, 2026Jinpeng Lu, Dexu Zhu, Haoyuan Shi +8Physical Consistency in Video GenerationWorld Models

  3. SurgVista: Long-Horizon Surgical World Modeling with Plausible Instrument-Tissue Dynamics

    Jun 18, 2026Wentao Pan, Wuyang Li, Shengyuan Liu +3World ModelsVideo Prediction

  4. SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

    Jun 17, 2026Wei-Cheng Tseng, Gashon Hussein, Yuzhu Dong +9Action-Conditioned Video GenerationPhysical Consistency in Video Generation

  5. ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control

    Jun 15, 2026Junjian Zhang, Hao Tan, Ruonan Li +3World Models for RoboticsAdversarial Robustness

  6. Mind-Studio: Executable World Models with Lookahead Evaluation for Partially Observable Games

    Jun 14, 2026Yifei Dong, Mingen Zheng, Linquan Wu +2World Model LearningLLM-Based Program Synthesis

  7. How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position

    Jun 13, 2026Yang Yu, Shiyuan Zhang, Yifei Sheng +2Model-Based PlanningWorld Model Evaluation

  8. ReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving

    Jun 12, 2026Zhiyuan Zhang, Yanlun Peng, Jianing Zhang +7Autonomous Driving BenchmarksWorld Models for Autonomous Driving

  9. WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

    Jun 11, 2026Arnav Kumar Jain, Yilin Wu, Jesse Farebrother +2World Model LearningLong-Horizon Robotic Manipulation

  10. Certified World Models: Predictability Across Configuration, Horizon, and Resolution

    Jun 11, 2026Hongbo WangEquivariant Neural NetworksLatent World Models

  11. WorldOlympiad: Can Your World Model Survive a Triathlon?

    Jun 9, 2026Yuke Zhao, Wangbo Zhao, Weijie Wang +8Physical Consistency in Video GenerationLong-Horizon Video Generation

  12. Echo-Memory: A Controlled Study of Memory in Action World Models

    Jun 8, 2026Wayne King, Zeyue Xue, Yuxuan Bian +13Action-Conditioned Video GenerationVideo Diffusion Models

  13. Bridging the Agent-World Gap: Text World Models for LLM-based Agents

    Jun 8, 2026Yixia Li, Hongru Wang, Peng Lai +13LLM World ModelsWorld Models

  14. ATM: Why Latent World Models Can Fail to Plan

    Jun 8, 2026Jiaheng Chen, Tinghe Zhang, Yucheng Xiao +5World Model-Based PlanningModel-Based Planning

  15. Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation

    Jun 3, 2026Luca Zanatta, Grzegorz Malczyk, Kostas AlexisWorld Models for RoboticsWorld Models

  16. RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

    Jun 1, 2026Huiqiong Li, Jiayu Wang, Zhiting Mei +3World ModelsAI Safety Evaluation

  17. Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends

    May 31, 2026Jiuming Liu, Chaojun Ni, Mengmeng Liu +7Action-Conditioned Video GenerationLong-Horizon Video Generation

  18. MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

    May 30, 2026Shengjun Zhang, Zhang Zhang, Simin Huang +11Long-Horizon Video GenerationWorld Model Evaluation

  19. World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications

    May 28, 2026Arif Hassan Zidan, Yi Pan, Hanqi Jiang +23World Model LearningWorld Model-Based Planning

  20. Physically Viable World Models: A Case for Query-Conditioned Embodied AI

    May 28, 2026Adam J. Thorpe, Stepan Tretiakov, Cheng-Hsi Hsiao +6World Models for RoboticsCausal World Models

  21. Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences

    May 28, 2026Benjamin Walker, Terry LyonsRecurrent Neural NetworksState Tracking

  22. MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

    May 28, 2026Tianzhuo Yang, Zihan Shen, Zirui Mi +7World Models for RoboticsAction-Conditioned World Models

  23. World Models for Robotic Manipulation: A Survey

    May 27, 2026Fangyuan Wang, Ziyuan Wang, Guorui Pei +15World Model LearningWorld Models for Robotics

  24. What-If World: A Causal Benchmark for General World Models in Embodied Scenarios

    May 26, 2026Kunlin Cai, Rui Song, Jinghuai Zhang +7Physical Consistency in Video GenerationCausal World Models

  25. WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation

    May 25, 2026Kaining Ying, Hengrui Hu, Siyu Ren +6World Model EvaluationInteractive World Models

  26. World Models as Group Actions

    May 23, 2026Zijie Wang, Wei Zhang, Weiming Zhang +4World Model LearningAction-Conditioned World Models