ROWBench: Do Video Models Render What the Program Specifies?
Organizations: Alaya Lab
Abstract
Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be checked against the observable consequences of program execution. An extensible framework constructs scenes, controls behaviors, and can render each camera view in different representations, such as coarse 3D, and bounding boxes. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Grounded in these records, PROWBench evaluates entity control, long-horizon memory, and, with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate, adherence to the prescribed timeline and the visual realization of timestamped engine-recorded events.
Figures & tables
| Method | Input | Entity control | Logic / State | Camera control | Quality | Temporal | Text | ||||||||
| IoU | CErr | ISR | LRA | gISR | gLRA | Rot | Trans | CamMC | Img | Aes | Warp | Smooth | CLIP | ||
| Verified-FF (40 pairs) | |||||||||||||||
| EchoWM ( Zhang et al., 2026 ) | 0.351 | 0.055 | 0.506 | 0.685 | 0.168 | 0.241 | 2.2 | 0.151 | 0.166 | 67.6 | 0.607 | 0.619 | 0.987 | 25.2 | |
| SANA-WM ( Zhu et al., 2026 ) | 0.309 | 0.070 | 0.590 | 0.569 | 0.185 | 0.180 | 2.7 | 0.145 | 0.168 | 68.6 | 0.612 | 0.871 | 0.985 | 25.1 | |
| LingBot 2.0 ( Robbyant Team et al., 2026 ) | 0.327 | 0.063 | 0.733 | 0.794 | 0.241 | 0.256 | 2.8 | 0.100 | 0.133 | 69.9 | 0.615 | 1.026 | 0.978 | 25.4 | |
| AlayaWorld v1.1 ( Team et al., 2026 ) | 0.303 | 0.073 | 0.583 | 0.622 | 0.174 | 0.185 | 9.4 | 0.229 | 0.356 | 66.8 | 0.592 | 0.530 | 0.990 | 25.0 | |
| Method | Input | Entity control | Logic / State | Memory | Camera control | |||||||
| IoU | CErr | ISR | LRA | gISR | gLRA | Re-IoU | Re-State | Rot | Trans | CamMC | ||
| C2R ( Gomez-Nogales et al., 2026 ) | 0.212 | 0.066 | 0.519 | 0.718 | 0.127 | 0.170 | 0.111 | 0.405 | 41.9 | 0.571 | 1.119 | |
| CWM ( Chen et al., 2026b ) | 0.210 | 0.070 | 0.504 | 0.854 | 0.111 | 0.182 | 0.053 | 0.886 | 30.2 | 0.381 | 0.793 | |
| LynnReal-Omni ( Mao et al., 2026b ) | 0.280 | 0.054 | 0.515 | 0.897 | 0.149 | 0.247 | 0.103 | 0.528 | 22.7 | 0.273 | 0.610 | |
| Seedance 2.5 ( ByteDance Seed Team, 2026 ) | 0.175 | 0.077 | 0.497 | 0.777 | 0.087 | 0.132 | 0.131 | 0.650 | 39.4 | 0.242 | 0.819 | |
| MiniMax-H3 ( MiniMax-AI, 2026 ) | 0.206 | 0.088 | 0.553 | 0.844 | 0.121 | 0.179 | 0.059 | 0.833 | 31.8 | 0.351 | 0.814 | |
| Method | Input | Entity control | Logic / State | Appearance | Camera control | |||||||||
| IoU | CErr | Miss | ISR | LRA | gISR | gLRA | Comp. | Cons. | MEt3R | Rot | Trans | CamMC | ||
| C2R ( Gomez-Nogales et al., 2026 ) | 0.414 | 0.039 | 0.192 | 0.637 | 0.750 | 0.271 | 0.318 | 32.2 | 25.5 | 0.397 | 2.2 | 0.155 | 0.173 | |
| CWM ( Chen et al., 2026b ) | 0.439 | 0.037 | 0.064 | 0.812 | 0.950 | 0.345 | 0.414 | 92.9 | 88.1 | 0.422 | 2.0 | 0.047 | 0.071 | |
| LynnReal-Omni ( Mao et al., 2026b ) | 0.490 | 0.038 | 0.047 | 0.700 | 0.900 | 0.347 | 0.435 | 87.6 | 89.7 | 0.448 | 0.4 | 0.023 | 0.026 | |
| Seedance 2.5 ( ByteDance Seed Team, 2026 ) | 0.341 | 0.049 | 0.089 | 0.812 | 0.963 | 0.284 | 0.331 | 82.2 | 81.0 | 0.499 | 1.0 | 0.076 | 0.083 | |
| MiniMax-H3 ( MiniMax-AI, 2026 ) | 0.524 | 0.029 | 0.049 | 0.825 | 0.938 | 0.428 | 0.489 | 94.0 | 88.6 | 0.400 | 0.9 | 0.034 | 0.043 | |
| Method | Input | Entity control | Logic / State | Camera control | Quality | Temporal | Text | ||||||||
| IoU | CErr | ISR | LRA | gISR | gLRA | Rot | Trans | CamMC | Img | Aes | Warp | Smooth | CLIP | ||
| C2R ( Gomez-Nogales et al., 2026 ) | 0.437 | 0.041 | 0.450 | 0.725 | 0.191 | 0.330 | 2.7 | 1.008 | 1.043 | 69.6 | 0.544 | 0.301 | 0.994 | 24.3 | |
| CWM ( Chen et al., 2026b ) | 0.413 | 0.039 | 0.650 | 0.900 | 0.287 | 0.381 | 0.7 | 0.740 | 0.744 | 72.6 | 0.562 | 0.157 | 0.992 | 25.6 | |
| PWM Huang et al. (2026a) | 0.602 | 0.021 | 0.850 | 0.875 | 0.523 | 0.529 | 1.1 | 0.348 | 0.359 | 71.0 | 0.579 | 0.182 | 0.993 | 26.0 | |
| LynnReal-Omni ( Mao et al., 2026b ) | 0.552 | 0.033 | 0.625 | 0.975 | 0.340 | 0.535 | 1.3 | 0.563 | 0.584 | 75.0 | 0.527 | 0.613 | 0.988 | 25.2 | |
| MiniMax-H3 ( MiniMax-AI, 2026 ) | 0.526 | 0.035 | 0.600 | 0.900 | 0.302 | 0.475 | 5.1 | 0.770 | 0.830 | 73.9 | 0.556 | 0.107 | 0.995 | 25.4 | |
| Method | Input | Entity control | Logic / State | Camera control | Quality | Temporal | Text | ||||||||
| IoU | CErr | ISR | LRA | gISR | gLRA | Rot | Trans | CamMC | Img | Aes | Warp | Smooth | CLIP | ||
| MiniMax-H3 ( MiniMax-AI, 2026 ) | 0.514 | 0.030 | 0.625 | 0.802 | 0.304 | 0.391 | 4.8 | 0.072 | 0.156 | 70.1 | 0.655 | 0.570 | 0.987 | 25.1 | |
| MiniMax-H3 ( MiniMax-AI, 2026 ) | 0.507 | 0.031 | 0.683 | 0.665 | 0.338 | 0.333 | 4.5 | 0.079 | 0.156 | 70.7 | 0.650 | 0.573 | 0.987 | 25.1 | |
| MiniMax-H3 ( MiniMax-AI, 2026 ) | 0.503 | 0.033 | 0.733 | 0.690 | 0.347 | 0.339 | 5.0 | 0.065 | 0.155 | 70.7 | 0.652 | 0.574 | 0.987 | 25.1 | |
| Benchmark | Behavior specification | State / trajectory reference | Task / event reference | Replay | Views | Proxies |
| VBench [ 19 ] | Text prompts | – | Prompt semantics | – | – | – |
| WorldScore [ 11 ] | Text; camera paths | Target camera paths | Text-specified dynamics | – | – | – |
| WorldModelBench [ 24 ] | Text instructions | – | Instructions; physics criteria | – | – | – |
| MIND [ 57 ] | Agent navigation; camera rotation | Actor and camera poses | – | – | – | – |
| WorldArena [ 44 ] | Text; robot actions | Robot states and actions | Task goals; simulator outcomes | Sim. | – | – |
| Omni-WorldBench [ 55 ] | Text interactions; optional camera paths | Camera targets where provided | Specified object effects and event order | – | – | – |
| Verified-FF | Unverified-FF | |
| References | <Picture 1> (first frame) and <Video 1> (proxy) | <Video 1> only |
| Task | keyframe completion and reference generation; the generated clip must start exactly from the provided first frame | reference generation |
| Appearance source | Qwen3.6-27B description of the verified first frame | authored scene description |
| Style source | authored style line; the first frame also visually exhibits this style | authored style line |
| Control rule | for overall style and the appearance of subjects visible in the first frame, the first frame takes precedence over text | the appearance of every subject is specified by text |
| Structural skeleton | identical: subject definitions, proxy bindings, action timeline, and measured facts | |
| Methods | Prompt format |
| MiniMax-H3, LynnReal-Omni | base prompt verbatim, including the control block, tagged references, subject bindings with legend colors, and the complete timestamped timeline |
| Seedance 2.5 | Verified-FF: converted to Seedance’s reference syntax ( @Image1 , @Video1 ); retains the style line, the appearance of every subject, a “who is who” list identifying each box by color name and legend color, proxy semantics, the timestamped timeline, and field-of-view information; removes the control block, subject tags, retention analysis, and soundscape. Unverified-FF: base prompt with the control block, as for MiniMax-H3 (see System prompt ) |
| C2R, Cosmos-Transfer2.5, EchoWM, SANA-WM, LingBot 2.0, AlayaWorld v1.1 | converted to the single-paragraph caption style used to train C2R: an opening description of the style, setting, number of people, and appearance of every subject, followed by the action progression in prose without timestamps and a final quality sentence; removes tags, proxy terminology, and camera-specific text, and is shortened to satisfy C2R’s 512-token text limit |
| CWM | converted to CWM’s caption format, consisting of appearance and scene , a timestamped protagonist action timeline , and other visible actors ; removes subject tags and proxy-specific wording, while CWM adds its own system prompt for the mixed proxy render and RGB anchor frame |
| Method | Conditioning | Supplied spatial input | Output | FPS | Frames |
| EchoWM [ 60 ] | RGB first frame, camera trajectory | image | 24 | 121 | |
| SANA-WM [ 63 ] | RGB first frame, camera trajectory | image | 16 | 80 | |
| LingBot-World 2.0 [ 43 ] | RGB first frame, camera trajectory | image | 16 | 81 | |
| AlayaWorld v1.1 [ 46 ] | RGB first frame, camera trajectory; chunk-autoregressive generation, 4 rounds, 4 denoising steps per round | image, resized and cropped to | 24 | 120 | |
| Cosmos-Transfer2.5 [ 1 ] | Depth, class-level segmentation | , 24 fps, 121 frames | 24 | 121 | |
| C2R [ 15 ] | Coarse 3D proxy video | , 16 fps, 81 frames | 16 | 81 |