Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{autonomous video-to-simulation} as a software engineering task in which an agent observes an embodied video, constructs the corresponding simulated environment and robot behavior, and iteratively refines the result through execution feedback. To evaluate this capability, we introduce \textbf{Video2World}, a benchmark comprising 222 reconstruction instances derived from 189 robot and human demonstration videos. Video2World measures reconstructed worlds along geometric fidelity, dynamic fidelity, and functional correctness, capturing spatial perception, physical reasoning, and executable interaction. Evaluating 9 frontier coding-agent systems reveals a sharp improvement in Task success beginning with Claude Opus 5, rising from below 5% to over 15%, while substantial gaps to human-assisted reconstruction remain. We further find that worlds that look better could work worse: better visual fidelity does not always lead to higher task success. This echoes the broader gap between perceptual realism and factual correctness observed in generative models.
Figures & tables
Benchmark
Input
Embodied
Physics
Spatial
Agent
Coding
Evaluation Task
SWE-bench ( 2024 )
Code Repository
×
×
×
×
✓
Code Generation
VSI-Bench ( 2025 )
RGB Video
×
×
✓
×
×
Question Answer
RoboDojo ( 2026b )
Interactive Env.
✓
✓
×
×
×
Robot Manipulation
CaP-X ( 2026 )
Interactive Env.
✓
✓
×
✓
✓
Robot Manipulation
RLE-Bench ( 2026 )
Interactive Env.
✓
✓
✓
✓
✓
Robotics Engineering
EmbodiedSWE-Bench ( 2026 )
Interactive Env.
✓
✓
×
✓
✓
Task Completion
Table 1: Comparison with related benchmarks. ✓ yes and × no denote whether the benchmark focuses on evaluating the corresponding capability: Embodied for embodied reasoning and planning; Physics for physical-world dynamics and interaction; Spatial for spatial understanding; Agent for agentic planning and iterative interaction; Coding for executable code generation.
Figure 1: Video2World benchmarks the conversion of embodied videos into simulated worlds. Given an RGB video and simulator tools, a coding agent constructs the environment and robot behavior, then revises its implementation using execution feedback. Evaluation examines geometric fidelity, physically generated object motion, and task completion.
Figure 2: Overview of Video2World. A coding agent constructs the scene and robot behavior in a sandbox, using video inspection, simulator APIs, asset resources, and execution feedback. Submitted reconstructions are evaluated against hidden scene geometry, demonstrated object motion, and task requirements. Evaluation separately measures geometric fidelity, dynamic fidelity of object motion, and functional correctness under physical execution.
V2WScore
Build
Functionality
Geometry
Dynamics
Method
Score (0–100) ↑
Rate (%) ↑
Success (%) ↑
Progress (0–1) ↑
Scene CD (cm) ↓
Shape CD (cm) ↓
Size err. (cm) ↓
T-APE (cm) ↓
R-APE (deg) ↓
T-RPE (cm) ↓
Human-assisted Ref
74.24
100.0
58.8
0.70
2.19
0.15
0.43
9.40
41.09
2.42
Claude Fable 5.1
48.52
93.1
25.5
0.53
5.35
1.00
2.78
18.71
104.61
8.25
GPT-6 Astra
43.55
93.1
10.7
0.33
6.23
0.69
1.62
20.78
95.08
6.19
Claude Opus5
41.56
92.8
16.0
0.34
6.01
1.14
3.13
19.37
110.85
8.42
Kimi K3
33.19
92.0
4.8
0.21
8.08
1.41
4.23
27.28
111.16
14.29
Table 2: Overall performance on Video2World. V2WScore combines Build, functionality, geometry, and dynamics into a single 0–100 score, macro-averaged across task families. Build measures executability; Functionality reports task success and stage-wise progress; Geometry measures scene- and object-level reconstruction accuracy; and Dynamics reports translation APE, rotation APE, and translation RPE. The human-assisted reference additionally uses source information and manual refinement. Best and second-best automated results are shown in bold and underlined, respectively; ties share the same formatting. Metric definitions, deformation and partial-observation terms, aggregation, and missing-data handling are detailed in Appendices 9.7 and 9.4 .
Figure 3: Human in the loop reference and evaluation (a) Human preference ratings of coding models and the human-assisted reference, based on 243 judgments from 12 raters across 39 cases. Error bars indicate 95% confidence intervals. (b) Source demonstrations and simulated executions for two representative tasks. Each method executes its own behavior in its reconstructed environment. Outcomes and task progress are reported below each simulated sequence; red boxes highlight failed executions.
Figure 4: Geometric fidelity versus functional correctness. (a) System-level Shape CD against Task Success; both are macro-averaged over task families, and Task Success counts failed builds as failures. (b) Task Success across geometric-error quartiles formed within each model and task family, averaged equally across families and then models. Dashed lines mark the mean of each panel; error bars show 95% source-video bootstrap intervals. The initial relation is the object–receiver relation before execution, since the terminal relation enters the success predicate.
Figure 5: Robot Effector Motion Beyond Object-Centric Evaluation. (a) Source versus simulated effector statistics for 40 human-assisted reference instances from 35 videos, selected by task success in the original evaluation and colored by source. We compare gripper TCPs on DROID and human wrists with robot hand roots on human-video sources. Diagonals indicate equality; ρ is Spearman correlation. (b) Simulated-minus-source differences normalized by the source sample SD. Boxes show median/IQR; whiskers use 1.5× IQR. Shading marks ±0.25 SD; arrows count off-scale points; the right column gives median absolute differences. See Appendix 9.6 for definitions.
Resource
Interface specification
Source observation
RGB video; resolution, frame count, and the sampling information stated in the task brief. No depth or measured source calibration.
Public specification
Target simulator and robot; simulator documentation, robot assets, joint limits, control conventions, and submission requirements.
Workspace tools
Shell execution; reading, writing, and editing files; filename and text search. Image-capable file reads expose inspected source frames or candidate renders.
Construction feedback
Program output and errors; package-validation errors; candidate physical rollouts, rendered observations, and available execution diagnostics.
Withheld information
Reference scene and object meshes, measured source trajectories, source robot logs, hidden camera calibration, reference submissions, and benchmark scores.
Table A1: Agent-visible resources and withheld information. Public resources are selected by the assigned configuration. Candidate observations concern the agent’s own reconstruction.
Configuration
N
Simulator / robot
Submitted control
Furniture assembly
64
SAPIEN / Panda
TCP position, orientation, and gripper; 7 values at 5 Hz.
DROID tasks
20
SAPIEN / Panda with Robotiq 2F-85
TCP position, orientation, and gripper; 7 values at 5 Hz.
Human video to arm
33
SAPIEN / Panda
TCP position, orientation, and gripper; 7 values at 5 Hz.
Human video to hands
33
SAPIEN / paired Wuji hands
Two palm streams of 7 values and two finger streams of 20 angles; synchronized timestep.
Planar pushing
3
SAPIEN / xArm7 with pusher
TCP position and orientation at 5 Hz; seventh value unused.
RoboDojo
25
Isaac Sim / dual X5 or dual xArm7
14 or 16 joint/gripper values at 25 Hz.
Table A2: Registered candidate interfaces. N is the number of reconstruction instances; the total is 222. Frequencies refer to submitted controls, not physics integration. For each floating hand, the seventh palm value records the open/closed convention; its twenty finger angles determine the actual posture.
Artifact
Required content
protocol.json
Version 3.0 manifest: target configuration, scene or scene-file reference, camera, robot, action-file references, timestep, and provenance.
actions.npy
Finite control matrix with the embodiment-specific width in Table A2 ; Cartesian streams are stored as 64-bit floating-point arrays.
Additional hand arrays
For paired hands, a second N×7 palm stream and two N×20 finger-angle streams, with four distinct file references.
scene.json / scene.xml
Native Isaac submissions use a separate scene description; MuJoCo submissions use an MJCF scene with the supplied robot and explicit object, robot, and target roles.
assets/
Locally referenced visual and collision geometry. Native mesh submissions provide vertices and face indices, with metric scale incorporated into the vertices.
expected/ obj_poses.npy
Camera-frame T×7 object-pose estimates for the camera-frame rigid and human-video profiles. These are declarations, not commands; native joint and MuJoCo interfaces do not require them.
Table A3: Submission artifacts. T denotes source-video frames and N submitted control rows; these lengths need not agree.
Capability
Agent-specified input
Agent-visible feedback
Shell execution
Program or command; optional execution limit
Standard output, execution errors, exit status, and timeout information.
File and image reading
File or directory; optional text range
File contents, directory entries, or an image observation.
Filename search
Filename pattern and search location
Matching file paths.
Text search
Text pattern and search scope
Matching passages and file locations.
File creation
File location and contents
Confirmation, diagnostics, or a write error.
Text replacement
File and requested replacement
Modification result or an error.
Table A4: Coding capabilities available to the agent. All operations are restricted to the permitted workspace and public resources. Model-specific tool definitions are supplied by the pinned agent framework.
Operation
Input
Agent-visible feedback
Effect on the episode
Frame extraction
Source video and sampling specification
Frame references, timestamps, and image metadata
Makes frames available for inspection; the candidate is unchanged.
Physical replay
Candidate scene and robot controls
Rendered rollout and available execution diagnostics
Supports inspection and revision of the reconstruction; does not submit or score it.
Validation
Candidate reconstruction and task configuration
Missing artifacts and violations of the public submission requirements
Construction continues; no candidate is submitted.
Checkpointing
Candidate reconstruction
Validation feedback and confirmation of preservation
A valid candidate is retained as a fallback during the permitted refinement.
Submission
Candidate reconstruction
Acceptance or validation errors
Acceptance freezes the final candidate and ends generation; an invalid submission consumes one opportunity.
Table A5: Benchmark operations during construction. The interface separates visual and physical observations, structural feedback, and candidate delivery. An accepted submission terminates generation; successful validation alone does not.
Elapsed time
Construction and delivery policy
About 30 min
Complete an initial evidence-based candidate; validate it and repair missing or malformed artifacts.
120 min
Restrict remaining work to completing and repairing delivery.
150 min
Close the optional refinement allowance.
165 min
Stop generation and new delivery requests; select the latest valid protected checkpoint if no final submission has been accepted.
180 min
Complete process termination, integrity verification, and publication of the selected candidate.
Table A6: Construction schedule for benchmark-managed episodes. Times are cumulative within one attempt. The first milestone is a prompt instruction; the generation and publication deadlines are enforced by the executor.
Model
Model identifier
Provider
Effort
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash
novita/fp8
–
GLM-5.3 Flash
z-ai/glm-5.3-flash
novita/fp8
–
Kimi K3
moonshotai/kimi-k3
moonshotai/mxfp4
–
GPT-5.6
openai/gpt-5.6-sol
openai
–
Gemini 3.8 Flash
google/gemini-3.8-flash
google-vertex/ global
High
Claude Opus 5
anthropic/claude-opus-5
anthropic
–
Table A7: Model configurations. The first seven rows use OpenCode and the model-identifier prefix openrouter/ . A dash denotes an unspecified reasoning-effort setting.
Video source
Clips
Instances
Target embodiment
FurnitureBench
64
64
Panda
DROID
20
20
Panda/Robotiq (18); cloth gripper (2)
RoboDojo
25
25
Dual X5 or dual xArm7
HOI4D
9
18
Panda; paired Wuji hands
HOT3D
7
14
Panda; paired Wuji hands
DexYCB
11
22
Panda; paired Wuji hands
Table A8: Composition of the construction inventory. Clips count distinct staged video windows; instances count clip–configuration pairs. Embodiments include the end effector. Shared recordings and reuse across embodiments are retained in the provenance.
Source
Available source evidence
Benchmark annotation work
FurnitureBench
CAD, tags, calibrated cameras, and robot logs.
Register part poses; fit grasp offsets to bridge occlusions during verified holding.
DROID
Calibrated RGB-D, multiple views, and robot kinematics.
Fit object geometry and poses; reconstruct supports and cavities; track visible cloth surfaces.
RoboDojo
Native assets and synchronized simulation states.
Export geometry, cameras, and trajectories from the same execution as the video.
HOI4D
Instance CAD, object boxes, object poses, and camera trajectories.
Compose object and camera poses in a gravity-aligned frame; model supporting surfaces.
HOT3D
Object CAD and tracked object/camera poses.
Register trajectories; identify supports and static/dynamic roles.
DexYCB
YCB meshes, calibrated views, and object/hand poses.
Convert pose units and frames; estimate support height from initial mesh bottoms.
Table A9: Evidence used to construct the hidden annotations. The middle column identifies available source information; the last identifies the processing and modeling needed for benchmark evaluation. Reconstructed twins supply modeled targets.
Figure A1: Complete task-family inventory. Each bar is one family in the fixed construction manifest, and its label gives the number of video–configuration pairs. The 39 counts sum to 222 instances. Arm and hand variants count separately; their underlying source clips are shared. Family identifiers F01–F39 index the accompanying catalogue. Counts describe the construction inventory rather than a model’s successful builds or the metric-specific evaluable subset.
Figure A2: From a video demonstration to an object-level evaluation specification. The manually drawn outlines identify the same bottle in frames A–C; they are visual guides, not segmentation annotations. Corresponding points on its annotated height trajectory identify the initial support, transport, and receiving support; dashed lines mark the annotated pickup and release. The conditions at right define acceptable terminal placement, supplementing the required transfer. The curve is computed from source object poses, while the conditions are the evaluation targets; neither is a robot-execution result.
Figure A3: A complete evaluated reconstruction. Top: three frames from the agent’s RGB input. Middle: three frames of its submitted controls executed in its reconstructed scene; the displayed times use the recorded observation clock. Bottom left: annotated and executed bottle-height changes after the object-frame correspondence used for scoring. The shaded tail lies beyond the demonstration’s motion-comparison window, while functional evaluation still uses its release and terminal state. Bottom right: per-frame translation discrepancy on the demonstration clock, whose mean is 7.77 cm. The recorded interaction succeeds despite temporal and geometric discrepancies. Full videos, annotations, controls, execution records, and scoring inputs accompany this example.
Source
Additional annotations and assets
Use in reference construction
FurnitureBench
Part CAD, calibrated part/robot motion, assembly targets, simulator state.
Reuse part and collision assets; initialize grasps and transport; refine contact, alignment, and release.
Initialize scene and controls from source records; fit object-guided motion and refine execution. The two cloth references retain their declared assisted grasp mechanism.
RoboDojo
Native scenes, robot and object assets, task rules, recorded controls and states.
Adapt the recorded native task execution to the benchmark interface.
HOI4D
Instance CAD, object and camera trajectories, fitted support, available hand observations.
Retarget object transport to the arm or hands; optimize contact using annotated motion.
HOT3D
Object CAD, tracked object/camera motion, available hand poses, support estimates.
Initialize grasp and transport, then refine wrist/palm and finger controls.
DexYCB
YCB meshes, calibrated views, object/hand poses, estimated support height.
Reuse geometry and pose observations to construct and refine lifting behavior.
Table A10: Information available to human-assisted reconstruction beyond RGB. Inputs vary by source and instance. The final column describes their use in constructing the selected reference; it does not claim that all assets were reconstructed from video. The evaluated coding agents do not receive these source-specific annotations or reference executions.
Figure A4: Source-level reference outcomes and recorded construction effort. Left: outcomes of the selected references under the current evaluation snapshot, including failures and suspended task targets. These are instance counts; the headline score instead gives equal weight to task families. Right: trial counts from the two audited campaigns, with the number of covered configurations relative to that source’s inventory. Unconsolidated histories are explicitly marked. Trial records do not measure human person-hours or complete manual-iteration counts, and coverage differs across sources.
Figure A5: Evaluating the requested interaction separately from geometric fidelity. Each row shows the demonstrated outcome and three frames from the scored Astra visual-feedback execution, with manually drawn endpoint outlines identifying the manipulated object. (a) The marker remains partly in the container; robot withdrawal does not establish object extraction. (b) The bottle is reoriented and placed upright, satisfying the task predicate despite the reported shape and size errors. Execution frames are 0, 36, and 72 for (a), and 0, 48, and 96 for (b). Judgments use the complete saved rollout; frames are cropped for visibility and are not temporally aligned with the source endpoint.
Category
Criterion
Count
Share
HAR pass
—
113
50.9%
A1 Physical-validity rejection
valid=false
19
8.6%
A2 No passing execution in family
family HAR 0/n
18
8.1%
B Near miss
margin ≤1.5× tolerance
7
3.2%
C Controller failure
valid, family has passes
65
29.3%
Table A11: Attribution of human-assisted reference outcomes on the 222-instance inventory.
Method
S0 all (215)
S2 primary (179)
S1 HAR-pass (112)
Human-assisted Ref
52.1
71.6
100.0
Fable-5.1
19.1
20.6
24.2
Claude Opus 5
12.1
11.7
14.1
GPT-6 Astra
8.1
9.8
9.7
Kimi K3
3.6
4.2
3.7
DeepSeek V4.1 Flash
3.0
2.3
2.4
Table A12: Task Success (%, per-instance) under three evaluation denominators. S2 is the primary set used in Table .
Cell
Behavior
Scene
Own success
agent actions
agent scene
HAR actions × agent scene
HAR actions, translated, open-loop
agent scene
Agent × HAR scene (submitted)
agent actions, unchanged
HAR scene
Agent × HAR scene (anchored) †
agent actions, XYZ-translated
HAR scene
HAR × HAR
HAR actions
HAR scene
Probe × scene
state-aware scripted controller
either scene
Table A13: Cross-execution cells. Each transplanted behavior is inserted into the target scene package and scored by that instance’s final family evaluator using the corresponding main-experiment extrinsics.
Set
Cell
Pooled
Macro
Rows
A (20 inst.)
Own success
24.3 [14.9, 34.6]
25.9
177
Probe × agent scene
13.1 [6.3, 20.8]
7.3
175
HAR actions × agent scene
12.1 [5.8, 19.5]
6.9
174
Agent × HAR scene (submitted)
2.3 [0.6, 4.5]
1.4
174
Agent × HAR scene (anchored) †
15.5 [7.3, 24.6]
8.7
174
Probe × HAR scene
100 (condition)
100
180
Table A14: Cross-execution success (%) pooled over nine models, with instance-clustered 95% confidence intervals, family macro-average, and number of available rows. Sets A and B are conditioned diagnostic subsets, not the full Video2World evaluation population.
Set A (172 rows)
Set B (385 rows)
Own
Scene / Behavior test
rows
share
rows
share
Success
✓/ ✓
10
23%
12
19%
✓/ ×
7
16%
7
11%
× / ✓
7
16%
14
22%
× / ×
19
44%
30
48%
Failure
✓/ ✓
2
2%
4
1%
Table A15: Outcome decomposition (row counts). The scene test is probe success in the agent scene for Set A and HAR-action success in the agent scene for Set B. The behavior test is anchored agent-action success in the HAR scene. These are transfer diagnostics rather than intrinsic binary labels of scene or behavior quality.
Model
Own
Probe × agent
HAR × agent
Agent × HAR (subm.)
Agent × HAR (anch.) †
Fable-5.1
70.0
40.0
25.0
5.0
35.0
Claude Opus 5
42.1
21.1
21.1
5.3
10.5
GPT-6 Astra
35.0
5.0
30.0
5.0
25.0
GPT-5.6 Sol
15.0
15.8
10.0
0.0
15.0
DeepSeek V4.1 Flash
15.0
5.0
5.3
0.0
15.8
Gemini 3.8 Flash
15.0
15.0
10.0
5.0
0.0
Table A16: Set-A cross-execution success (%) per model; 20 valid rows per model unless noted. Instance-clustered 95% intervals span approximately 20–40 percentage points, so per-model differences are descriptive.
Figure A6: Robot Experience Beyond the Source Video. Direct reuse and agent adaptation across translated object and target layouts. The same nine source videos (three per task family) contribute at every displacement, with 180 paired layout evaluations per model. Each source uses the same two table directions and both moved-object roles at every magnitude. Both directions were verified using the human-assisted reconstruction jointly across all five magnitudes before new candidate evaluations. Success is averaged within video and then across task families. Shading shows 95% paired source-video bootstrap intervals within family. The 0 cm points are verified original baselines without adaptation.
API cost (USD)
Agent time
API calls
Eval. GPU
Method
mean
median
max
est. total
mean (min)
mean
mean (min)
Fable-5.1 †
14.03
12.90
45.19
3,115
37.5
–
10.1
GPT-6 Astra †
6.40
3.91
24.58
1,421
14.6
–
9.7
Claude Opus 5
13.83
10.03
61.07
3,070
32.6
59
8.7
Kimi K3
5.51
4.81
13.27
1,223
82.2
72
7.4
GPT-5.6 Sol
1.39
1.00
7.89
309
3.7
17
8.2
Table A17: Per-episode cost of each method. API cost is in USD; agent time is wall-clock time from the first model call to termination; evaluation time is GPU-minutes on one L40S. Estimated totals extrapolate the reported per-episode means over all 222 benchmark instances.
Figure A7: Task completion despite pose disagreement. DROID block stacking with a Panda–Robotiq embodiment. The source and candidate execution are shown at three task phases. The top-view inset compares the submitted and annotated block orientations. Successful stacking coexists with a large object-frame rotation error; the rigidly aligned shape score does not measure the submitted world orientation.
Figure A8: Endpoint proximity does not establish the required interaction. DROID shape-sorter insertion. The block is transferred into the submitted open container, but the receiver lacks the apertured top visible in the annotation. The final relative-position error passes its tolerance; top-face coverage fails its separate requirement. The source and candidate rows show task phases, rather than equal timestamps.
Figure A9: Nine reconstructions of the same bowl-nesting demonstration. Top: the source video’s initial and terminal configurations. The nine model panels show the last frame of each submission’s evaluation review video, using a common crop within this instance. Scene CD and translation APE are in centimetres. Opus, GPT-5.6, and DeepSeek succeed; the other six fail the receiving relation. Fable’s smaller surface error than Opus does not imply successful placement. Scores and decisions come from the complete execution records, including terminal settling; the displayed video endpoints are qualitative views rather than substitutes for those records.
Figure A10: Nine reconstructions of the same can-lifting demonstration. Source frames and model panels follow Figure A9 , with Scene CD and translation APE in centimetres. Opus and Kimi succeed. Qwen and Astra meet the ungated terminal lift condition but violate inverse-kinematics feasibility and robot-penetration checks, respectively. DeepSeek and Fable have valid executions without the required lift. The remaining three violate both checks. In particular, an elevated can in a rendered frame does not establish physical validity, which is determined from the execution record.
Figure A11: Nine models, two tasks, and different reasons for failure. (a) Terminal receiver-relative position error for black-bowl nesting. (b) Terminal can-height change for DexYCB lifting. Dashed lines mark the task-specific positional and height thresholds, which are components of the complete predicates. Green denotes Task Success, grey an unmet task condition, and red rejection by physical-validity checks. Qwen and Astra exceed the lift threshold but are rejected for inverse-kinematics infeasibility and robot penetration, respectively. Left-pointing markers denote out-of-range negative displacements, with their values stated explicitly. No aggregate model ranking is inferred from these two selected cases.
Figure A12: Successful containment with spatial and temporal mismatch. HOT3D food-package placement by a Franka arm. Top: source and candidate execution at selected task phases. Bottom left: the annotated scene in grey and the reconstruction in colour, viewed from the side in a common frame. Bottom right: displacement of each recorded object origin from its own initial position. The shaded interval is the 1.6 s demonstration window; the candidate execution continues to 7.2 s. This displacement plot illustrates timing and is not itself an APE curve.
Figure A13: One demonstration, two target robots, and distinct execution outcomes. Top: the RoboDojo source video’s initial and terminal arrangements. Middle and bottom: terminal states of Astra and Fable with X5 and xArm7 arms. Each panel is drawn from its own recorded physical execution; insets magnify the outlined receiver region. Astra’s X5 objects remain stationary, while its xArm7 submission succeeds. Fable succeeds with X5; its xArm7 stack satisfies the object-stacking subpredicate but fails robot return and gripper-coupling checks. Scene CD and translation APE are shown in centimetres. A visually plausible final stack therefore does not certify the complete interaction.
World models have emerged as a powerful paradigm for building interactive simulation environments, with recent video-based approaches demonstrating impressive progress in generating visually plausible dynamics. However, because these models typically infer dynamics from video and represent them in latent states, they do not explicitly enforce physical constraints. As a result, the generated video rollouts are not physically plausible, exhibiting unstable contacts, distorted shapes, or inconsistent motion. In this paper, we present an agentic framework constructing physics-based world models through executable simulation code. The framework coordinates planning, code generation, visual review, and physics analysis agents. The planning agent converts the natural language prompt into a structured scene plan, the code agent implements it as executable simulation code, and the visual review agent provide visual feedback while the physics analysis agent checks physical consistency. The code is iteratively revised based on the feedback until the simulation matches the prompt reqirements and physical constraints. Experimental results show that our framework outperforms advanced video-based models in physical accuracy, instruction fidelity and visual quality, which could be applied to various scenarios including driving simulation and embodied robot tasks.
Hongyu Wang, Jingquan Wang, Bocheng Zou +2
Department of Mechanical & Aerospace Engineering, University of Wisconsin-Madison · School of Computer, Data, and Information Sciences, University of Wisconsin-Madison
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.
Kaixin Ding, Xi Chen, Minghong Cai +9
The University of Hong Kong · Kling Team, Kuaishou Technology · The Chinese University of Hong Kong +1
Video world models have emerged as promising candidates for high-fidelity world models, offering the potential to synthesize high-quality videos capturing fine-grained interactions between agents and their environments conditioned on multi-modal user inputs. Their impressive capabilities address many of the long-standing challenges faced by physics-based simulators, driving broad adoption in many problem domains, e.g., robotics. For example, video models can generate photorealistic, physically consistent deformable-body simulation without making prohibitive simplifying assumptions, which is a major bottleneck in physics-based simulation. Moreover, video models can serve as foundation world models that capture the dynamics of the world in a fine-grained and expressive way. They thus overcome the limited expressiveness of language-only abstractions in describing intricate physical interactions. However, despite their potential, video models still generate physics-violating future predictions, often manifesting as hallucinations. In this survey, we provide a review of video models and their applications as embodied world models in robotics, including efficient data generation and policy learning, dynamics and rewards modeling in reinforcement learning, policy evaluation, and visual planning. Further, we highlight important challenges hindering the trustworthy integration of video models, such as poor instruction following, hallucinations like violations of physics, unsafe content generation, in addition to significant data and compute overhead. We present potential future directions to address these open research challenges to motivate research and ultimately facilitate broader applications, especially in safety-critical settings. We provide a curated bibliography at https://github.com/irom-princeton/awesome-robotics-video-world-model-papers .