Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos
Organizations: Aether AI · University of California, San Diego
Abstract
Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{autonomous video-to-simulation} as a software engineering task in which an agent observes an embodied video, constructs the corresponding simulated environment and robot behavior, and iteratively refines the result through execution feedback. To evaluate this capability, we introduce \textbf{Video2World}, a benchmark comprising 222 reconstruction instances derived from 189 robot and human demonstration videos. Video2World measures reconstructed worlds along geometric fidelity, dynamic fidelity, and functional correctness, capturing spatial perception, physical reasoning, and executable interaction. Evaluating 9 frontier coding-agent systems reveals a sharp improvement in Task success beginning with Claude Opus 5, rising from below 5% to over 15%, while substantial gaps to human-assisted reconstruction remain. We further find that worlds that look better could work worse: better visual fidelity does not always lead to higher task success. This echoes the broader gap between perceptual realism and factual correctness observed in generative models.
Figures & tables
| Benchmark | Input | Embodied | Physics | Spatial | Agent | Coding | Evaluation Task |
|---|---|---|---|---|---|---|---|
| SWE-bench ( 2024 ) | Code Repository | ✓ | Code Generation | ||||
| VSI-Bench ( 2025 ) | RGB Video | ✓ | Question Answer | ||||
| RoboDojo ( 2026b ) | Interactive Env. | ✓ | ✓ | Robot Manipulation | |||
| CaP-X ( 2026 ) | Interactive Env. | ✓ | ✓ | ✓ | ✓ | Robot Manipulation | |
| RLE-Bench ( 2026 ) | Interactive Env. | ✓ | ✓ | ✓ | ✓ | ✓ | Robotics Engineering |
| EmbodiedSWE-Bench ( 2026 ) | Interactive Env. | ✓ | ✓ | ✓ | ✓ | Task Completion |
| V2WScore | Build | Functionality | Geometry | Dynamics | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | Score (0–100) | Rate (%) | Success (%) | Progress (0–1) | Scene CD (cm) | Shape CD (cm) | Size err. (cm) | T-APE (cm) | R-APE (deg) | T-RPE (cm) |
| Human-assisted Ref | 74.24 | 100.0 | 58.8 | 0.70 | 2.19 | 0.15 | 0.43 | 9.40 | 41.09 | 2.42 |
| Claude Fable 5.1 | 48.52 | 93.1 | 25.5 | 0.53 | 5.35 | 1.00 | 2.78 | 18.71 | 104.61 | 8.25 |
| GPT-6 Astra | 43.55 | 93.1 | 10.7 | 0.33 | 6.23 | 0.69 | 1.62 | 20.78 | 95.08 | 6.19 |
| Claude Opus5 | 41.56 | 92.8 | 16.0 | 0.34 | 6.01 | 1.14 | 3.13 | 19.37 | 110.85 | 8.42 |
| Kimi K3 | 33.19 | 92.0 | 4.8 | 0.21 | 8.08 | 1.41 | 4.23 | 27.28 | 111.16 | 14.29 |
| Resource | Interface specification |
|---|---|
| Source observation | RGB video; resolution, frame count, and the sampling information stated in the task brief. No depth or measured source calibration. |
| Public specification | Target simulator and robot; simulator documentation, robot assets, joint limits, control conventions, and submission requirements. |
| Workspace tools | Shell execution; reading, writing, and editing files; filename and text search. Image-capable file reads expose inspected source frames or candidate renders. |
| Construction feedback | Program output and errors; package-validation errors; candidate physical rollouts, rendered observations, and available execution diagnostics. |
| Withheld information | Reference scene and object meshes, measured source trajectories, source robot logs, hidden camera calibration, reference submissions, and benchmark scores. |
| Configuration | Simulator / robot | Submitted control | |
|---|---|---|---|
| Furniture assembly | 64 | SAPIEN / Panda | TCP position, orientation, and gripper; 7 values at 5 Hz. |
| DROID tasks | 20 | SAPIEN / Panda with Robotiq 2F-85 | TCP position, orientation, and gripper; 7 values at 5 Hz. |
| Human video to arm | 33 | SAPIEN / Panda | TCP position, orientation, and gripper; 7 values at 5 Hz. |
| Human video to hands | 33 | SAPIEN / paired Wuji hands | Two palm streams of 7 values and two finger streams of 20 angles; synchronized timestep. |
| Planar pushing | 3 | SAPIEN / xArm7 with pusher | TCP position and orientation at 5 Hz; seventh value unused. |
| RoboDojo | 25 | Isaac Sim / dual X5 or dual xArm7 | 14 or 16 joint/gripper values at 25 Hz. |
| Artifact | Required content |
|---|---|
| protocol.json | Version 3.0 manifest: target configuration, scene or scene-file reference, camera, robot, action-file references, timestep, and provenance. |
| actions.npy | Finite control matrix with the embodiment-specific width in Table A2 ; Cartesian streams are stored as 64-bit floating-point arrays. |
| Additional hand arrays | For paired hands, a second palm stream and two finger-angle streams, with four distinct file references. |
| scene.json / scene.xml | Native Isaac submissions use a separate scene description; MuJoCo submissions use an MJCF scene with the supplied robot and explicit object, robot, and target roles. |
| assets/ | Locally referenced visual and collision geometry. Native mesh submissions provide vertices and face indices, with metric scale incorporated into the vertices. |
| expected/ obj_poses.npy | Camera-frame object-pose estimates for the camera-frame rigid and human-video profiles. These are declarations, not commands; native joint and MuJoCo interfaces do not require them. |
| Capability | Agent-specified input | Agent-visible feedback |
|---|---|---|
| Shell execution | Program or command; optional execution limit | Standard output, execution errors, exit status, and timeout information. |
| File and image reading | File or directory; optional text range | File contents, directory entries, or an image observation. |
| Filename search | Filename pattern and search location | Matching file paths. |
| Text search | Text pattern and search scope | Matching passages and file locations. |
| File creation | File location and contents | Confirmation, diagnostics, or a write error. |
| Text replacement | File and requested replacement | Modification result or an error. |
| Operation | Input | Agent-visible feedback | Effect on the episode |
|---|---|---|---|
| Frame extraction | Source video and sampling specification | Frame references, timestamps, and image metadata | Makes frames available for inspection; the candidate is unchanged. |
| Physical replay | Candidate scene and robot controls | Rendered rollout and available execution diagnostics | Supports inspection and revision of the reconstruction; does not submit or score it. |
| Validation | Candidate reconstruction and task configuration | Missing artifacts and violations of the public submission requirements | Construction continues; no candidate is submitted. |
| Checkpointing | Candidate reconstruction | Validation feedback and confirmation of preservation | A valid candidate is retained as a fallback during the permitted refinement. |
| Submission | Candidate reconstruction | Acceptance or validation errors | Acceptance freezes the final candidate and ends generation; an invalid submission consumes one opportunity. |
| Elapsed time | Construction and delivery policy |
|---|---|
| About 30 min | Complete an initial evidence-based candidate; validate it and repair missing or malformed artifacts. |
| 120 min | Restrict remaining work to completing and repairing delivery. |
| 150 min | Close the optional refinement allowance. |
| 165 min | Stop generation and new delivery requests; select the latest valid protected checkpoint if no final submission has been accepted. |
| 180 min | Complete process termination, integrity verification, and publication of the selected candidate. |
| Model | Model identifier | Provider | Effort |
| DeepSeek V4.1 Flash | deepseek/deepseek-v4.1-flash | novita/fp8 | – |
| GLM-5.3 Flash | z-ai/glm-5.3-flash | novita/fp8 | – |
| Kimi K3 | moonshotai/kimi-k3 | moonshotai/mxfp4 | – |
| GPT-5.6 | openai/gpt-5.6-sol | openai | – |
| Gemini 3.8 Flash | google/gemini-3.8-flash | google-vertex/ global | High |
| Claude Opus 5 | anthropic/claude-opus-5 | anthropic | – |
| Video source | Clips | Instances | Target embodiment |
|---|---|---|---|
| FurnitureBench | 64 | 64 | Panda |
| DROID | 20 | 20 | Panda/Robotiq (18); cloth gripper (2) |
| RoboDojo | 25 | 25 | Dual X5 or dual xArm7 |
| HOI4D | 9 | 18 | Panda; paired Wuji hands |
| HOT3D | 7 | 14 | Panda; paired Wuji hands |
| DexYCB | 11 | 22 | Panda; paired Wuji hands |
| Source | Available source evidence | Benchmark annotation work |
|---|---|---|
| FurnitureBench | CAD, tags, calibrated cameras, and robot logs. | Register part poses; fit grasp offsets to bridge occlusions during verified holding. |
| DROID | Calibrated RGB-D, multiple views, and robot kinematics. | Fit object geometry and poses; reconstruct supports and cavities; track visible cloth surfaces. |
| RoboDojo | Native assets and synchronized simulation states. | Export geometry, cameras, and trajectories from the same execution as the video. |
| HOI4D | Instance CAD, object boxes, object poses, and camera trajectories. | Compose object and camera poses in a gravity-aligned frame; model supporting surfaces. |
| HOT3D | Object CAD and tracked object/camera poses. | Register trajectories; identify supports and static/dynamic roles. |
| DexYCB | YCB meshes, calibrated views, and object/hand poses. | Convert pose units and frames; estimate support height from initial mesh bottoms. |
| Source | Additional annotations and assets | Use in reference construction |
|---|---|---|
| FurnitureBench | Part CAD, calibrated part/robot motion, assembly targets, simulator state. | Reuse part and collision assets; initialize grasps and transport; refine contact, alignment, and release. |
| DROID | Fitted object/support geometry, calibration, robot kinematics, object motion; visible cloth surfaces. | Initialize scene and controls from source records; fit object-guided motion and refine execution. The two cloth references retain their declared assisted grasp mechanism. |
| RoboDojo | Native scenes, robot and object assets, task rules, recorded controls and states. | Adapt the recorded native task execution to the benchmark interface. |
| HOI4D | Instance CAD, object and camera trajectories, fitted support, available hand observations. | Retarget object transport to the arm or hands; optimize contact using annotated motion. |
| HOT3D | Object CAD, tracked object/camera motion, available hand poses, support estimates. | Initialize grasp and transport, then refine wrist/palm and finger controls. |
| DexYCB | YCB meshes, calibrated views, object/hand poses, estimated support height. | Reuse geometry and pose observations to construct and refine lifting behavior. |
| Category | Criterion | Count | Share |
|---|---|---|---|
| HAR pass | — | 113 | 50.9% |
| A1 Physical-validity rejection | valid=false | 19 | 8.6% |
| A2 No passing execution in family | family HAR | 18 | 8.1% |
| B Near miss | margin tolerance | 7 | 3.2% |
| C Controller failure | valid, family has passes | 65 | 29.3% |
| Method | S0 all (215) | S2 primary (179) | S1 HAR-pass (112) |
|---|---|---|---|
| Human-assisted Ref | 52.1 | 71.6 | 100.0 |
| Fable-5.1 | 19.1 | 20.6 | 24.2 |
| Claude Opus 5 | 12.1 | 11.7 | 14.1 |
| GPT-6 Astra | 8.1 | 9.8 | 9.7 |
| Kimi K3 | 3.6 | 4.2 | 3.7 |
| DeepSeek V4.1 Flash | 3.0 | 2.3 | 2.4 |
| Cell | Behavior | Scene |
| Own success | agent actions | agent scene |
| HAR actions agent scene | HAR actions, translated, open-loop | agent scene |
| Agent HAR scene (submitted) | agent actions, unchanged | HAR scene |
| Agent HAR scene (anchored) † | agent actions, XYZ-translated | HAR scene |
| HAR HAR | HAR actions | HAR scene |
| Probe scene | state-aware scripted controller | either scene |
| Set | Cell | Pooled | Macro | Rows |
| A (20 inst.) | Own success | 24.3 [14.9, 34.6] | 25.9 | 177 |
| Probe agent scene | 13.1 [6.3, 20.8] | 7.3 | 175 | |
| HAR actions agent scene | 12.1 [5.8, 19.5] | 6.9 | 174 | |
| Agent HAR scene (submitted) | 2.3 [0.6, 4.5] | 1.4 | 174 | |
| Agent HAR scene (anchored) † | 15.5 [7.3, 24.6] | 8.7 | 174 | |
| Probe HAR scene | 100 (condition) | 100 | 180 |
| Set A (172 rows) | Set B (385 rows) | ||||
|---|---|---|---|---|---|
| Own | Scene / Behavior test | rows | share | rows | share |
| Success | ✓/ ✓ | 10 | 23% | 12 | 19% |
| ✓/ | 7 | 16% | 7 | 11% | |
| / ✓ | 7 | 16% | 14 | 22% | |
| / | 19 | 44% | 30 | 48% | |
| Failure | ✓/ ✓ | 2 | 2% | 4 | 1% |
| Model | Own | Probe agent | HAR agent | Agent HAR (subm.) | Agent HAR (anch.) † |
|---|---|---|---|---|---|
| Fable-5.1 | 70.0 | 40.0 | 25.0 | 5.0 | 35.0 |
| Claude Opus 5 | 42.1 | 21.1 | 21.1 | 5.3 | 10.5 |
| GPT-6 Astra | 35.0 | 5.0 | 30.0 | 5.0 | 25.0 |
| GPT-5.6 Sol | 15.0 | 15.8 | 10.0 | 0.0 | 15.0 |
| DeepSeek V4.1 Flash | 15.0 | 5.0 | 5.3 | 0.0 | 15.8 |
| Gemini 3.8 Flash | 15.0 | 15.0 | 10.0 | 5.0 | 0.0 |
| API cost (USD) | Agent time | API calls | Eval. GPU | ||||
|---|---|---|---|---|---|---|---|
| Method | mean | median | max | est. total | mean (min) | mean | mean (min) |
| Fable-5.1 † | 14.03 | 12.90 | 45.19 | 3,115 | 37.5 | – | 10.1 |
| GPT-6 Astra † | 6.40 | 3.91 | 24.58 | 1,421 | 14.6 | – | 9.7 |
| Claude Opus 5 | 13.83 | 10.03 | 61.07 | 3,070 | 32.6 | 59 | 8.7 |
| Kimi K3 | 5.51 | 4.81 | 13.27 | 1,223 | 82.2 | 72 | 7.4 |
| GPT-5.6 Sol | 1.39 | 1.00 | 7.89 | 309 | 3.7 | 17 | 8.2 |