STRIKE: Learning Visual State Transitions for Physical World Modeling
Organizations: Applied Intuition · University of Southern California · University of California, Berkeley
Abstract
Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pairing them with transition descriptions and temporal offsets. An image-based transition model learns to predict the next scene configuration from the current image, a local transition specification, and elapsed time. At inference, a pretrained vision-language planner predicts time transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A separately trained dynamic model then generates the complete rollout conditioned on these states and their temporal locations. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements of STRIKE over the corresponding video-backbone baselines in benchmark measures of physical consistency and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.
Figures & tables
| Model | # params | EFLOPs | Score | S-IoU | ST-IoU | WS-IoU |
|---|---|---|---|---|---|---|
| Wan2.2 ( Wan Team, 2025 ) | 14B | 49.100 | 34.2 | 50.3 | 25.2 | 29.9 |
| Cosmos-Predict-2.5 ( Ali et al., 2025 ) | 2B | 2.560 | 32.2 | 45.5 | 27.2 | 30.1 |
| CogVideoX ( Yang et al., 2025b ) | 5B | 6.640 | 30.5 | 40.4 | 30.3 | 24.1 |
| Cosmos3-Nano ( NVIDIA, 2026a ) | 16B | 11.802 | 27.6 | 39.7 | 20.2 | 22.1 |
| CoECT ( Wang et al., 2026 ) | 20B+5B | 29.600 | 19.9 | 27.6 | 23.2 | 14.8 |
| Wan2.2-5B ( Wan Team, 2025 ) | 5B | 2.739 | 19.4 | 25.0 | 19.4 | 16.3 |
| Method | # param | Mechanics | Optics | Thermal | Material | Average |
|---|---|---|---|---|---|---|
| VideoCrafter2 ( Chen et al., 2024 ) | 1.4B | 40.00 | 58.00 | 28.89 | 34.17 | 42.08 |
| LaVie ( Wang et al., 2025b ) | 0.91B | 29.17 | 50.67 | 23.33 | 32.50 | 35.63 |
| Open-Sora v2 ( Peng et al., 2025 ) | 11B | 55.00 | 66.67 | 47.78 | 50.83 | 56.25 |
| LTX-Video ( HaCohen et al., 2024 ) | 13B | 45.83 | 65.33 | 43.33 | 37.50 | 49.37 |
| DreamWorld ( Tan et al., 2026 ) | 1.3B | 54.17 | 64.67 | 51.11 | 43.33 | 54.17 |
| Cosmos3-Nano ( NVIDIA, 2026a ) | 16B | 57.50 | 75.33 | 48.89 | 46.67 | 58.75 |
| Method | L2 | CD | IoU |
|---|---|---|---|
| CausalMotion | 0.1292 | 0.3435 | 0.0966 |
| CogVideoX-5B | 0.1180 | 0.3017 | 0.1481 |
| Cosmos-Predict2.5 | 0.1373 | 0.3871 | 0.1467 |
| Cosmos3-Nano | 0.1250 | 0.3437 | 0.1361 |
| LaMo-5B | 0.1113 | 0.2888 | 0.1590 |
| LTX-Video | 0.1181 | 0.3071 | 0.1017 |
| Method | EWMScore | Trajectory | Interaction | Perspectivity | Instruction | Semantic |
|---|---|---|---|---|---|---|
| OpenDW | 44.26 | 0.0220 | 0.2556 | 0.6096 | 0.2080 | 0.8685 |
| Ctrl-World | 60.16 | 0.2217 | 0.5268 | 0.7776 | 0.4960 | 0.8782 |
| CogVideoX-5B | 58.04 | 0.2301 | 0.5432 | 0.7952 | 0.5180 | 0.8934 |
| Wan2.2-5B | 59.16 | 0.2372 | 0.5088 | 0.8040 | 0.4560 | 0.8866 |
| Ours (CogVideoX-5B) | 60.51 | 0.3306 | 0.5900 | 0.8160 | 0.5916 | 0.8975 |
| Ours (Wan2.2-5B) | 62.53 | 0.3368 | 0.6012 | 0.8448 | 0.6012 | 0.8943 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Type | Size |
| Transition-model supervision | ||
| WISA-80K | Real | 59,298 episodes |
| NVIDIA PhysicalAI | Simulation | 39,873 episodes |
| PhyCo/Kubric | Simulation | 31,109 episodes |
| PICA-100K | Synthetic | 105,085 image pairs |
| Conditional video-dynamics supervision | ||
| You are labeling transition state frames for training a physics-aware video generation model. You will receive {N} storyboard images, one per sample. Each storyboard image contains labeled candidate frames. For each sample, do this in order: 1. Determine whether the video contains a visible physical state change. 2. Identify the main objects, scene elements, and main physical or visual process. 3. Infer event_structure using a concise general label, such as single_transition, sequential_transitions, repeated_motion, camera_motion, scene_change, or unclear. 4. Write ordered_subevents as the minimal ordered list of visible subevents. Merge simultaneous changes into one subevent and split clearly sequential changes into separate subevents. 5. Build state_plan around moments that matter for transition state selection. State descriptions should mention visible states or relations such as onset of motion, contact or collision, deformation, liquid or granular flow, opening or closing, placement, stacking, release, stable result, camera reveal, or view change. 6. Reject clips that are mostly static and have no useful visible state/view progression. Return strict JSON only with one item per sample_id. |
| Solid–Solid | Solid–Fluid | Fluid–Fluid | Overall | |||||||
| Method | # params | EFLOPs | SA | PC | SA | PC | SA | PC | SA | PC |
| VideoCrafter2 | 1.4B | 0.441 | 27.3 | 25.2 | 47.3 | 24.7 | 49.1 | 30.9 | 39.2 | 25.9 |
| Open-Sora v2 | 11B | 9.471 | 53.9 | 13.3 | 72.6 | 24.0 | 65.5 | 27.3 | 63.7 | 20.2 |
| LTX-Video | 13B | 1.379 | 39.9 | 17.5 | 70.6 | 22.6 | 49.1 | 25.5 | 54.4 | 20.9 |
| DreamWorld | 1.3B | 9.830 | 58.7 | 12.6 | 74.0 | 27.4 | 70.9 | 32.7 | 67.2 | 22.1 |
| Cosmos3-Nano | 16B | 14.033 | 67.1 | 13.3 | 82.2 | 23.3 | 78.2 | 25.5 | 75.3 | 19.5 |
| Select transition state frames from the candidate frames for each sample. Decision procedure: • First read the caption, source metadata, event_plan.event_structure, event_plan.ordered_subevents, and event_plan.state_plan. • Choose transition state frames for distinct visible states in the main event, not for generic beginning/middle/end coverage. • For a single clear transition, choose the earliest candidate where the defining changed state is visible; include an earlier setup state only when it is needed to understand the change. • For sequential transitions, select one frame per visible subevent at the earliest candidate where that subevent’s state is clearly established. • For contact, collision, attachment, containment, placement, stacking, opening, or closing, choose the first candidate where the relation is visibly established, not a later redundant proof frame. • For deformation, breakage, spilling, pouring, smoke, fire, liquid, granular motion, or other continuous processes, cover the onset and the most informative result or peak state. Add an intermediate keyframe only when the process has a materially different middle state. • For object motion, choose frames where the object’s state or relation has changed meaningfully. Avoid frames that differ only by a small position shift unless that shift completes the event. • For camera motion without object state change, choose frames that capture distinct viewpoints or newly revealed content. Rules: • Choose only candidate_id values in the sample’s candidate table. • Do not select static initial frames, near-duplicates, or redundant middle frames unless they represent a distinct visible state. • A transition state frame should be an event-defining visual state: onset, contact, relation established, state transformed, view changed, peak action, or stable result. • Selected frames must be chronological. • Each selected frame must represent a distinct event-relevant visible state or relation. • For each transition state frame provide: slot, candidate_id, frame_index, time_sec, state_text, selection_reason, confidence. • Also provide transition_text between adjacent transition state frames. • In coverage_notes, explicitly write: event_structure=<label>; selected_subevents= <comma-separated selected visible states>; omitted= <why any expected state was not selected or ‘none’>. |
| System prompt You are a strict quality inspector for a physics-focused video dataset. Judge only what is actually visible in the given frame(s). Be skeptical: if the described state is not clearly visible, answer false. Always reply with a single JSON object and nothing else. User prompt for the first transition state frame This image is transition state frame {slot_pos} of {total} extracted from a video. Video context: physical category = ‘‘{label}’’; caption = ‘‘{caption}’’ The transition state frame annotation claims the scene has reached this state: ‘‘{state_text}’’ Question: does the image show that the scene has actually reached the claimed state? Judge the physical state of the scene, not image style or camera work. Reply with JSON only: {{‘‘match’’: true or false, ‘‘confidence’’: 0.0--1.0, ‘‘reason’’: ‘‘<at most 20 words>’’}} User prompt for a subsequent transition state frame You are shown two transition state frames from the same video. Image 1 is the PREVIOUS transition state frame (reference only). Image 2 is the CURRENT transition state frame {slot_pos} of {total}. Video context: physical category = ‘‘{label}’’; caption = ‘‘{caption}’’ The annotation claims that by the CURRENT transition state frame the scene has reached this state: ‘‘{state_text}’’ Question: comparing against image 1 where the claim describes change or motion, does image 2 show that the scene has actually reached the claimed state? Judge the physical state of the scene, not image style or camera work. Reply with JSON only: {{‘‘match’’: true or false, ‘‘confidence’’: 0.0--1.0, ‘‘reason’’: ‘‘<at most 20 words>’’}} |
| Method | # param | Mechanics | Optics | Thermal | Material | Average |
|---|---|---|---|---|---|---|
| Open-Sora v2 ( Peng et al., 2025 ) | 11B | 55.00 | 66.67 | 47.78 | 50.83 | 56.25 |
| LTX-Video ( HaCohen et al., 2024 ) | 13B | 45.83 | 65.33 | 43.33 | 37.50 | 49.37 |
| Cosmos3-Nano ( NVIDIA, 2026a ) | 16B | 57.50 | 75.33 | 48.89 | 46.67 | 58.75 |
| CausalMotion ( Zhuang et al., 2026 ) | 13B | 73.33 | 75.33 | 65.56 | 64.17 | 70.21 |
| Ours (CogVideoX-2B-T2V) | 20B+2B | 60.00 | 73.33 | 63.33 | 66.04 | 65.68 |