World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions. Magic-W0 represents interaction as a Structured World Transition consisting of Current State, Transition, and Future State. Current State combines vision-language context with Current 3D Geometry; Transition is represented by 3D Motion capturing action-induced three-dimensional changes; and Future State is represented by Future Semantics describing task-relevant outcomes. To couple prediction and control, we propose a layer-aligned world-action interaction architecture in which evolving action hypotheses condition world-transition prediction, while predicted world representations continuously inform action generation. Magic-W0 is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision for geometry, 3D motion, and future semantics from pre-trained visual models. Inference-time interventions show that structured world representations respond systematically to changes in candidate actions and that action-related information propagates through shared 3D representations into future semantic predictions. On RoboDojo-Sim, Magic-W0 achieves an average Score of 27.10, the highest among the compared WAMs. Across multiple real-robot tasks, it also demonstrates strong downstream performance after fine-tuning with limited downstream data, supporting generalization and rapid adaptation.
Figures & tables
Figure 1 : Overview of Magic-W0, a large-scale, cross-embodiment structured world–action foundation model for physical intelligence. The model jointly learns physical state transitions and continuous control from diverse embodied experience.
Figure 2 : Comparison of model architectures. (a) Action-centered VLAs primarily rely on implicit world understanding. (b) Pixel-based WAMs predict future observations alongside actions. (c) Latent WAMs predict generic future representations alongside actions. (d) Magic-W0 constructs Structured World Transition from Current 3D Geometry, 3D Motion, and Future Semantics, with bidirectional interaction through layer-aligned joint attention and an action expert.
Figure 3 : Overall framework of Magic-W0. Structured World Transition comprises Current State, Transition, and Future State. Current State combines vision–language context with Current 3D Geometry; Transition is represented by 3D Motion; Future State is represented by Future Semantics. Current 3D Geometry and 3D Motion share the hidden states of the 3D stream, while a separate semantic stream models Future Semantics. Both streams interact with the action expert through layer-aligned joint attention, with vision–language context supplied by the VLM. Track4World and DINOv3 provide latent supervision only during training.
Figure 4 : Layer-aligned world–action interaction. Each block contains three within-stream Gated DeltaNet layers and one attention layer. Six repetitions yield 24 aligned depths. The VLM retains causal computation at attention layers, while the semantic stream, 3D stream, and action expert access joint keys and values through joint attention. In the visibility matrix, rows indicate query sources and columns indicate key–value sources. Blue denotes visible connections; white denotes masked connections. Orange dashed cells in the Act row and Sem/3D columns indicate connections that can be randomly masked during training. Only one is selected per masking event; semantic and 3D query visibility remains unchanged.
Figure 5 : Approximate numbers of valid manipulation episodes used for pre-training, aggregated by acquisition domain. Appendix Table 3 reports dataset-level counts. Both individual counts and the total are approximate.
Figure 6 : Example object and action words in trajectories with task descriptions. Font size reflects mention frequency in the text.
Figure 7 : Slot occupancy in the 34-dimensional state–action interface. q denotes joint angles, g gripper opening, p end-effector position, and r the 6D rotation encoding. Subscripts L and R distinguish sides; blank slots are zero-padded and masked. Egocentric human data have no corresponding robot embodiment, so all 14 joint slots are invalid. The example UMI, real-robot, and simulation trajectories share the same occupancy and are grouped in one row. Each uses two arms with six degrees of freedom, leaving the seventh joint slot of each arm empty.
Figure 8 : Representative RoboDojo-Sim tasks. The five panels illustrate Generalization, Memory, Precision, Long-Horizon, and Open capabilities. Frames are taken from official RoboDojo task demonstrations [ 9 , 46 ] .
Method
Type
Open- source
Generalization
Precision
Long- Horizon
Memory
Open
Average
GPT-6-Astra [ 41 ]
Agent
No
33.36/30.50
12.65/4.00
21.45/8.25
43.04/38.67
34.36 /31.00
28.97/22.48
VLAct [ 45 ]
VLA
Yes
9.54/6.28
20.57/15.17
20.12/13.67
0.66/0.56
2.37/2.25
10.65/7.58
StarVLA-PI_v3 [ 45 ]
VLA
Yes
11.22/8.05
17.77/12.50
18.46/11.00
4.59/4.00
2.03/2.00
10.81/7.51
InternVLA-A1.5 [ 45 ]
VLA
Yes
10.35/6.83
15.23/10.17
23.80/13.75
4.93/3.56
1.43/1.42
11.15/7.14
π0.5 [ 45 ]
VLA
Yes
13.38/8.17
12.40/5.50
23.54/14.67
5.89/4.67
1.98/1.67
11.44/6.93
Spatial Forcing [ 45 ]
VLA
Yes
14.12/9.34
17.32/10.58
23.26/14.58
5.43/4.11
1.78/1.58
12.38/8.04
Table 1 : RoboDojo-Sim comparison. Agent results precede the VLA and WAM groups. Horizontal rules separate groups and distinguish Magic-W0. VLA and WAM entries are sorted by ascending Average Score, with Magic-W0 placed last in the WAM group. Yes in the Open-source column means that both code and model weights are publicly available. Evaluation follows the official protocol. The highest Score in each column is bold.
Method
Spatial
Object
Goal
Long
Avg. SR
Octo [ 39 ]
78.9
85.7
84.6
51.1
75.1
OpenVLA [ 22 ]
84.7
88.4
79.2
53.7
76.5
SpatialVLA [ 44 ]
88.2
89.9
78.6
55.5
78.1
GR00T-N1 [ 38 ]
94.4
97.6
93.0
90.6
93.9
π0 +FAST [ 42 ]
96.4
96.8
88.6
60.2
85.5
π0 [ 4 ]
96.8
98.8
95.8
85.2
94.1
Table 2 : Comparison on LIBERO.
Figure 9 : Video keyframes and success rates for five real-robot tasks. Rows show six keyframes each for clothes folding, bottle uprighting, pen storage, object storage, and kitchen storage. Task SRs for π0.5 and Magic-W0 appear on the right. Row titles are English instructions composed from the task content and are not the original prompts used in the experiments.
Figure 10 : Effects of action replacement on structured world heads. (a) Normalized output changes Dm(τ) for the three heads. The horizontal reference line represents the mean output distance between samples within a batch; the dashed line shows the action expert velocity-field control. (b) Future Semantics loss before and after action replacement. Percentages are batch-averaged values of Lsem(x~τ)/Lsem(xτ)−1 . Larger τ indicates less ground-truth action information in xτ . Error bars denote standard deviations across batches.
Figure 11 : Effects of cross-stream connection masking on Future Semantics. (a) Semantic loss increases after masking different connections, averaged over all flow intervals. Percentages are relative to the full model. (b) Results for the same interventions at different flow times τ . Arrows indicate information flow; all cross-stream edges removed denotes simultaneous removal of all connections between expert streams. Error bars denote standard deviations across batches.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Data category
Dataset
Valid episodes
Egocentric human
EgoSuite
360k
EgoDex
330k
UMI
Hy-Embodied-0.5-VLA-Data
231k
Simulation
InternData-A1
374k
RoboTwin 2.0
73.2k
Real robot
AgiBotWorld-Beta
500k
Appendix
Table 3: Composition of the embodied pre-training corpus.
Data category
Dataset
Samples
Open-source datasets
EO-Data1.5M
1.42M
Robo2VLM-1
685k
In-house annotations
Annotations constructed from open-source datasets
500k
Appendix
Table 4: Composition of the vision–language supervised fine-tuning corpus.
Figure 12: End-effector pose projections from three source categories under a unified convention. Each source uses its own camera poses and intrinsics to project end-effector poses onto observations. UMI poses are obtained from tracked devices and inverse kinematics; real-robot poses from joint states and URDF; simulation poses directly from simulator records. Red, green, and blue indicate local x , y , and z axes, respectively. Figure 13 shows projections for egocentric human manipulation.
Figure 13: Raw UMI and egocentric human manipulation records contain no proprioceptive states for the target robot, requiring constructed supervision. UMI tracked poses yield bimanual joint angles and gripper openings through IK, while retaining end-effector poses. Human manipulation videos use hand keypoints to derive virtual end-effector poses, camera-frame trajectories, and continuous gripper openings.
Figure 14: Calibration when camera parameters are unavailable. Real-robot data fit bimanual projections using moving foreground regions and URDF geometry. UMI data extract device candidates in the tabletop region and optimize camera parameters using correspondences between tracked poses and images.
Check
Content and criteria
Action completeness
Dimensions, value ranges, frame-to-frame increments, and constant signals
Media quality
Video readability, image dimensions, mean grayscale intensity, dark-pixel proportion, and sharpness
Geometric visibility
Camera-frame end-effector depth and image-boundary constraints
Extreme values
Per-trajectory tolerance band [q.01−αΔq,q.99+αΔq] , where Δq=q.99−q.01 ; out-of-band values are soft candidates
Spikes
Smoothing residuals and second-/third-order differences assessed against their respective robust scales; both the residual and at least one higher-order difference must exceed thresholds
End-effector continuity
Linear and angular velocities between adjacent poses, and rotation representation validity
Appendix
Table 5: Offline quality checks, listed in execution order. Checks and thresholds are configured by data source.
Figure 15: Quality-check examples. (a) Smoothing residuals and second-/third-order differences relative to their respective robust scales. (b) State–action lag. (c) Quantile tolerance bands. (d) End-effector projections. (e) Frame-to-frame joint increments and flagged stationary intervals. (f) End-effector linear velocity. The 4× and 2m/s thresholds are illustrative references; actual criteria depend on the source. The right side of each panel’s title bar indicates where its output is used. Spikes and end-effector discontinuities are hard flags repaired by interpolation where enabled. Quantile bands produce soft flags used only in trajectory-level scoring. State–action lag is diagnostic only. End-effector visibility and boundary stationarity flags enter anchor gating and timestep masks.