KineWorld: Action-Induced Transport Fields for Embodied World Modeling
Authors: Ziying Song, Yuchen Liu, Zhuoran Xu, Ziyang Liu, Jian Jin, Jiangtao Su, Haibao Yu, Lei Yang, +1 more
Organizations: Nanyang Technological University · North University of China · Psibot · The Hong Kong Polytechnic University · China Academy of Information and Communications Technology · University of Hong Kong · Peking University
Embodied world models predict the visual consequences of candidate actions before execution. However, existing action-conditioned world models often adopt uniformly weighted visual generation objectives that can be misaligned with embodied prediction needs. Even with explicit motion conditioning, these objectives can underemphasize spatially sparse changes that are critical to interaction. We propose KineWorld, a transport-aware world-modeling framework that extends robot kinematics from motion conditioning to the spatial allocation of generative supervision. Kinematic Transport Lifting (KTL) constructs renderer-derived, camera-aligned transport fields from commanded robot motion. Transport-Aware World Diffusion (TAWD) calibrates their motion support on the video-latent grid and reweights future-RGB flow matching through a normalized mixture of uniform and transport-focused distributions. We train KineWorld using ALOHA-AgileX bimanual manipulation data from RoboTwin 2.0. KineWorld achieves an EWMScore-P of 68.95 in single-view evaluation and a TWB-Score of 54.82 in multi-view evaluation. These results support a shift from appearance fitting toward action-consequence modeling for embodied decision-making.
Figures & tables
Figure 1: (a) Token-conditioned WAMs use numerical trajectories or learned action embeddings for video generation ( Zhu et al., 2024 ; Guo et al., 2025 ; Liao et al., 2025 ) , but typically apply uniform RGB supervision , potentially underemphasizing interaction-critical local changes . (b) Auxiliary-cue WAMs incorporate motion-focused auxiliary cues , such as masks ( Chen et al., 2026c ) , action images ( Zhen et al., 2026 ) , optical flow ( Chen et al., 2026b ) , or projected kinematic fields ( Yang et al., 2026 ) , to improve motion localization. These cues alone do not directly change the spatial weighting of the primary RGB objective, potentially leaving motion localization and supervision priorities misaligned . (c) KineWorld introduces action-guided RGB supervision by lifting camera-aligned kinematic transport onto the video latent grid and reweighting flow matching, enabling action-consequence-oriented supervision while preserving the total token weight.
Figure 2: Overview of KineWorld. Kinematic Transport Lifting (KTL) converts the commanded robot rollout into camera-aligned transport and injects its latent representation as a clean structural condition. Transport-Aware World Diffusion (TAWD) calibrates the transport support and derives mean-one token weights for the flow-matching objective, emphasizing action-responsive regions under a fixed weight budget. The weighting branch is used only during training, while the inherited action expert remains unchanged.
Method
Overall
Generation Quality
Embodied Capability
EWM↑
Vis.↑
Mot.↑
Cont.↑
Phys.↑
3D↑
Ctrl.↑
JF_World ( DreamX Team et al., 2026 )
64.85
66.52
30.66
56.91
69.27
97.88
88.06
BWM-Super ( BWM Team, 2026 )
64.30
67.20
30.16
58.36
64.15
97.35
87.19
BWM-Turbo ( BWM Team, 2026 )
63.96
66.99
30.10
57.22
64.46
97.74
86.05
FlowWAM-FiveAges ( Chen et al., 2026b )
63.87
66.60
29.92
54.47
65.97
98.48
88.09
WoVR_Plus ( Jiang et al., 2026 )
57.89
58.37
30.52
57.91
47.95
85.31
80.75
Table 1: Single-View Reference Profiles. EWM denotes EWMScore-P ( Shang et al., 2026 ) ; metric definitions and aggregation appear in Appendix A.3.1 . These aggregates do not constitute a shared RoboTwin 2.0 evaluation. Source status is documented in Appendix A.3 .
Method
Overall Fidelity
View PSNR (dB)
Boundary PSNR (dB)
PSNR↑
SSIM↑
N\mbox−MAE↓
Mot. Corr.↑
Head↑
Left↑
Right↑
First↑
Last↑
KineWorld (white wrist flow)
35.35
0.9622
16.532
0.9132
36.10
36.13
33.82
42.01
32.32
KineWorld (all-view RAFT)
42.76
0.9860
5.713
0.9953
41.88
43.28
43.13
42.65
42.76
Table 5: KTL Ablation. Three-view validation diagnostic on 100 episodes (300 videos per variant). Dataset scope and metric definitions appear in Appendix A.8.1 .
Figure 3: Retained qualitative comparison between Wan2.2 and KineWorld. Decoded episode-550 frames at t=1.0 , 2.0 , 3.0 , 4.0 , and 5.0 s from 24-fps videos. This is a descriptive rollout comparison, not an action- or checkpoint-matched ablation.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Partition
Episode indices per task
Episodes
Membership
Development training
0–35
1,800
Disjoint from validation
Validation
36–39
200
Included in combined training
Combined training
0–39
2,000
Development ∪ validation
Held-out
40–49
500
Disjoint from both training sets
Appendix
Table A.1: RoboTwin 2.0 Clean-50 Data Splits. Indices are local to each of the 50 tasks. Counts refer to unique episodes; combined training includes the validation partition.
Method
Overall
Visual
Motion
Content
Physics
3D
Control
EWM
AES
IMG
JEPA
DYN
FL
MS
BG
PC
SC
INT
TA
DEP
PER
IF
SA
JF_World ( DreamX Team et al., 2026 )
64.85
43.38
63.25
92.93
22.90
5.81
63.26
84.52
14.29
71.93
81.40
57.15
98.55
97.20
85.60
90.53
BWM-Super ( BWM Team, 2026 )
64.30
44.39
62.38
94.83
22.01
5.77
62.70
84.31
18.92
71.86
79.20
49.09
97.51
97.20
85.20
89.19
BWM-Turbo ( BWM Team, 2026 )
63.96
44.69
60.27
96.00
22.23
5.76
62.31
82.97
17.33
71.37
78.60
50.32
98.48
97.00
82.80
89.31
FlowWAM-FiveAges ( Chen et al., 2026b )
63.87
44.00
64.68
91.12
22.26
5.80
61.69
79.62
15.08
68.71
79.60
52.33
99.15
97.80
87.20
88.98
WoVR Plus ( Jiang et al., 2026 )
57.89
40.34
46.78
87.99
22.92
5.80
62.83
86.99
15.46
71.27
67.20
28.69
85.02
85.60
73.00
88.50
Appendix
Table A.2: Complete Single-View Reference Profiles. EWM : EWMScore-P; AES : Aesthetic Quality; IMG : Image Quality; JEPA : JEPA Similarity; DYN : Dynamic Degree; FL : Flow Score; MS : Motion Smoothness; BG : Background Consistency; PC : Photometric Consistency; SC : Subject Consistency; INT : Interaction Quality; TA : Trajectory Accuracy; DEP : Depth Accuracy; PER : Perspectivity; IF : Instruction Following; SA : Semantic Alignment. Original score scales are preserved. † denotes the author-supplied candidate profile.
Method
AES↑
IMG↑
BG↑
SC↑
PC↑
INT↑
PER↑
IF↑
DYN
FL
Wan2.2-TI2V-5B ( Wan Team, 2025 )
42.10
44.42
85.17
80.90
0.96
68.68
88.52
76.00
–
–
KineWorld
45.26
61.21
87.19
75.53
16.43
79.20
97.30
85.84
0.53
1.88
Appendix
Table A.3: Complete Single-View Local Diagnostics. Additional components for Table 4 ; DYN and FL retain the local evaluator scale.
Figure A.1: Single-view local diagnostic profiles. Reported AES–IF values from Table A.3 , retaining the normalized score scale and metric abbreviations in Table A.2 . Scores are compared within each metric. The profiles describe the reported local systems rather than a checkpoint-matched KTL or TAWD ablation.
Figure A.2: Conceptual illustration of the KineWorld pipeline. KTL renders robot states and uses RAFT to construct camera-aligned transport, which conditions the generator. TAWD detaches the transport path, aggregates calibrated activity on the latent grid, and converts it into a transport distribution πi and mean-one weights wi=Nπi for future-RGB flow-matching errors. TAWD changes training-time supervision only; it adds no inference steps or parameters.
Figure A.3: Conceptual visualization of transport-support construction. Two manipulation scenes illustrate camera views, rendered robot motion, schematic 3D fields, image-plane transport, and normalized support. KTL estimates transport using RAFT between rendered robot states rather than analytic 3D-field projection. The final column illustrates ρtr , not the mean-one loss weights w . These panels are illustrative rather than measured experiment outputs.
Figure A.4: Analytical illustration of TAWD token-weight allocation. For the constructed exponential support over N=1000 future RGB latent tokens at β=1 , (a) shows descending mean-one weights and (b) their cumulative share relative to uniform weighting. The top 10% receive 36.63% of the total token weight. The green curve shows the general upper bound (1+q)/2 for q>0 ; the cumulative share has C0=0 . These curves illustrate the weighting rule, not measured training statistics or loss/gradient shares.
Method
Appearance
Motion
Consistency
AES↑
IMG↑
DYN↑
FL↑
MS↑
BG↑
PS↑
SC↑
Linear expansion
41.86
45.43
22.47
0.37
58.75
85.00
1.02
81.30
VFIMamba (no TTA)
40.30
45.50
41.65
8.37
63.43
90.50
54.22
86.97
RIFE HDv3
39.53
45.95
43.23
10.85
63.47
90.42
65.62
87.13
Appendix
Table A.4: Interpolation Strategy. Diagnostic on the same 16 episodes with fixed generated keyframes. Linear expansion uses linear frame interpolation; VFIMamba and RIFE HDv3 use learned interpolation.
Figure A.5: Interpolation diagnostics on 16 local episodes. The methods expand the same generated keyframes; values match Table A.4 . AES, IMG, DYN, FL, MS, BG, and SC follow Table A.2 ; PS denotes Photometric Smoothness. This comparison evaluates temporal reconstruction with the generator fixed, rather than the isolated effect of TAWD.
Temporal Variant
Local Motion Metrics (%)
DYN↑
FL↑
PS↑
Keyframe-6 (6 fps)
45.80
12.20
26.40
Linear expansion
22.47
0.37
1.02
Keyframe hold-4
33.09
9.69
13.78
VFIMamba (no TTA)
41.65
8.37
54.22
RIFE HDv3
43.23
10.85
65.62
Appendix
Table A.5: Temporal Reconstruction. Comparison on the same 16 local episodes with the generator fixed. Keyframe-6 reports the sparse generated keyframes directly at 6 fps, whereas the remaining variants produce 24-fps outputs. Bold marks the best results among the 24-fps reconstruction variants. DYN , FL , and PS denote Dynamic Degree, Flow Score, and Photometric Smoothness.
Method
Pixel Difference
Repetition
Optical Flow
A\mbox−MAD
CPF
FL\mbox−MAD
NDP
FB Mean
FB P95
Jerk
Linear expansion
0.006884
0.049167
0.096238
0.588374
0.109187
0.418630
0.026805
VFIMamba (no TTA)
0.008584
0.038254
0.096187
0.570205
0.324903
1.399471
0.078122
RIFE HDv3
0.008832
0.039097
0.096191
0.563069
0.337443
1.428315
0.075933
Appendix
Table A.6: Frame Diagnostics for the 16-episode interpolation ablation. Definitions, interpretation, and metric provenance are provided in Appendix A.8.3 .
Method
Paired Local VLM Metrics (%)
INT↑
PER↑
IF↑
Linear expansion
67.40
87.40
73.40
RIFE HDv3
66.60
87.80
72.80
RIFE − Linear (pp)
−0.80
+0.40
−0.60
Appendix
Table A.7: Representative-100 paired comparison on identical episodes. INT , PER , and IF denote Interaction Quality, Perspectivity, and Instruction Following. Values are local VLM-judge percentages.
Metric
RIFE versus Linear (episodes)
Improved
Tied
Regressed
Interaction Quality
8
80
12
Perspectivity
8
86
6
Instruction Following
12
73
15
Appendix
Table A.8: Paired Outcome counts for RIFE versus linear expansion on Representative-100. Each row sums to 100; ties denote unchanged VLM-judge scores.
Method
Measured Time (s)
Relative Efficiency
Mean↓
Total↓
Speedup↑
VFIMamba (no TTA)
1168.24
18691.88
1.00 ×
RIFE HDv3
6.02
96.32
194.07 ×
Appendix
Table A.9: Interpolation Efficiency measured on 16 episodes. Mean is per-episode time, Total is summed accelerator time, and Speedup is relative to VFIMamba.
Flow\mbox−Condition Scale
Local Motion Diagnostics
DYN↑
FL↑
1.00
0.224695
0.003716
1.25
0.227275
0.003652
1.50
0.235240
0.003520
Appendix
Table A.10: Inference-Time Flow-Condition Scale on 16 local episodes. DYN and FL denote Dynamic Degree and Flow Score. This is an auxiliary inference diagnostic rather than a KineWorld component ablation.
A. Tri-view consistency components
Method
NPSNR↑
SSIM↑
VLM-I↑
VLM-II↑
VLM-III↑
VQA↑
BWM
73.43
87.47
85.11
83.94
95.31
65.94
WoVR_Plus
74.62
87.91
83.79
83.94
94.84
69.11
DreamDojo
59.78
71.58
77.14
73.26
88.19
47.82
Motus
60.13
76.31
63.68
67.56
87.22
45.33
Genie Envisioner
58.74
74.44
69.40
57.13
77.93
36.72
Appendix
Table A.11: Complete retained multi-view component profiles. All seven methods and numerical entries are preserved from the supplied comparison; these are external aggregates, not verified RoboTwin 2.0 local re-evaluations. The normalized score scale and preference direction are unchanged; NPSNR is not PSNR in dB. Bold denotes the largest retained values at the reported precision, not a verified local ranking.
Figure A.6: Additional qualitative rollout comparisons (1/3). Wan2.2 and KineWorld outputs for retained episode identifiers 45, 125, 212, 300, and 336. Each row shows the decoded input followed by five uniformly spaced rollout frames.
Figure A.7: Additional qualitative rollout comparisons (2/3). Wan2.2 and KineWorld outputs for retained episode identifiers 391, 488, 586, 704, and 732, using the same deterministic frame-sampling rule as Figure A.6 .
Figure A.8: Additional qualitative rollout comparisons (3/3). Wan2.2 and KineWorld outputs for retained episode identifiers 780, 865, 938, 976, and 982, using the same deterministic frame-sampling rule as Figure A.6 .
Figure A.9: Episode 1, bottle manipulation. Columns are Left wrist, Head, and Right wrist; within each cell, the left image is GT and the right image is KineWorld. Eleven uniformly sampled frame indices are shown.
Figure A.10: Episode 50, bimanual food placement. Columns are Left wrist, Head, and Right wrist; within each cell, the left image is GT and the right image is KineWorld. Eleven uniformly sampled frame indices are shown.
Figure A.11: Episode 100, switch interaction. Columns are Left wrist, Head, and Right wrist; within each cell, the left image is GT and the right image is KineWorld. Eleven uniformly sampled frame indices are shown.