Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling
Authors: Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu, Yifan Sun, Xiancong Ren, Jianxin Sun, Yong Dai, +1 more
Organizations: Beijing Humanoid Robot Innovation Center · Beijing Institute of Technology · Harbin Institute of Technology, Shenzhen · The University of Hong Kong · China University of Mining & Technology, Beijing
Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs. End-effector visualizations provide an alternative but do not specify the full articulated configuration needed for robot execution. We present Dream4ACT, a world model built for joint video-action modeling across embodiments. To unify action representations across embodiments, we introduce a shared visual action interface, called action views, which render target joint configurations from four prescribed virtual cameras using URDF-based forward kinematics. This shared visual representation preserves embodiment-specific articulated geometry while allowing observation and action sequences to share a video autoencoder and diffusion transformer. Through masked flow-matching, our model supports forward dynamics, inverse dynamics, and joint observation--action generation within a single jointly trained model by varying which future sequences are corrupted. To recover executable action sequences from predicted action views, we propose a training-free, URDF-constrained multiview recovery mechanism, without a learned embodiment-specific decoder. Dream4ACT achieves an average success rate of 88.98% on RoboTwin~2.0 and an overall score of 65.66 on TriWorldBench, supporting effective closed-loop manipulation and competitive action-conditioned multiview prediction through the visual action interface.
Figures & tables
Figure 1 : Overview of Dream4ACT. Top: A shared action-view interface supports forward dynamics, inverse dynamics, and joint generation, with training-free recovery of joint targets. Bottom left: Examples of multi embodiments real-world task execution. Bottom right: Success rates and scores show strong performance in both robotic manipulation and action-conditioned video generation.
Figure 2 : Architecture of Dream4ACT. (a) URDF-based rendering produces four action views. (b) A shared video VAE and diffusion transformer model RGB observations and action views with VLM-derived semantic conditioning, here we show the joint generation mode. (c) Training-free multiview action solver recovers joint targets from predicted views.
Sample candidate and render corresponding action-views using Equation 4 .
Calculate score Et′ using Equation 16 .
Table 2 : Algorithm for training-free action recovery.
Method
Clean
Randomized
Avg.
π0.5 ( Intelligence et al., 2025 )
82.74
76.76
79.80
Motus ( Bi et al., 2026 )
88.66
87.02
87.80
LingBot-VA ( Li et al., 2026b )
92.90
91.50
92.20
Fast-WAM ( Yuan et al., 2026 )
91.88
91.78
91.80
Ours
90.50
87.46
88.98
Table 3: Success rates (%) on RoboTwin2.0 over 50 clean and 50 randomized tasks. Table 10 presents a comparison of success rates for each individual task.
Model
TVC
TA
P3D
MQ
TC
VQ
Overall
Ctrl-World ( Guo et al., 2026 )
57.42
43.72
34.67
29.28
46.17
16.64
38.98
Motus ( Bi et al., 2026 )
66.70
49.69
34.60
24.63
26.56
16.26
42.35
Genie Envisioner ( Liao et al., 2026 )
62.39
33.17
54.00
20.17
32.18
17.46
40.73
DreamDojo ( Gao et al., 2026 )
69.63
56.24
43.84
27.96
60.84
21.02
51.72
BWM ( Team and others, 2026 )
81.87
86.05
60.40
41.29
62.81
31.42
65.54
Ours
81.63
84.22
61.30
41.66
64.88
31.42
65.66
Table 4: TriWorldBench results . Dream4ACT is evaluated on the official 500-episode test set and reports six key indicators.
Robot
Clean
Random
Position (mm) ↓
Rotation ( ∘ ) ↓
Avg ↑
Aloha-Agilex
81.94
82.45
1.39
1.57
82.19
Piper
81.55
82.32
1.66
1.71
81.94
ARX-X5
82.19
80.26
1.22
1.95
81.23
Franka-Panda
64.52
62.45
4.11
9.45
63.48
UR5-Xsg
31.87
30.06
11.85
21.39
30.97
Table 5: Success rates (%) and end-effector recovery errors of the same jointly trained checkpoint across five embodiments on 31 shared RoboTwin 2.0 tasks. We use 25 trials per task and setting, Avg combines Clean and Randomized results. Recovery errors use GT action views with five held-out trajectories per task and 40 future frames per window, aggregated with equal task weighting.
Robot
Place_block
Wipe_plate
Stack_blocks
Storage_item
Common
All
Aloha-Agilex
85.0
90.0
20.0
50.0
87.5
61.3
TienYi2.5 Pro
90.0
80.0
30.0
65.0
85.0
66.3
Franka
90.0
85.0
–
–
87.5
–
UR5e
85.0
95.0
–
–
90.0
–
Table 6: Real-world success rates (%) using the same jointly trained checkpoint across four robot platforms. Each embodiment–task pair is evaluated over 20 trials under scene variations, including randomized background and randomized distractor objects. Common averages the first two tasks; All averages all tasks evaluated on each platform. “–” denotes a task not evaluated on that platform.
Figure 3 : Comparison of conditioning representations. The same recorded joint configurations are represented as camera-aligned skeleton projections or action-view renderings from four prescribed virtual cameras.
Method
DROID
PSNR ↑
SSIM ↑
LPIPS ↓
Skeleton rendering
23.19
0.899
0.105
Action-view rendering (Ours)
24.33
0.906
0.093
Table 7: RGB prediction fidelity on 50 held-out DROID trajectories from CLVR and RAD. Both variants condition on recorded joint states and predict 41-frame sequences.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Trainable adapter modulates main view and instruction information.
Hyperparameter
Value
Training Resource
16×Nvidia B200
Batch Size
16
Optimizer
AdamW
Weight Decay
1×10−2
Learning Rate
1×10−5
Learning Rate Scheduler
Cosine Scheduler
Appendix
Table 8: Training hyperparameters.
Figure 5 : Real-world robots hardware configuration. Franka and UR embodiment use one front camera and one wrist camera, while Aloha-Agilex and TienYi use one head camera and two wrist cameras.
Embodiment
Position (mm) ↓
Rotation ( ∘ ) ↓
Avg. success (%) ↑
Aloha-Agilex
1.39
1.57
82.19
Piper
1.66
1.71
81.94
ARX-X5
1.22
1.95
81.23
Franka-Panda
4.11
9.45
63.48
UR5-Xsg
11.85
21.39
30.97
Appendix
Table 9: Action recovery errors and task success across five embodiments on 31 shared RoboTwin 2.0 tasks. Avg. success averages Clean and Randomized evaluations, with 25 trials per task, per embodiment, and per setting.
Simulation Task
π0.5
Motus
LingBot-VA
Fast-WAM
Ours
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Adjust Bottle
100%
99%
89%
93%
90%
94%
100%
100%
91%
97%
Beat Block Hammer
96%
93%
95%
88%
96%
98%
99%
97%
100%
96%
Blocks Ranking RGB
92%
85%
99%
97%
99%
98%
100%
100%
96%
90%
Blocks Ranking Size
49%
26%
75%
63%
94%
96%
94%
98%
62%
58%
Click Alarmclock
98%
89%
100%
100%
99%
100%
100%
100%
100%
100%
Appendix
Table 10 : Per-task success rates (%) on the 50-task RoboTwin 2.0 benchmark under clean and randomized (Rand.) settings. Baseline scores are taken from publicly reported results. Bold marks the highest score for each task and setting.
Simulation Task
Aloha
Franka
Piper
ARX-X5
UR5
Clean
Random
Clean
Random
Clean
Random
Clean
Random
Clean
Random
Beat Block Hammer
100%
100%
60%
64%
60%
76%
56%
68%
56%
48%
Blocks Ranking RGB
0%
0%
0%
0%
4%
0%
0%
0%
0%
0%
Blocks Ranking Size
64%
64%
68%
56%
80%
84%
52%
60%
12%
8%
Grab Roller
100%
96%
96%
96%
100%
100%
100%
100%
32%
48%
Handover Mic
80%
80%
40%
28%
96%
100%
100%
96%
48%
40%
Appendix
Table 11 : Per-task success rates (%) across five robot embodiments on 31 shared RoboTwin 2.0 tasks under clean and randomized settings. Each task is evaluated over 25 episodes per embodiment and setting.
Figure 6 : Execution progress of real-world manipulation tasks. Each row shows representative temporal snapshots for one task, illustrating how the robot complete place_block, wipe_plate, stack_blocks, and storage_item.