World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points and learns from successful and failed futures to support classifier-free guidance. An additional value head estimates task progress from video representations and activates guidance when progress stalls during inference. Across five DexJoCo tasks, DEWO improves average success across all three WAM formulations. Ablations show that visual supervision from successful and failed continuations improves both prediction and control beyond action supervision alone. On four real-world tasks across Wuji and Sharpa, 3 x 3 grid evaluations show that two rounds of deployment learning increase success from 51.0% to 71.7% in cells with at least one initial success, a gain of 20.7 percentage points. These findings support continued predictive learning for improving control through deployment experience, making world modeling an active part of WAM adaptation.
Figures & tables
Figure 1: Learning signals for post-deployment policy improvement. Direct action supervision ( top ) learns from action targets. Value-based reinforcement learning ( middle ) uses value estimates to improve the policy. DEWO ( bottom ) learns from observed visual futures and uses the learned success condition to guide actions.
Figure 2: Overview of Direct Experience World-Model Optimization (DEWO). Replay exploration locates interaction turning points and collects successful and failed continuations from matched contexts. Their visual futures train outcome-conditioned world modeling, while successful behavior supplies action targets. At deployment, a value head monitors task progress and selectively activates classifier-free guidance using the base and success-conditioned action predictions.
Model
Method
Water
Fold
Hammer
Pick
Pinch
Avg.
Plant
Glasses
Nail
Bucket
Tongs
A. Matched post-deployment comparisons
π0.5
Initial
73.3
58.0
74.7
78.7
62.7
69.5
+ SFT
75.3
59.3
76.7
85.3
22.7
63.9
+ RECAP
75.3
51.3
80.0
84.0
57.3
69.6
+ DSRL
76.7
54.7
81.3
86.0
22.7
64.3
Table 1: Simulation success rates (%) over three 50-trial sets per task; Avg. weights tasks equally. A: matched adaptation comparisons. B: prediction–action formulation ablation; FastWAM repeats A. C: DEWO-S construction from pretrained components (—: no subsequent adaptation). Per-set results are in Appendix F .
Figure 3: Spatial and iterative real-world evaluation. (a) Two hands, four objects, three rounds; image labels Round #1/#2/#3 denote R0/R1/R2. (b) Pooled success for eight centers, 29 R0-nonzero cells (including centers), 43 R0-zero cells, and all 72 cells. Groups stay fixed across rounds, with ten trials per cell (720 per round). (c) Full-grid rates by object and hand; π0.5 is a separate reference.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Pool or source
DEWO
DEWO-S
Expert episodes included in D0
0
500
Collected successful episodes included in D0
106
106
D0 subtotal
106
606
Dscan
1,396
1,396
D+
125
125
Dfail
134
134
Appendix
Table 2: Five-task training pools before window expansion. Source-entry counts differ from frame and interaction counts. The expert pool contains 100 successful episodes per task.
Setting
DEWO
DEWO-S
Initialization
Five-task FastWAM checkpoint at 55,000 updates
Pretrained VideoDiT and ActionDiT components
Trainable parameters
Text cross-attention K/V residual adapters and value head
Full video/action DiTs and MoT, proprioception/outcome encoders, and value head
Frozen components
MoT, VideoDiT, ActionDiT backbone weights
No additional freezing
Adapter
Rank 16, scale parameter 16; video and action experts
None
Maximum updates
10,000
100,000
Checkpoint save interval
2,500 updates
5,000 updates
Appendix
Table 3: Optimization settings for five-task FastWAM adaptation with DEWO and Joint WAM construction with DEWO-S.
Setting
Value
Simulation evaluation
3 sets × 50 scenes per task
Initial simulation collection
50 seeds × 4 attempts per task
Simulation / real continuations per anchor
K=10 / K=4
Simulation query start / interval
Step 96 / 24 steps
Maximum queries per failed trajectory
20 anchors
Intermediate recoverability-drop threshold
At least 4/10
Appendix
Table 4: Collection, temporal, and deployment settings. The guidance interval includes replans 10–24 for all five FastWAM DEWO simulation tasks.
Role
Video (DEWO)
Video (DEWO-S)
Action
Value
D0
1
1
1
1
D+
1
1
1
1
Dscan
0
0
0
1
Dfail
1
1
0
1
Appendix
Table 5: Per-role supervision weights. Action and value columns apply to both methods. Value is supervised separately; failed action targets are masked.
Spatial group
Trials/round
R0
R1
R2
Center
80
64
71
71
Initially nonzero off-center
210
84
138
137
Initially zero off-center
430
0
0
0
All initially nonzero (including center)
290
148
209
208
All cells
720
148
209
208
Appendix
Table 6: Pooled spatial success counts and fixed denominators. The initially nonzero subtotal combines the center and nonzero off-center rows.
Table 11: Per-evaluation-set DEWO-S success rates (%). Parentheses give successful trials / total trials. The final column reports the mean and population standard deviation over the three 50-scene evaluation sets, with pooled trial counts.
Method
Task
Eval. 0
Eval. 1
Eval. 2
Mean
Initial
Water Plant
68 (34/50)
66 (33/50)
64 (32/50)
66.0 (99/150)
Fold Glasses
76 (38/50)
64 (32/50)
64 (32/50)
68.0 (102/150)
Hammer Nail
26 (13/50)
24 (12/50)
26 (13/50)
25.3 (38/150)
Pick Bucket
88 (44/50)
84 (42/50)
84 (42/50)
85.3 (128/150)
Pinch Tongs
80 (40/50)
76 (38/50)
78 (39/50)
78.0 (117/150)
All five tasks
67.6 (169/250)
62.8 (157/250)
63.2 (158/250)
64.5 (484/750)
Appendix
Table 12: Per-evaluation-set FACT success rates (%). Each method uses one fixed checkpoint across three 50-scene evaluation sets per task. Parentheses give successful trials / total trials. The final column reports the mean over the three sets with pooled trial counts; the all-task rows aggregate 250 trials per set and 750 trials overall.
Emb.
Object
TL
TC
TR
ML
C
MR
BL
BC
BR
Overall
A. π0.5 reference
Wuji
Water Bottle
0
4
1
0
10
6
0
0
0
21/90 (23.3%)
Tape
0
0
0
4
10
0
3
9
0
26/90 (28.9%)
Eraser
0
0
0
2
5
0
1
3
0
11/90 (12.2%)
Tennis Ball
0
0
0
0
4
0
0
2
0
6/90 (6.7%)
Sharpa
Water Bottle
0
5
1
0
10
8
0
0
0
24/90 (26.7%)
Appendix
Table 13: Real-world spatial evaluation. Each spatial entry reports complete grasp-and-place successes out of 10 trials; Overall reports the total out of 90 trials and the corresponding success rate. Panels A–D show π0.5 and DEWO-S Rounds 0–2, respectively.
World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). In VDM-based WAMs, observations are first encoded by a VAE, and the resulting compressed latents are then processed by large video diffusion backbones to extract effective features for action generation. However, this paradigm ties WAM performance and training cost to large-scale video generation pretraining, limiting WAM efficiency and scalability. In this paper, we theoretically and empirically investigate how visual representations affect action generation in WAMs. Our results show that predictive embeddings from Joint-Embedding Predictive Architecture (JEPA) encoders better support action generation than compressed VAE latents, with I-JEPA performing best in our encoder comparison. Based on these findings, we propose LeWAM, which conditions action generation on JEPA embeddings and models environment evolution by predicting future embeddings in the same space, without relying on a video diffusion backbone. We further find that imitation learning matches demonstrated actions but does not distinguish better actions from worse ones, even though small action deviations can greatly affect task success. To address this limitation without additional environment interaction or the human oversight required for resets and safety, we introduce Demonstration-Guided DPO (DemoDPO), an offline preference refinement stage that derives preference supervision directly from demonstrations. With only 0.4B trainable parameters, LeWAM achieves an average success rate of 92.28% on RoboTwin 2.0, comparable to that of state-of-the-art VLAs and WAMs, and maintains practical effectiveness on real-world manipulation tasks.
Xueji Fang, Boqiang Duan, Hua Wu +2
Zhejiang University · Westlake University · Baidu Inc.
Post-training is essential for turning pretrained generalist robot policies into reliable task-specific controllers, but existing human-in-the-loop pipelines remain tied to physical execution: each correction requires robot time, scene setup, resets, and operator supervision in the real world. Meanwhile, action-conditioned world models have been studied mainly for imagination, synthetic data generation, and policy evaluation. We propose \textbf{Human-in-the-World-Model (Hi-WM)}, a post-training framework that uses a learned world model as a reusable corrective substrate for failure-targeted policy improvement. A policy is first rolled out in closed loop inside the world model; when the rollout becomes incorrect or failure-prone, a human intervenes directly in the model to provide short corrective actions. Hi-WM caches intermediate states and supports rollback and branching, allowing a single failure state to be reused for multiple corrective continuations and yielding dense supervision around behaviors that the base policy handles poorly. The resulting corrective trajectories are then added back to the training set for post-training. We evaluate Hi-WM on three real-world manipulation tasks spanning both rigid and deformable object interaction, and on two policy backbones. Hi-WM improves real-world success by 37.9 points on average over the base policy and by 19.0 points over a world-model closed-loop baseline, while world-model evaluation correlates strongly with real-world performance (r = 0.953). These results suggest that world models can serve not only as generators or evaluators, but also as effective corrective substrates for scalable robot post-training.
Yaxuan Li, Zhongyi Zhou, Yefei Chen +5
Current Robotics · Tsinghua University · Peking University +1
Recent World-Action (WA) models demonstrate strong generalization ability and data efficiency, but they typically rely on expert trajectories for training. This reliance limits their ability to acquire fine-grained manipulation skills beyond the demonstration distribution and prevents them from continuously improving through real-world interaction. To address these limitations, we propose WAM-RL, a reinforcement learning framework that enables joint optimization of the world model and the action model through online interaction with the environment. By allowing the two components to co-evolve, our approach enhances fine-grained control and adaptability. Specifically, a WA model consists of a world model and an actor. We design a tailored reinforcement learning method with hierarchical optimization to coordinate their improvement. On the methodological side, we systematically investigate the effects of applying reinforcement learning to the action model, as well as online training of the world model within an RL setting. Our experiments reveal a key insight: optimizing only the actor yields improvements on short-horizon tasks, but fails to provide significant gains on long-horizon tasks. In contrast, jointly optimizing both the world model and the actor is critical for achieving strong performance in long-horizon settings. Our work is the first to introduce reinforcement learning into the World-Action paradigm, and provides insights into how online optimization of both the action head and the world model impacts overall performance.
Zezhong Qian, Xiaowei Chi, Yu Qi +3
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Northeastern University · Tsinghua University