Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately 2.5× and 1.6× their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: https://evo-wam.github.io/.
Figures & tables
Figure 1: EVO-WAM improves world action models through verified experience. Left: unseen tasks in simulation and long-horizon composite tasks in the real world. Middle: the generate-verify-improve cycle, where learning from verified rollouts improves task execution. Right: Cosmos3’s performance gains over successive rounds in simulation and the real world.
Figure 2: Overview of EVO-WAM. (1) Generate: WAMr−1 autoregressively generates video-action-state trajectories from an initial scene and task instruction. (2) Verify: a VLM identifies task-completing prefixes, and an IDM checks their video-action consistency. (3) Improve: verified prefixes accumulated across rounds are combined with the original training data to obtain WAMr , which generates candidates for the next round.
Figure 3: Video-action verification. (a) A VLM checks subgoals one at a time through description and judgment. (b) An IDM reconstructs actions from the generated videos and compares them with WAM-generated actions at the fixed visual endpoint. Prefixes passing this check undergo endpoint vote confirmation before entering training.
Figure 4: Unseen-task improvement in RoboTwin 2.0. (a) Example task scenes. (b) Cosmos3’s per-task success rates before self-training (Round 0) and after four rounds (Round 4).
Method
Object on scale
Stamp seal
Bread in basket
Place cup
Stack 3 blocks
A left of B
Cans in box
Average
Vision-Language-Action Models
π0.5
17.0
14.0
14.0
44.0
0.0
29.0
0.0
16.9
LingBot-VLA
13.0
0.5
29.0
17.0
0.0
21.5
2.0
11.9
StarVLA-OFT
0.0
0.0
8.5
21.5
0.0
0.5
0.0
4.4
World Action Models
Fast-WAM
3.5
1.0
22.5
31.5
0.0
0.0
0.0
8.4
Table 1: RoboTwin 2.0 results on unseen tasks. Success rates (%) on seven tasks unseen during base-model training. EVO-WAM Cosmos3 achieves 68.0% average success, compared with 31.6% for the strongest baseline Cosmos3. Baselines are trained for 34K steps; EVO-WAM models start from 30K checkpoints and are evaluated at 34K steps. Bold and underlined values indicate the best and second-best results.
Figure 5: Real-robot tasks. Initial and goal scenes for three unseen composite tasks.
Model
Stack Bowls
Place Ducks
Load the Air Fryer
Average
π0.5
10.0
10.0
0.0
6.7
DreamZero
60.0
0.0
0.0
20.0
Cosmos3
40.0
10.0
10.0
20.0
EVO-WAM Cosmos3 (Ours)
80.0
60.0
90.0
76.7
Table 2: Real-robot results. Success rates (%) on three unseen long-horizon composite tasks. EVO-WAM Cosmos3 achieves 76.7% average success, compared with 20.0% for Cosmos3. The Cosmos3 baseline and our Round 2 model are both evaluated at 31K steps. Bold and underlined values indicate the best and second-best results across all methods.
Figure 6: Placing two ducks. Top and bottom show real executions before and after EVO-WAM. The middle shows generated candidates and their visual-goal and video-action consistency checks; prefixes passing both checks are used to update the policy.
Model
Round 0
Round 1
Round 2
Round 3
Round 4
RoboTwin
EVO-WAM DreamZero (Ours)
28.5
36.5 (+8.0)
42.3 (+13.8)
45.1 (+16.6)
46.4 (+17.9)
EVO-WAM Cosmos3 (Ours)
26.9
58.3 (+31.4)
66.6 (+39.8)
63.6 (+36.7)
68.0 (+41.1)
Real world
EVO-WAM Cosmos3 (Ours)
20.0
60.0 (+40.0)
76.7 (+56.7)
73.3 (+53.3)
76.7 (+56.7)
Table 3: Improvement over successive self-training rounds. Average success rates (%) on seven unseen RoboTwin 2.0 tasks and three unseen real-world tasks. Parentheses show gains over Round 0 in percentage points (pp), computed before rounding. Round 0 denotes the 30K self-training initialization. Each simulation round adds 1K updates, and each real-world round adds 500 updates.
Verification
VLM
R0
R1
R2
R3
R4
VLM only
Qwen3.8-Flash-Next
26.9
37.9 (+11.1)
42.7 (+15.9)
39.6 (+12.7)
43.7 (+16.9)
VLM + IDM (Ours)
Qwen3.8-Flash-Next
26.9
58.3 (+31.4)
66.6 (+39.8)
63.6 (+36.7)
68.0 (+41.1)
VLM + Simulator
Qwen3.8-Flash-Next
26.9
66.4 (+39.5)
69.3 (+42.4)
73.2 (+46.4)
72.7 (+45.9)
VLM + IDM
Qwen3.5-27B
26.9
52.7 (+25.9)
57.8 (+30.9)
62.7 (+35.9)
65.7 (+38.9)
Table 4: Effect of verification on self-training. Success rates (%) on unseen RoboTwin 2.0 tasks across rounds, comparing action verification criteria and task-completion VLMs.
Evaluation
Cosmos3
EVO-WAM Cosmos3
New-scene generalization
24.9
70.4
Seen-task retention
85.8
84.8
Table 5: Generalization and retention (%) of EVO-WAM.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Effect of data accumulation. Success in adaptation scenes across self-improvement rounds under the two configurations described above.
Training data
Round 0
Round 1
Round 2
Round 3
Round 4
Latest-round data only
26.9
58.6
60.5
57.1
53.6
Accumulated data
26.9
58.3
66.6
63.6
68.0
Appendix
Table 6: Effect of data accumulation. Success rates (%) corresponding to Figure 7 , with 1,400 trials per round. Accumulated data reproduces Table 3 . Both series use IDMs trained on the 43 seen RoboTwin tasks.
Clean
Randomized
Average
Task
Adapt.
New
Adapt.
New
Adapt.
New
Place object on scale
79.0
84.0
77.0
75.0
78.0
79.5
Stamp seal
48.0
49.0
61.0
59.0
54.5
54.0
Place bread in basket
74.0
79.0
71.0
85.0
72.5
82.0
Place empty cup
93.0
89.0
87.0
86.0
90.0
87.5
Stack three blocks
34.0
49.0
37.0
45.0
35.5
47.0
Appendix
Table 7: Per-task evaluation across scene sets. Success rates (%) of the final EVO-WAM Cosmos3 model at 34K steps. Adapt. and New denote adaptation and new scenes; each set has 100 Clean and 100 Randomized trials per task (1,400 total). Average pools both conditions.
Figure 8: Initial configurations for physical evaluation. Three layouts for Stack Bowls, three for Place Ducks, and two for Load the Air Fryer. Each task has ten trials per evaluated policy: 3/3/4 across the three-layout tasks and 5/5 across the two air-fryer layouts.
VLA
WAM
EVO-WAM Cosmos3 (Ours)
Layout
π0.5
DreamZero
Cosmos3
R1
R2
R3
R4
Stack Bowls
Layout 1
0/3
3/3
1/3
3/3
3/3
3/3
3/3
Layout 2
0/3
3/3
2/3
3/3
3/3
2/3
2/3
Layout 3
1/4
0/4
1/4
1/4
2/4
2/4
3/4
Total
1/10
6/10
4/10
7/10
8/10
7/10
8/10
Appendix
Table 8: Real-robot results by layout. Successful trials over attempts for the models in Table 2 . Layout numbers match Figure 8 . The Cosmos3 baseline is evaluated at 31K steps. R1–R4 denote the self-training rounds; Table 2 uses R2.
Figure 9: Stacking three bowls. The instruction is: “Stack the pink bowl on the blue bowl, then lift both together onto the white bowl.” Before shows the bowls still separated. The imagined trajectories illustrate different stacking orders and verification outcomes; After shows the two successive stacking operations.
Figure 10: Loading the air fryer. The robot must pull the red handle to open the drawer and transfer the bread from the plate into it. Before shows the bread still held by the gripper. Rollout 1 distorts the drawer and bread during transfer, while Rollout 3 passes verification. After shows drawer opening, bread transfer, and release.
Setting
EVO-WAM Cosmos3 , RoboTwin
EVO-WAM DreamZero , RoboTwin
EVO-WAM Cosmos3 , real robot
Starting step
30,000
30,000
30,000
Global batch size
256
256
256
Updates per round
1,000
1,000
500
Rounds
4
4
4
Learning rate
10−5
2×10−5
10−5
Recorded/generated
1:1
1:1
1:1
Appendix
Table 9: Self-training configuration. Learning rates are constant within each round. The recorded/generated ratio is the training sampling ratio.
Round
New prefixes
Cumulative pool
Updates
EVO-WAM Cosmos3 — RoboTwin
1
531
531
1,000
2
1,400
1,931
1,000
3
1,482
3,413
1,000
4
1,592
5,005
1,000
EVO-WAM DreamZero — RoboTwin
Appendix
Table 10: Retained training-pool sizes. New prefixes are trajectories retained for training after filtering and exclusions. Cumulative pool size sums new prefixes through the current round. Each retained trajectory contributes one prefix. Updates are optimizer steps per round.
Figure 11: Generated completion and failed action execution. Top: the video generated by the 30,000-step Cosmos3 policy. Bottom: simulator execution of the same generated action sequence. Columns use the same action steps and show the main view above the two wrist views. The generated stapler reaches the scale; the replayed stapler remains on the table. The last column is the end of the 194-action replay. The visually accepted 120-action prefix has an IDM consistency error of 0.00632283, exceeding the fixed threshold of 0.00418487, and is therefore rejected by the IDM.
Verification
Precision ↑
Recall ↑
FPR ↓
VLM
68.0
64.4
12.8
VLM + IDM
86.0
44.2
3.0
Appendix
Table 11: Verifier reliability against simulator replay (%). Both methods use the same 700 candidates. Precision and recall treat replay success as the positive label; FPR is the fraction of replay failures accepted.
Figure 12: A recorded verification trace for stacking bowls. The initial VLM scan proposes endpoints at step 165 (11 s) for the first subgoal and step 320 (21.33 s) for the second. The 320-action prefix passes the IDM threshold, after which each endpoint receives two further assessments. Both subgoals receive Accept/Accept/Reject, satisfying the two-of-three rule. Five complete 64-action windows are scored, and all 320 actions are retained for training. The images show wrist and right external views; assessment also uses the left view.
Setting
RoboTwin, Cosmos3
RoboTwin, DreamZero
DROID, Cosmos3
Action dimensions
14
14
8
Actions per window
1≤L≤64
L∈{3,6,…,72}
1≤L≤64
Video frames
L+1
L/3+1
L+1
Video/action rate
15/15 Hz
5/15 Hz
15/15 Hz
Normalized clipping
None
None
[−5,5]
Reconstruction steps
4
4
4
Appendix
Table 12: IDM settings. L is the number of actions in a verification window; the video includes its starting observation. Standardization uses recorded-data means and standard deviations.
Setting
Cosmos3, RoboTwin
DreamZero, RoboTwin
Cosmos3, DROID
Video/action rate
15/15 Hz
5/15 Hz
15/15 Hz
New frames/actions
64/64
24/72
64/64
Recent video context
32 frames
8 frames
32 frames
Recent video latents
8
2
8
Appendix
Table 13: Autoregressive generation settings. Output counts exclude the starting observation. Recent visual context is in addition to the initial anchor.
Figure 13: Object consistency during autoregressive rollout. Top: without global sink and context. Bottom: with both components. Columns show matched timestamps relative to each clip’s start. The yellow block is initially visible and is occluded by the arm at 10 s. At 12 s and 22 s, it is missing in the top row and preserved in the bottom row. Each frame shows the main camera view cropped from the source video; boxes and enlarged insets highlight the same fixed image region in both rows.
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China · Singapore Management University, Singapore · Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China