EVO-WAM: Evolving World Action Models through Video-Action Verification
Organizations: HITSZ · SLAI · THU · JD · HKUST · PKU · SJTU · HKUSTGZ · HKU · UBC · CUHK · CUHKSZ
Abstract
Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately and their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: https://evo-wam.github.io/.
Figures & tables
| Method | Object on scale | Stamp seal | Bread in basket | Place cup | Stack 3 blocks | A left of B | Cans in box | Average |
|---|---|---|---|---|---|---|---|---|
| Vision-Language-Action Models | ||||||||
| 17.0 | 14.0 | 14.0 | 44.0 | 0.0 | 29.0 | 0.0 | 16.9 | |
| LingBot-VLA | 13.0 | 0.5 | 29.0 | 17.0 | 0.0 | 21.5 | 2.0 | 11.9 |
| StarVLA-OFT | 0.0 | 0.0 | 8.5 | 21.5 | 0.0 | 0.5 | 0.0 | 4.4 |
| World Action Models | ||||||||
| Fast-WAM | 3.5 | 1.0 | 22.5 | 31.5 | 0.0 | 0.0 | 0.0 | 8.4 |
| Model | Stack Bowls | Place Ducks | Load the Air Fryer | Average |
|---|---|---|---|---|
| 10.0 | 10.0 | 0.0 | 6.7 | |
| DreamZero | 60.0 | 0.0 | 0.0 | 20.0 |
| Cosmos3 | 40.0 | 10.0 | 10.0 | 20.0 |
| EVO-WAM Cosmos3 (Ours) | 80.0 | 60.0 | 90.0 | 76.7 |
| Model | Round 0 | Round 1 | Round 2 | Round 3 | Round 4 |
|---|---|---|---|---|---|
| RoboTwin | |||||
| EVO-WAM DreamZero (Ours) | 28.5 | 36.5 (+8.0) | 42.3 (+13.8) | 45.1 (+16.6) | 46.4 (+17.9) |
| EVO-WAM Cosmos3 (Ours) | 26.9 | 58.3 (+31.4) | 66.6 (+39.8) | 63.6 (+36.7) | 68.0 (+41.1) |
| Real world | |||||
| EVO-WAM Cosmos3 (Ours) | 20.0 | 60.0 (+40.0) | 76.7 (+56.7) | 73.3 (+53.3) | 76.7 (+56.7) |
| Verification | VLM | R0 | R1 | R2 | R3 | R4 |
|---|---|---|---|---|---|---|
| VLM only | Qwen3.8-Flash-Next | 26.9 | 37.9 (+11.1) | 42.7 (+15.9) | 39.6 (+12.7) | 43.7 (+16.9) |
| VLM + IDM (Ours) | Qwen3.8-Flash-Next | 26.9 | 58.3 (+31.4) | 66.6 (+39.8) | 63.6 (+36.7) | 68.0 (+41.1) |
| VLM + Simulator | Qwen3.8-Flash-Next | 26.9 | 66.4 (+39.5) | 69.3 (+42.4) | 73.2 (+46.4) | 72.7 (+45.9) |
| VLM + IDM | Qwen3.5-27B | 26.9 | 52.7 (+25.9) | 57.8 (+30.9) | 62.7 (+35.9) | 65.7 (+38.9) |
| Evaluation | Cosmos3 | EVO-WAM Cosmos3 |
|---|---|---|
| New-scene generalization | 24.9 | 70.4 |
| Seen-task retention | 85.8 | 84.8 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Training data | Round 0 | Round 1 | Round 2 | Round 3 | Round 4 |
|---|---|---|---|---|---|
| Latest-round data only | 26.9 | 58.6 | 60.5 | 57.1 | 53.6 |
| Accumulated data | 26.9 | 58.3 | 66.6 | 63.6 | 68.0 |
| Clean | Randomized | Average | ||||
|---|---|---|---|---|---|---|
| Task | Adapt. | New | Adapt. | New | Adapt. | New |
| Place object on scale | 79.0 | 84.0 | 77.0 | 75.0 | 78.0 | 79.5 |
| Stamp seal | 48.0 | 49.0 | 61.0 | 59.0 | 54.5 | 54.0 |
| Place bread in basket | 74.0 | 79.0 | 71.0 | 85.0 | 72.5 | 82.0 |
| Place empty cup | 93.0 | 89.0 | 87.0 | 86.0 | 90.0 | 87.5 |
| Stack three blocks | 34.0 | 49.0 | 37.0 | 45.0 | 35.5 | 47.0 |
| VLA | WAM | EVO-WAM Cosmos3 (Ours) | |||||
| Layout | DreamZero | Cosmos3 | R1 | R2 | R3 | R4 | |
| Stack Bowls | |||||||
| Layout 1 | 0/3 | 3/3 | 1/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| Layout 2 | 0/3 | 3/3 | 2/3 | 3/3 | 3/3 | 2/3 | 2/3 |
| Layout 3 | 1/4 | 0/4 | 1/4 | 1/4 | 2/4 | 2/4 | 3/4 |
| Total | 1/10 | 6/10 | 4/10 | 7/10 | 8/10 | 7/10 | 8/10 |
| Setting | EVO-WAM Cosmos3 , RoboTwin | EVO-WAM DreamZero , RoboTwin | EVO-WAM Cosmos3 , real robot |
|---|---|---|---|
| Starting step | 30,000 | 30,000 | 30,000 |
| Global batch size | 256 | 256 | 256 |
| Updates per round | 1,000 | 1,000 | 500 |
| Rounds | 4 | 4 | 4 |
| Learning rate | |||
| Recorded/generated | 1:1 | 1:1 | 1:1 |
| Round | New prefixes | Cumulative pool | Updates |
| EVO-WAM Cosmos3 — RoboTwin | |||
| 1 | 531 | 531 | 1,000 |
| 2 | 1,400 | 1,931 | 1,000 |
| 3 | 1,482 | 3,413 | 1,000 |
| 4 | 1,592 | 5,005 | 1,000 |
| EVO-WAM DreamZero — RoboTwin | |||
| Verification | Precision | Recall | FPR |
|---|---|---|---|
| VLM | 68.0 | 64.4 | 12.8 |
| VLM + IDM | 86.0 | 44.2 | 3.0 |
| Setting | RoboTwin, Cosmos3 | RoboTwin, DreamZero | DROID, Cosmos3 |
|---|---|---|---|
| Action dimensions | 14 | 14 | 8 |
| Actions per window | |||
| Video frames | |||
| Video/action rate | 15/15 Hz | 5/15 Hz | 15/15 Hz |
| Normalized clipping | None | None | |
| Reconstruction steps | 4 | 4 | 4 |
| Setting | Cosmos3, RoboTwin | DreamZero, RoboTwin | Cosmos3, DROID |
|---|---|---|---|
| Video/action rate | 15/15 Hz | 5/15 Hz | 15/15 Hz |
| New frames/actions | 64/64 | 24/72 | 64/64 |
| Recent video context | 32 frames | 8 frames | 32 frames |
| Recent video latents | 8 | 2 | 8 |