CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces
Organizations: Xi'an Jiaotong University · Horizon Robotics · Huazhong University of Science and Technology · University of Science and Technology of China
Abstract
Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple action chunks. To bridge semantic context with physical prediction and control, CogWAM employs progress-conditioned WORLD and ACTION queries that selectively extract task-relevant information for future-world prediction and action generation. During training, the Semantic State provides shared task-progress context for both branches, while inference removes the future-prediction branch and directly generates actions from observations and the maintained state. We further introduce semantic training strategies to improve transition learning and closed-loop conditioning. Without additional robot-action pretraining, CogWAM achieves 15.56 / 11.70 % Score/SR on RoboDojo and state-of-the-art performance on BiCoord, while real-world experiments demonstrate closed-loop dual-arm manipulation with 16.4 fewer Semantic State regenerations than step-wise updating.
Figures & tables
| Method | Generalization | Precision | Long-Horizon | Memory | Open | Average |
| w/ Prior Embodied Robot-Data Pre-training | ||||||
| X-VLA Zheng et al. (2026) | 10.47 / 6.78 | 18.32 / 12.00 | 16.53 / 9.75 | 4.76 / 3.56 | 0.55 / 0.50 | 10.13 / 6.52 |
| InternVLA-A1.5 Ma et al. (2026) | 10.35 / 6.83 | 15.23 / 10.17 | 23.80 / 13.75 | 4.93 / 3.56 | 1.43 / 1.42 | 11.15 / 7.14 |
| Black et al. (2025) | 13.38 / 8.17 | 12.40 / 5.50 | 23.54 / 14.67 | 5.89 / 4.67 | 1.98 / 1.67 | 11.44 / 6.93 |
| Spatial Forcing Li et al. (2026e) | 14.12 / 9.34 | 17.32 / 10.58 | 23.26 / 14.58 | 5.43 / 4.11 | 1.78 / 1.58 | 12.38 / 8.04 |
| Hy-Embodied-0.5-VLA Zhang et al. (2026c) | 11.78 / 8.39 | 13.81 / 8.00 | 25.74 / 14.92 | 13.37 / 12.11 | 0.65 / 0.58 | 13.07 / 8.80 |
| Method | Clean Table | Cook | Exchange Mics | Exchange Pots | Match Blocks With Signs | Put Objects Cabinet | Average (18 Tasks) | |
| RDT Liu et al. (2025) | 17.5 / 0.0 | 31.0 / 9.0 | 35.0 / 23.0 | 96.0 / 92.0 | 7.0 / 1.0 | 49.0 / 31.0 | 34.8 / 16.9 | |
| OpenVLA-OFT Kim et al. (2025) | 25.0 / 2.0 | 25.0 / 10.0 | 67.5 / 66.0 | 56.0 / 53.0 | 8.7 / 5.0 | 49.5 / 26.0 | 36.5 / 23.1 | |
| Black et al. (2024) | 46.2 / 6.0 | 29.5 / 14.0 | 59.5 / 52.0 | 60.5 / 52.0 | 18.7 / 8.0 | 24.0 / 6.0 | 43.2 / 27.2 | |
| Black et al. (2025) | 61.5 / 24.0 | 48.0 / 31.0 | 62.0 / 55.0 | 68.0 / 61.0 | 15.0 / 6.0 | 58.5 / 42.0 | 50.8 / 35.6 | |
| CogWAM (Ours) | 80.0 / 50.0 | 77.5 / 52.0 | 56.5 / 49.0 | 76.5 / 70.0 | 17.3 / 7.0 | 84.0 / 74.0 | 58.0 / 43.8 |
| Basic Task | Success | Generalization | Success |
| Place Objects | 20 / 20 | Spatial Location | 21 / 30 |
| Organize Utensils | 17 / 20 | Object Appearance | 20 / 30 |
| Put in Drawer | 18 / 20 | Distractor | 17 / 30 |
| Fill Pen Holder | 18 / 20 | Novel Objects | 14 / 30 |
| Place in Bag | 16 / 20 | – | |
| Block Sorting | 5 / 20 | – | |
| Semantic-model calls | Inference latency | Transition | |||
| Update strategy | Total | Regen. | Mean (ms) | Dec. tok. | delay (ms) |
| Synchronous | 65.4 | 65.4 | 927.8 | 45 | 152 |
| Asynchronous (1 Hz) | 22.1 | 22.1 | 309.3 | 45 | 502 |
| Event-triggered (ours) | 65.4 | 0 4.0 | 146.7 | 0 0 | 152 |
| Components | RoboDojo Evaluation | ||||||||
| Method | Pred. | VLM | State | Gen. | Prec. | Long-H. | Mem. | Open | Avg. |
| Action MoT | – | – | – | 8.00 / 5.20 | 18.20 / 12.60 | 21.62 / 12.00 | 3.50 / 2.00 | 2.10 / 1.90 | 10.68 / 6.74 |
| MoT | – | – | 11.73 / 8.83 | 21.24 / 14.00 | 21.88 / 13.64 | 5.72 / 4.67 | 1.05 / 1.00 | 12.32 / 8.43 | |
| CogWAM | 15.53 / 12.17 | 24.45 / 19.00 | 28.14 / 19.00 | 7.65 / 6.33 | 2.05 / 2.00 | 15.56 / 11.70 | |||
| Method | Gen. | Prec. | Long-H. | Mem. | Open | Avg. |
| Fast-WAM Yuan et al. (2026) | 2.33 / 1.11 | 1.96 / 0.00 | 9.14 / 5.17 | 3.55 / 3.44 | 0.42 / 0.42 | 3.48 / 2.03 |
| Fast-WAM + CogWAM Interface | 12.28 / 9.00 | 19.61 / 13.50 | 23.69 / 14.75 | 9.17 / 8.00 | 1.93 / 1.75 | 13.33 / 9.40 |
| CogWAM | 15.53 / 12.17 | 24.45 / 19.00 | 28.14 / 19.00 | 7.65 / 6.33 | 2.05 / 2.00 | 15.56 / 11.70 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| VLM backbone | World–Action MoT | Visual teacher | ||
| RynnBrain1.1-2B | World stream | Action stream | DINOv3 ViT-B/16 | |
| Parameters | 2.72 B | 442 M | 1.01 B | 85.7 M |
| Trained | yes | yes | yes | frozen |
| Layers | 24 | 30 | 12 | |
| Attention heads | 0 8 | 24 | 12 | |
| Hidden dim | 2048 | 0 512 | 1024 | 0 768 |
| Configuration | Value |
| Optimizer | AdamW, , |
| Weight decay | |
| Learning rate (VLM backbone) | |
| Learning rate (planner queries) | |
| Learning rate (action model) | |
| LR schedule | cosine, warmup steps, floor |
| Task | Train ep. | Train frames | Val ep. | Val frames |
| Place Objects | 1,0 66 | 0 83,314 | 0 1 | 0 1,244 |
| Organize Utensils | 1,0 87 | 112,181 | 0 5 | 0 6,725 |
| Put in Drawer | 1, 437 | 729,929 | 23 | 37,273 |
| Fill Pen Holder | 1,477 | 220,962 | 78 | 10,950 |
| Place in Bag | 1, 954 | 392,681 | 49 | 19,131 |
| Block Sorting | 1,210 | 410,721 | 64 | 21,353 |
| Method / Task | Average | Balance Roller | Build Bridge | Build Tower With Blocks | Clean Table | Collect Pens | Cook | Divide Block Tower | Exchange Mics |
| RDT Liu et al. [2025] | 34.8 / 16.9 | 43.0 / 15.0 | 50.2 / 47.0 | 1.0 / 0.0 | 17.5 / 0.0 | 68.8 / 15.0 | 31.0 / 9.0 | 12.8 / 2.0 | 35.0 / 23.0 |
| OpenVLA-OFT Kim et al. [2025] | 36.5 / 23.1 | 74.5 / 49.0 | 2.0 / 2.0 | 0.0 / 0.0 | 25.0 / 2.0 | 46.2 / 5.0 | 25.0 / 10.0 | 11.1 / 0.0 | 67.5 / 66.0 |
| Black et al. [2024] | 43.2 / 27.2 | 81.5 / 68.0 | 47.0 / 44.0 | 5.0 / 2.0 | 46.2 / 6.0 | 81.8 / 47.0 | 29.5 / 14.0 | 10.8 / 0.0 | 59.5 / 52.0 |
| Black et al. [2025] | 50.8 / 35.6 | 89.8 / 82.0 | 54.3 / 51.0 | 21.6 / 11.0 | 61.5 / 24.0 | 84.8 / 57.0 | 48.0 / 31.0 | 14.8 / 1.0 | 62.0 / 55.0 |
| CogWAM (Ours) | 58.0 / 43.8 | 96.5 / 95.0 | 60.2 / 58.0 | 35.0 / 19.0 | 80.0 / 50.0 | 87.2 / 66.0 | 77.5 / 52.0 | 18.0 / 1.0 | 56.5 / 49.0 |
| Without Robot Pre-training | With Robot Pre-training | |||||
| Task | Fast-WAM | Fast-WAM+Interface | StarVLA- | CogWAM (Ours) | GalaxeaVLA (G0.5) | |
| Generalization | ||||||
| stack_bowls (Std.) | 5.87 / 2.67 | 72.20 / 68.00 | 15.27 / 10.67 | 72.20 / 68.00 | 76.00 / 72.00 | 64.07 / 58.67 |
| stack_bowls (Rand.) | 0.80 / 0.00 | 3.00 / 0.00 | 1.20 / 0.00 | 4.80 / 0.00 | 14.20 / 4.00 | 20.33 / 13.33 |
| push_T (Std.) | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| push_T (Rand.) | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |