CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces
Authors: Sen Wang, Liu Liu, Xinjiang Wang, Zequn Chen, Haoyi Jiang, Taojun Ding, Tingyang Xiao, Zhizhong Su, +2 more
Organizations: Xi'an Jiaotong University · Horizon Robotics · Huazhong University of Science and Technology · University of Science and Technology of China
Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple action chunks. To bridge semantic context with physical prediction and control, CogWAM employs progress-conditioned WORLD and ACTION queries that selectively extract task-relevant information for future-world prediction and action generation. During training, the Semantic State provides shared task-progress context for both branches, while inference removes the future-prediction branch and directly generates actions from observations and the maintained state. We further introduce semantic training strategies to improve transition learning and closed-loop conditioning. Without additional robot-action pretraining, CogWAM achieves 15.56 / 11.70 % Score/SR on RoboDojo and state-of-the-art performance on BiCoord, while real-world experiments demonstrate closed-loop dual-arm manipulation with 16.4 fewer Semantic State regenerations than step-wise updating.
Figures & tables
Figure 1 : Overview of CogWAM. (a) Event-triggered updates maintain a persistent Semantic State while avoiding redundant updates and stale states. (b) The Semantic State couples semantic planning with world prediction and action generation through a closed-loop execution loop. (c) Benchmark results and real-world success–latency trade-off.
Figure 2 : Architecture of CogWAM. An event-triggered KEEP / UPDATE mechanism maintains a persistent Semantic State during closed-loop execution. WORLD and ACTION queries condition parallel branches for future multi-view DINO feature prediction and continuous action generation.
Method
Generalization
Precision
Long-Horizon
Memory
Open
Average
w/ Prior Embodied Robot-Data Pre-training
X-VLA Zheng et al. (2026)
10.47 / 6.78
18.32 / 12.00
16.53 / 9.75
4.76 / 3.56
0.55 / 0.50
10.13 / 6.52
InternVLA-A1.5 Ma et al. (2026)
10.35 / 6.83
15.23 / 10.17
23.80 / 13.75
4.93 / 3.56
1.43 / 1.42
11.15 / 7.14
π0.5 Black et al. (2025)
13.38 / 8.17
12.40 / 5.50
23.54 / 14.67
5.89 / 4.67
1.98 / 1.67
11.44 / 6.93
Spatial Forcing Li et al. (2026e)
14.12 / 9.34
17.32 / 10.58
23.26 / 14.58
5.43 / 4.11
1.78 / 1.58
12.38 / 8.04
Hy-Embodied-0.5-VLA Zhang et al. (2026c)
11.78 / 8.39
13.81 / 8.00
25.74 / 14.92
13.37 / 12.11
0.65 / 0.58
13.07 / 8.80
Table 1 : Multi-Task Evaluation on RoboDojo. Each cell reports Score / Success Rate (in %). Methods are grouped by whether prior embodied robot-data pre-training is used before RoboDojo-specific training. Bold denotes the best result within each group.
Method
Clean Table
Cook
Exchange Mics
Exchange Pots
Match Blocks With Signs
Put Objects Cabinet
⋯
Average (18 Tasks)
RDT Liu et al. (2025)
17.5 / 0.0
31.0 / 9.0
35.0 / 23.0
96.0 / 92.0
7.0 / 1.0
49.0 / 31.0
⋯
34.8 / 16.9
OpenVLA-OFT Kim et al. (2025)
25.0 / 2.0
25.0 / 10.0
67.5 / 66.0
56.0 / 53.0
8.7 / 5.0
49.5 / 26.0
⋯
36.5 / 23.1
π0 Black et al. (2024)
46.2 / 6.0
29.5 / 14.0
59.5 / 52.0
60.5 / 52.0
18.7 / 8.0
24.0 / 6.0
⋯
43.2 / 27.2
π0.5 Black et al. (2025)
61.5 / 24.0
48.0 / 31.0
62.0 / 55.0
68.0 / 61.0
15.0 / 6.0
58.5 / 42.0
⋯
50.8 / 35.6
CogWAM (Ours)
80.0 / 50.0
77.5 / 52.0
56.5 / 49.0
76.5 / 70.0
17.3 / 7.0
84.0 / 74.0
⋯
58.0 / 43.8
Table 2 : Multi-Task Evaluation on BiCoord. Each cell reports Stage-wise Success Rate / Success Rate (in %). We show the six tasks with the longest average expert trajectories, while Avg. is computed over all 18 tasks. Complete per-task results are provided in Appendix Tab. A4 .
Figure 3 : Qualitative Analysis of Semantic State Progression. Simulation rollouts illustrate how persistent Semantic States track task progress across action chunks, reducing subtask drift, repetition, and premature transitions. The real-world rollout visualizes the predicted KEEP / UPDATE decisions together with the corresponding Semantic States throughout execution.
Figure 4 : Effect of Semantic State and Update Strategies. Task performance on (a) RoboDojo (Score / SR) and (b) BiCoord (SR / SSR), shown with benchmark-specific vertical scales. w/o Semantic State conditions the policy only on the current observation and global task instruction, while the remaining variants differ in how the Semantic State is updated.
Basic Task
Success
Generalization
Success
Place Objects
20 / 20
Spatial Location
21 / 30
Organize Utensils
17 / 20
Object Appearance
20 / 30
Put in Drawer
18 / 20
Distractor
17 / 30
Fill Pen Holder
18 / 20
Novel Objects
14 / 30
Place in Bag
16 / 20
–
Block Sorting
5 / 20
–
Table 3: Real-world evaluation. Success counts across basic manipulation tasks and distribution shifts.
Figure 5 : Real-world setup and task examples. CogWAM is deployed on a dual-arm AgileX PIPER platform with two 6-DoF manipulators and parallel grippers, using synchronized RGB observations from one fixed head-mounted and two wrist-mounted Intel RealSense D435 cameras.
Semantic-model calls
Inference latency
Transition
Update strategy
Total
Regen. ↓
Mean (ms) ↓
Dec. tok. ↓
delay (ms) ↓
Synchronous
65.4
65.4
927.8
45
152
Asynchronous (1 Hz)
22.1
22.1
309.3
45
502
Event-triggered (ours)
65.4
0 4.0
146.7
0 0
152
Table 4: Deployment cost of update strategies. Total denotes semantic-model calls per rollout, while Regen. counts calls that regenerate the Semantic State; Dec. tok. denotes tokens decoded per replanning step. Transition delay is the lag between a subtask boundary and the first replanning step at which the strategy can observe it. Latency is measured on an NVIDIA RTX 5090.
Figure 6 : World and Action gradients inside the shared VLM, by depth. (a) Gradient alignment, with 95% bootstrap intervals over 120 batches. CogWAM is significantly positive at every depth and decays inward from the visual tower; the dual-DiT baseline with its gradient path restored is indistinguishable from zero throughout. (b) World-gradient magnitude relative to the action gradient. As trained, the dual-DiT baseline detaches the world branch, so no world gradient reaches the trunk and GA is undefined.
Components
RoboDojo Evaluation
Method
Pred.
VLM
State
Gen.
Prec.
Long-H.
Mem.
Open
Avg.
Action MoT
–
–
–
8.00 / 5.20
18.20 / 12.60
21.62 / 12.00
3.50 / 2.00
2.10 / 1.90
10.68 / 6.74
MoT
✓
–
–
11.73 / 8.83
21.24 / 14.00
21.88 / 13.64
5.72 / 4.67
1.05 / 1.00
12.32 / 8.43
CogWAM
✓
✓
✓
15.53 / 12.17
24.45 / 19.00
28.14 / 19.00
7.65 / 6.33
2.05 / 2.00
15.56 / 11.70
Table 5: Stepwise decomposition of CogWAM on RoboDojo. Each cell reports Score / Success Rate (in %). Action MoT → MoT adds DINO-space future-world prediction; MoT → CogWAM further adds the CogWAM Interface.
Method
Gen.
Prec.
Long-H.
Mem.
Open
Avg.
Fast-WAM Yuan et al. (2026)
2.33 / 1.11
1.96 / 0.00
9.14 / 5.17
3.55 / 3.44
0.42 / 0.42
3.48 / 2.03
Fast-WAM + CogWAM Interface
12.28 / 9.00
19.61 / 13.50
23.69 / 14.75
9.17 / 8.00
1.93 / 1.75
13.33 / 9.40
CogWAM
15.53 / 12.17
24.45 / 19.00
28.14 / 19.00
7.65 / 6.33
2.05 / 2.00
15.56 / 11.70
Table 6: Transferability of the CogWAM Interface on RoboDojo. Each cell reports Score / Success Rate (in %). Applying the same semantic interface to Fast-WAM substantially improves performance without changing its underlying predictive formulation.
Figure 7 : Language grounding under visual perturbations. CogWAM follows instructions specifying different target pen holders: white, grey, and brown. The grey-holder setting introduces spatial perturbations of the target objects, while the brown-holder setting introduces additional distractor objects. Despite these variations, CogWAM consistently grounds the instructed target and completes the corresponding manipulation behavior.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
VLM backbone
World–Action MoT
Visual teacher
RynnBrain1.1-2B
World stream
Action stream
DINOv3 ViT-B/16
Parameters
2.72 B
442 M
1.01 B
85.7 M
Trained
yes
yes
yes
frozen
Layers
24
30
12
Attention heads
0 8
24
12
Hidden dim
2048
0 512
1024
0 768
Appendix
Table A1: Model architecture. The backbone row reports its 24-layer language model; its vision tower is also 24 layers. The two MoT streams traverse the same 30 layers but keep separate widths, attention projections, and feed-forward blocks; attention head dimension is 128 for both. At inference, the world-prediction stream and future-target encoding are removed, while the frozen DINO encoder is retained to extract current visual features.
Figure A1 : Illustration of attention masks in CogWAM.
Configuration
Value
Optimizer
AdamW, β=(0.9,0.95) , ϵ=10−8
Weight decay
10−8
Learning rate (VLM backbone)
1×10−5
Learning rate (planner queries)
1×10−4
Learning rate (action model)
1×10−4
LR schedule
cosine, 2,000 warmup steps, floor 5×10−7
Appendix
Table A2: Training hyperparameters. The semantic stream is a second, independent dataloader: every step trains both a physical batch and a balanced semantic batch.
Task
Train ep.
Train frames
Val ep.
Val frames
Place Objects
1,0 66
0 83,314
0 1
0 1,244
Organize Utensils
1,0 87
112,181
0 5
0 6,725
Put in Drawer
1, 437
729,929
23
37,273
Fill Pen Holder
1,477
220,962
78
10,950
Place in Bag
1, 954
392,681
49
19,131
Block Sorting
1,210
410,721
64
21,353
Appendix
Table A3: Training data. Per-task counts after the 5% episode-level validation split.
Figure A2 : Action learning over training. Panel (a) shows per-step training loss (light) with an exponential moving average (dark) on a log scale. Panels (b) and (c) show held-out action MSE and reverse KL at every validation checkpoint.
Method / Task
Average
Balance Roller
Build Bridge
Build Tower With Blocks
Clean Table
Collect Pens
Cook
Divide Block Tower
Exchange Mics
RDT Liu et al. [2025]
34.8 / 16.9
43.0 / 15.0
50.2 / 47.0
1.0 / 0.0
17.5 / 0.0
68.8 / 15.0
31.0 / 9.0
12.8 / 2.0
35.0 / 23.0
OpenVLA-OFT Kim et al. [2025]
36.5 / 23.1
74.5 / 49.0
2.0 / 2.0
0.0 / 0.0
25.0 / 2.0
46.2 / 5.0
25.0 / 10.0
11.1 / 0.0
67.5 / 66.0
π0 Black et al. [2024]
43.2 / 27.2
81.5 / 68.0
47.0 / 44.0
5.0 / 2.0
46.2 / 6.0
81.8 / 47.0
29.5 / 14.0
10.8 / 0.0
59.5 / 52.0
π0.5 Black et al. [2025]
50.8 / 35.6
89.8 / 82.0
54.3 / 51.0
21.6 / 11.0
61.5 / 24.0
84.8 / 57.0
48.0 / 31.0
14.8 / 1.0
62.0 / 55.0
CogWAM (Ours)
58.0 / 43.8
96.5 / 95.0
60.2 / 58.0
35.0 / 19.0
80.0 / 50.0
87.2 / 66.0
77.5 / 52.0
18.0 / 1.0
56.5 / 49.0
Appendix
Table A4 : Multi-Task Evaluation on BiCoord. Each cell reports Stage-wise Success Rate (SSR) / Success Rate (SR) in %. Avg. is averaged over all 18 tasks. Bold denotes the best result in each column.
Without Robot Pre-training
With Robot Pre-training
Task
Fast-WAM
Fast-WAM+Interface
StarVLA- α
CogWAM (Ours)
π0.5
GalaxeaVLA (G0.5)
Generalization
stack_bowls (Std.)
5.87 / 2.67
72.20 / 68.00
15.27 / 10.67
72.20 / 68.00
76.00 / 72.00
64.07 / 58.67
stack_bowls (Rand.)
0.80 / 0.00
3.00 / 0.00
1.20 / 0.00
4.80 / 0.00
14.20 / 4.00
20.33 / 13.33
push_T (Std.)
0.00 / 0.00
0.00 / 0.00
0.00 / 0.00
0.00 / 0.00
0.00 / 0.00
0.00 / 0.00
push_T (Rand.)
0.00 / 0.00
0.00 / 0.00
0.00 / 0.00
0.00 / 0.00
0.00 / 0.00
0.00 / 0.00
Appendix
Table A5 : Complete task-wise results on RoboDojo. Each cell reports Score / Success Rate (in %). For Generalization tasks, Std. and Rand. denote the standard and randomized evaluation settings, respectively. Methods are grouped according to whether prior embodied robot-data pre-training is used before RoboDojo-specific training. Within the group without robot pre-training, bold and underlined numbers denote the best and second-best results for each metric, respectively. Methods with robot pre-training are reported for reference and are excluded from the ranking. Fast-WAM+Interface and CogWAM (Ours) are our own evaluation logs. Bold denotes the best result in each column.
Figure A3 : Closed-loop rollouts on BiCoord. Bimanual tasks whose stages differ in which arm leads. Subtask text carries the arm assignment, so a transition changes both what is being done and which arm does it.
Figure A4 : Closed-loop rollouts on RoboDojo. Three episodes of increasing length. The memory row grows one entry per UPDATE , so the subtask the policy is conditioned on is always paired with an explicit record of what preceded it.
Figure A5 : Closed-loop rollouts on the real robot. Block Sorting and Fill Pen Holder . The latter repeats one grasp–hand-over–insert cycle many times, so the memory is the only signal distinguishing the third repetition from the first. Zoom in for a better view.
World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions. Magic-W0 represents interaction as a Structured World Transition consisting of Current State, Transition, and Future State. Current State combines vision-language context with Current 3D Geometry; Transition is represented by 3D Motion capturing action-induced three-dimensional changes; and Future State is represented by Future Semantics describing task-relevant outcomes. To couple prediction and control, we propose a layer-aligned world-action interaction architecture in which evolving action hypotheses condition world-transition prediction, while predicted world representations continuously inform action generation. Magic-W0 is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision for geometry, 3D motion, and future semantics from pre-trained visual models. Inference-time interventions show that structured world representations respond systematically to changes in candidate actions and that action-related information propagates through shared 3D representations into future semantic predictions. On RoboDojo-Sim, Magic-W0 achieves an average Score of 27.10, the highest among the compared WAMs. Across multiple real-robot tasks, it also demonstrates strong downstream performance after fine-tuning with limited downstream data, supporting generalization and rapid adaptation.
World Action Models (WAMs) predict visual futures and robot actions, yet they remain susceptible to task-incomplete imagination, where plausible, action-consistent predictions omit the transition needed for task completion. In this paper, we show that this failure is not inherent to the world model backbone, but emerges when adapted for short-chunk control, which can repeatedly favor plausible local continuations over task-completing transitions. To address this, we introduce Completion Aware Guidance (CAG), a training-free sampling method that guides generation toward task completion. Across representative WAMs, CAG improves success from 64% to 70% on a RoboTwin 2.0 subset and from 69% to 75% in zero-shot simulation, while reducing task-incomplete imagination from 79% to 40%.
Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets, thereby constraining action learning to the inductive biases preserved by a particular representation. To address this limitation, we propose CF-WAM, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections. The action-relevant constraints exposed by these projections accumulate across training steps, forcing WAM to capture the underlying state-transition structure that supports multiple projections of the same action-conditioned future. This dynamic mechanism also provides a natural cross-embodiment dynamics reference frame for Human and Robot learning. By jointly learning across different next-state parameterizations, heterogeneous Human and Robot experience can bypass appearance differences and directly contribute to shared state-transition learning, improving cross-embodiment generalization. Experiments show that CF-WAM improves both training efficiency and final control performance, while translating Human experience effectively into policy gains. CF-WAM achieves state-of-the-art performance on RoboCasa-GR1 with an average success rate of 82.50%, while reaching 82.65% on LIBERO-Plus and up to 84.00% in real-world evaluations.