EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action
Organizations: Yinwang Intelligent Technology Co. Ltd.
Abstract
Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embodied model whose asymmetric joint attention lets action tokens read semantic, current-visual, predicted-future, and action information at every layer while the perceptual experts retain their distinct roles. Without layer-wise supervision, EWAM develops an emergent depth-wise specialization: action queries attend mainly to vision-language features in shallow layers, to predicted future frames in intermediate layers, and to action tokens themselves in deep layers. This handoff replicates across tasks and is stable across denoising steps. Checkpoint tracking and causal interventions show that it is learned and that action generation depends on it. EWAM is pretrained in two separate regimes, one on cross-embodiment robot trajectories and one on human egocentric video. In simulation and real-robot experiments, it surpasses existing VLA, WAM, and hybrid baselines. Human egocentric data improve both cross-embodiment transfer and real-robot robustness, and subtask-phase supervision improves long-horizon completion. Together, these results suggest that unified embodied learning can induce an ordered internal progression from semantic understanding, through visual foresight, to action formation.
Figures & tables
| Training results (without pretraining) | Inference results (robot-pretrained model) | ||||
|---|---|---|---|---|---|
| Configuration | Score | Configuration | Score | ||
| Full unified model | 83 | – | Full unified model | 92.80 | – |
| Symmetric full attention | 81 | Mask all video, layers 0–9 | 91.98 | ||
| No VL access, layers 0–9 | 78 | Mask VL, layers 0–9 | 76.53 | ||
| No video access, layers 10–19 | 78 | Mask condition, layers 10–19 | 81.40 | ||
| Action only, layers 24–29 | 83 | Mask future, layers 10–19 | 3.55 | ||
| Task | No injection | Injection setting | Observed behavior or outcome | Success |
| Adjust Bottle | 5/5 | Left-facing (43, 46) | Donor-directed first grasp misses; both recover after cutoff | 2/2 |
| Right-facing (44, 45, 47) | Donor-consistent right arm instead of required left arm; no recovery | 0/3 | ||
| Place Empty Cup | 5/5 | All targets (43–47) | All injected rollouts fail | 0/5 |
| Stack Blocks Two | 5/5 | All targets (43–47) | All injected rollouts fail | 0/5 |
| Click Bell | 5/5 | All targets (43–47) | Only seed 46 remains successful | 1/5 |
| Grab Roller | 5/5 | All targets (43–47) | Only seed 45 remains successful | 1/5 |
| Session | TCP P95 (mm) | Checks | |
|---|---|---|---|
| Left | Right | ||
| 013312 | 56.79 | 2.92 | Pass |
| 013326 | 77.01 | 14.65 | Pass |
| 013335 | 42.53 | 1.66 | Pass |
| 013352 | 18.63 | 1.30 | Pass |
| 222650 | 4.07 | 19.69 | Pass |
| Paradigm | Model | C2C (%) | C2R (%) | Average (%) |
|---|---|---|---|---|
| VLA | starVLA [ 46 ] | 46.52 | 3.16 | 24.84 |
| GalaxeaVLA [ 25 ] | 62.70 | 12.72 | 37.71 | |
| Xiaomi Robotics-0 [ 9 ] | 62.90 | 18.20 | 40.55 | |
| X-VLA [ 75 ] | 68.00 | 20.90 | 44.45 | |
| ABot-M0 [ 61 ] | 57.40 | 30.36 | 43.88 | |
| [ 4 ] | 70.70 | 46.00 | 58.35 |
| Paradigm | Model | Clean (%) | Randomized (%) | Average (%) |
|---|---|---|---|---|
| VLA | [ 4 ] | 82.7 | 76.8 | 79.8 |
| ABot-M0 [ 61 ] | 86.1 | 85.1 | 85.6 | |
| LingBot-VLA [ 56 ] | 88.6 | 86.7 | 87.7 | |
| JoyAI-RA [ 71 ] | 90.5 | 89.3 | 89.9 | |
| HyVLA-0.5 [ 68 ] | 90.9 | 90.1 | 90.5 | |
| ACE-Ego-0 [ 28 ] | 91.1 | 90.6 | 90.9 |
| Paradigm | Model | Spatial | Object | Goal | Long | Average |
|---|---|---|---|---|---|---|
| VLA | [ 4 ] | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 |
| ABot-M0 [ 61 ] | 98.8 | 99.8 | 99.0 | 96.6 | 98.6 | |
| Qwen-VLA-Instruct [ 51 ] | – | – | – | – | 97.9 | |
| WAM | Fast-WAM [ 67 ] | 98.2 | 100.0 | 97.0 | 95.2 | 97.6 |
| LingBot-VA [ 29 ] | 98.5 | 99.6 | 97.2 | 98.5 | 98.5 | |
| LaWAM [ 11 ] | 99.4 | 99.6 | 98.4 | 97.0 | 98.6 |
| Condition | Trials | S | P | F | Success (%) |
|---|---|---|---|---|---|
| In-domain baseline | 10 | 8 | 2 | 0 | 80.0 |
| Cup instance | 10 | 7 | 3 | 0 | 70.0 |
| Placement | 10 | 7 | 2 | 1 | 70.0 |
| Background and clutter | 10 | 6 | 4 | 0 | 60.0 |
| All held-out conditions | 30 | 20 | 9 | 1 | 66.7 |
| Task | Setting | Grasp | Place | Stack | Full |
|---|---|---|---|---|---|
| Instructed Ranking | w/o subtask | 94.0 | 89.5 | – | 83.5 |
| w/ subtask | 93.5 | 97.5 | – | 91.0 | |
| Instructed Stacking | w/o subtask | 91.5 | – | 44.5 | 44.5 |
| w/ subtask | 98.5 | – | 58.5 | 58.5 | |
| Instructed Ranking & Stacking | w/o subtask | 60.0 | 37.0 | 24.5 | 1.5 |
| w/ subtask | 63.0 | 51.5 | 39.5 | 16.0 |
| Post-training data | Pretraining | Train loss | Held-out | Success (%) |
|---|---|---|---|---|
| 90 robot episodes | Robot | 0.055 | 0.0098 | 70 |
| 60 robot + 30 human | Robot | 0.068 | 0.0231 | 40 |
| 60 robot + 30 human | Human | 0.034 | 0.0164 | 60 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| VL-image | Future frame | |||
|---|---|---|---|---|
| Layers | Mass (%) | Enrichment | Mass (%) | Enrichment |
| 0–9 | 36.75 [36.45, 37.05] | 1.74 [1.73, 1.76] | 9.58 [9.49, 9.66] | 0.23 [0.23, 0.23] |
| 10 | 27.99 [27.72, 28.26] | 1.33 [1.32, 1.34] | 7.42 [7.34, 7.51] | 0.18 [0.17, 0.18] |
| 11 | 22.42 [22.22, 22.60] | 1.06 [1.06, 1.07] | 17.79 [17.64, 17.94] | 0.42 [0.42, 0.43] |
| 12 | 21.48 [21.30, 21.66] | 1.02 [1.01, 1.03] | 21.15 [20.89, 21.42] | 0.50 [0.50, 0.51] |
| 13 | 13.37 [13.15, 13.61] | 0.63 [0.62, 0.65] | 24.85 [24.65, 25.05] | 0.59 [0.58, 0.59] |
| VL-image | Future frame | Shallow mass (%) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Step | 0–9 | 10–19 | 20–29 | 0–9 | 10–19 | 20–29 | Text | Image | Crossover |
| 50 | 2.07 | 1.74 | 2.33 | 0.40 | 0.56 | 0.28 | 2.9 | 44.2 | — |
| 100 | 1.11 | 0.95 | 0.75 | 0.56 | 0.86 | 0.77 | 14.3 | 23.5 | 23 (71) |
| 150 | 1.10 | 0.61 | 0.61 | 0.53 | 0.89 | 0.81 | 22.6 | 23.3 | 11 (84) |
| 200 | 1.02 | 0.27 | 0.25 | 0.27 | 0.55 | 0.46 | 37.1 | 21.9 | 11 (95) |
| 250 | 1.04 | 0.21 | 0.31 | 0.15 | 0.46 | 0.45 | 45.8 | 22.0 | 12 (89) |
| Stage-level action-attention mass (%) | |||
|---|---|---|---|
| Stage | VL | Video | Action |
| Shallow (0–9) | 75.38 [75.24, 75.51] | 14.05 [13.93, 14.16] | 10.58 [10.54, 10.61] |
| Middle (10–19) | 36.55 [36.28, 36.79] | 24.31 [23.90, 24.76] | 39.14 [38.86, 39.41] |
| Deep (20–29) | 25.22 [25.02, 25.40] | 15.32 [14.83, 15.84] | 59.46 [59.00, 59.90] |
| Mechanism prevalence and transition location | |||
| Criterion | Coverage | Estimate | Uncertainty |
| Mask | Shallow (0–9) | Middle (10–19) | Deep (20–29) | Shallow/Deep |
|---|---|---|---|---|
| Asymmetric (ours) | 1.75 | 0.56 | 0.04 | 43.8 |
| Fully bidirectional | 0.81 | 0.70 | 0.67 | 1.2 |
| Protocol | Tasks | Evaluation/task | Test domain |
|---|---|---|---|
| Standard Clean | 50 | 100 | Clean |
| Standard Randomized | 50 | 100 | Randomized |
| Clean-to-Random | 50 | 100 | Randomized Hard |
| ACE-Ego-0 [ 28 ] | Fast-WAM [ 67 ] | WLA-0 [ 62 ] | EWAM | |||||
| Task | Clean | Rand. | Clean | Rand. | Clean | Rand. | Clean | Rand. |
| Adjust Bottle | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| Beat Block Hammer | 98 | 92 | 99 | 97 | 95 | 87 | 94 | 95 |
| Blocks Ranking RGB | 98 | 97 | 100 | 100 | 98 | 98 | 99 | 97 |
| Blocks Ranking Size | 89 | 91 | 94 | 98 | 93 | 85 | 76 | 85 |
| Click Alarmclock | 52 | 38 | 100 | 100 | 99 | 100 | 100 | 100 |
| Platform | Embodiment | Tasks |
|---|---|---|
| Franka | Single arm | Stack bowls; object into box |
| Dobot | Dual arm | Pour water; tidy desk; towel manipulation |
| Unitree G1-D | Humanoid dual arm | Kettle pouring; pour beans |
| Configuration | 10 steps (ms) | 5 steps (ms) |
|---|---|---|
| Baseline inference | 1668 | – |
| Without future-frame VAE decoding | 1128 | 578 |
| + graph-compatible RoPE and full-graph compilation | 461 | 254 |
| + chunk-overlap correction | 464 | 258 |
| Model | 5 steps (ms) | 10 steps (ms) | Per step (ms) |
| VLA | |||
| [ 4 ] | 50 | 60 | 1.9 |
| Xiaomi Robotics-1 [ 58 ] | 186 | 321 | 27.1 |
| LingBot-VLA-4B [ 56 ] | 293 | 504 | 42.3 |
| WAM | |||
| Fast-WAM [ 67 ] | 216 | 401 | 37.1 |
| Corpus | Source material | Hours | Anchor camera pose from |
|---|---|---|---|
| VITRA‑1M [ 30 ] | Ego4D, EPIC‑Kitchens, SSv2 | 283 | Inverted extrinsics |
| EgoDex [ 23 ] | Apple Vision Pro | 829 | Camera‑to‑world transform |
| Xperience [ 44 ] | Head‑mounted stereo rig | 972 | SLAM pose and calibration |
| Total | 2,084 |
| Mode | Robot Platform | Joint Dim. |
|---|---|---|
| Single Arm | Franka | 8 |
| UR5 | 7 | |
| ARX-5 | 7 | |
| Dual Arm | Franka | 16 |
| UR5 | 14 | |
| Agilex | 14 |