cs.ROSep 30, 2026

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Authors: Hao Wang, Jiajun Wen, Jingzhi Liu, Shuoshuo Xue, Zhiliang Chen, Min Lin, Yicheng Chang, Xiaoyu Guo, +16 more

Organizations: Yinwang Intelligent Technology Co. Ltd.

Abstract

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embodied model whose asymmetric joint attention lets action tokens read semantic, current-visual, predicted-future, and action information at every layer while the perceptual experts retain their distinct roles. Without layer-wise supervision, EWAM develops an emergent depth-wise specialization: action queries attend mainly to vision-language features in shallow layers, to predicted future frames in intermediate layers, and to action tokens themselves in deep layers. This handoff replicates across tasks and is stable across denoising steps. Checkpoint tracking and causal interventions show that it is learned and that action generation depends on it. EWAM is pretrained in two separate regimes, one on cross-embodiment robot trajectories and one on human egocentric video. In simulation and real-robot experiments, it surpasses existing VLA, WAM, and hybrid baselines. Human egocentric data improve both cross-embodiment transfer and real-robot robustness, and subtask-phase supervision improves long-horizon completion. Together, these results suggest that unified embodied learning can induce an ordered internal progression from semantic understanding, through visual foresight, to action formation.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. World Action Models: The Next Frontier in Embodied AI

    May 12, 2026Siyin Wang, Junhao Shi, Zhaoyang Fu +11Recent World-Action ModelsWorld Models

  2. JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

    Aug 10, 2026Yihan Lin, Jiawei He, Shifeng Bao +6Efficient World-Action ModelRobot Policies

  3. UniWAM: Unified World-Action Model

    Oct 1, 2026Jiayi Chen, Wenxuan Song, Jingbo Wang +13Action PredictionWorld Models