cs.CVSep 29, 2026

Rethinking Representations for World-Action Modeling

Authors: Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, +7 more

Organizations: Huazhong University of Science & Technology · Horizon Robotics · Fudan University · Xi'an Jiaotong University · D-Robotics

Abstract

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.

Figures & tables

Explore similar work

CardsList
  1. RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

    Jun 11, 2026Junke Wang, Qihang Zhang, Shuai Yang +5Efficient World-Action ModelDiscrete Action Tokenizers

  2. ReWorld: Learning Better Representations for World Action Models

    Jun 25, 2026Tianze Xia, Lijun Zhou, Kaixin Xiong +9Efficient World-Action ModelWorld Models

  3. Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

    Jun 10, 2026Lu Qiu, Yizhuo Li, Yi Chen +3Efficient World-Action ModelWorld Models