cs.AISep 29, 2026

Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation

Authors: Xiangcheng Zhan, Zirui Chen, Yicheng Zhao, Ziteng Gao, Shuo Yang

Organizations: Harbin Institute of Technology · Dalian University of Technology · Southern University of Science and Technology

Abstract

World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points and learns from successful and failed futures to support classifier-free guidance. An additional value head estimates task progress from video representations and activates guidance when progress stalls during inference. Across five DexJoCo tasks, DEWO improves average success across all three WAM formulations. Ablations show that visual supervision from successful and failed continuations improves both prediction and control beyond action supervision alone. On four real-world tasks across Wuji and Sharpa, 3 x 3 grid evaluations show that two rounds of deployment learning increase success from 51.0% to 71.7% in cells with at least one initial success, a gain of 20.7 percentage points. These findings support continued predictive learning for improving control through deployment experience, making world modeling an active part of WAM adaptation.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Latent evolving World Action Model

    Sep 23, 2026Xueji Fang, Boqiang Duan, Hua Wu +2Action GenerationWorld Models

  2. Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training

    Apr 23, 2026Yaxuan Li, Zhongyi Zhou, Yefei Chen +5Efficient World-Action ModelScalable Robot Learning

  3. WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT

    Jun 16, 2026Zezhong Qian, Xiaowei Chi, Yu Qi +3Efficient World-Action ModelFaster-Wam