cs.ROSep 27, 2026

Achieve What You Imagined: Learning to Align Actions with Visual Plans

Authors: Yuheng Qiao, Ziran Wei, Xiaohan Wang, Daqiang Guo, Yichen Luo, Zhibo Pang, Peng Zhou, Sichao Liu

Organizations: KTH · Beihang University · The Hong Kong University of Science and Technology (Guangzhou) · Peking University · Great Bay University

Abstract

World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences that reliably achieve them. Consequently, we treat the WAM-generated visual prediction as a goal-conditioned visual proposal rather than a directly executable plan. We use a frozen action-conditioned world model to predict action-conditioned consequences and construct feedback based on consistency between the two future predictions and alignment with the terminal goal. Leveraging this feedback, we employ Flow Policy Optimization (FPO) to optimize the action head of the WAM. This framework avoids online robot interaction and additional training of task-specific reward models. Across four real-world UR5 manipulation tasks, our method increases the mean success rate from 43.4% to 75.1%, compared with 61.4% for π0.5π_{0.5}. These results show that cross-model prediction discrepancy can provide useful feedback for improving robot policies under the evaluated manipulation tasks. Website: https://imagine-to-achieve.github.io/

Figures & tables

Explore similar work

CardsList
  1. PACT-WAM: Predicting Actions and Visual Foresight with Compact Temporal Encoding for Robot Manipulation

    Feb 5, 2026Yushan Liu, Jingjing Fan, Shoujie Li +3Robotic ManipulationForesight

  2. When to Trust Imagination: Adaptive Action Execution for World Action Models

    May 7, 2026Rui Wang, Yue Zhang, Canyang Chen +3Efficient World-Action ModelWorld Models

  3. τ0τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation

    May 31, 2026Pengfei Zhou, Shengcong Chen, Di Chen +17Efficient World-Action ModelVideo World Models