World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.
Figures & tables
Figure 1: Overview of AnyStep WAM. (a) Fixed budgets can waste computation on easy states or under-refine difficult ones. (b) Teacher-anchored interval supervision enables prediction across budgets. (c) Risk-benefit scheduling adapts denoising steps to scene difficulty and predicted fidelity.
Figure 2: Budget-aligned distillation and adaptive inference. (a) Teacher-trajectory displacements supervise finite-interval updates. (b) Shared LoRA adapters support multiple budgets, while teacher dynamics and student-teacher agreement provide offline risk and benefit targets. (c) Single-preview scheduling and recycling require exactly K⋆ denoising evaluations, without online teacher access or candidate rollouts.
Figure 3: Adaptive denoising rollouts. The left column shows representative observations. The middle column plots predicted video ( Cv ) and action ( Ca ) curvature risks, shaded regions marking high-risk stages. The right column shows student-teacher velocity similarities, combined by their minimum. Black boxes mark selected budgets.
Figure 4: Training stability and fixed-budget performance. (a) Per-sample target RMS distributions. (b) Gradient norms. (c) Average success rates at fixed denoising steps.
Figure 5: Real-world rollouts across six tasks on a Unitree G1D dual-arm robot. K denotes the denoising budget adaptively selected for each shown inference call.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Value
Optimizer
AdamW
Learning rate
2×10−5
Weight decay
0.01
Learning-rate schedule
Linear schedule
Training iterations
8,000
Training GPUs
4× NVIDIA H100
Appendix
Table 4: Hyperparameters and trainable parameter counts for LoRA-based AnyStep distillation. All original model parameters remain frozen during this stage.
Configuration
Value
Optimizer
AdamW
Learning rate
1×10−4
Weight decay
0.01
Learning-rate schedule
Cosine decay
Number of training GPUs
4
Batch size per GPU
64
Appendix
Table 5: Training hyperparameters of the adaptive scheduler. The teacher and AnyStep student remain frozen during scheduler training.
Figure 6: Internal architecture of the preview-based adaptive scheduler. Features and velocity predictions from a single-step AnyStep WAM preview are fused with instruction and robot-state context, projected into context tokens, and processed jointly with learned risk and budget queries by a shared Transformer. Query-specific heads predict modality-wise execution risk and the expected student-teacher similarity for each candidate denoising budget.
Task
Motus ( Bi et al., 2026 )
AnyStep-Motus
Clean
Rand.
Clean
Rand.
SR ↑
SR ↑
SR ↑
Steps ↓
SR ↑
Steps ↓
Adjust Bottle
89
93
89
5.52
95
4.53
Beat Block Hammer
95
88
97
3.49
93
4.03
Blocks Ranking RGB
99
97
100
2.82
95
3.11
Blocks Ranking Size
75
63
74
3.83
71
3.79
Appendix
Table 9: Task-wise comparison of Motus and AnyStep-Motus across 50 tasks. We report success rates (SR, %) under clean and randomized (Rand.) settings, and average denoising steps (Steps) for AnyStep-Motus. The final row reports the unweighted mean across tasks. Bold indicates the highest SR within each row and setting, including ties.
Task
FastWAM ( Yuan et al., 2026 )
AnyStep-FastWAM
Clean
Rand.
Clean
Rand.
SR ↑
SR ↑
SR ↑
Steps ↓
SR ↑
Steps ↓
Adjust Bottle
100
100
100
5.57
100
6.18
Beat Block Hammer
99
97
100
4.50
100
4.81
Blocks Ranking RGB
100
100
100
4.00
98
4.15
Blocks Ranking Size
94
98
93
4.55
92
4.81
Appendix
Table 10: Task-wise comparison of FastWAM and AnyStep-FastWAM across 50 tasks. We report success rates (SR, %) under clean and randomized (Rand.) settings, and average denoising steps (Steps) for AnyStep-FastWAM. The final row reports the aggregate results. Bold indicates the highest SR within each row and setting, including ties.
Task
LingBotVA ( Li et al., 2026b )
AnyStep-LingBotVA
Clean
Rand.
Clean
Rand.
SR ↑
SR ↑
SR ↑
Steps ↓
SR ↑
Steps ↓
Adjust Bottle
90
94
100
3.36
96
4.00
Beat Block Hammer
96
98
97
2.93
98
4.15
Blocks Ranking RGB
99
98
94
3.28
93
3.86
Blocks Ranking Size
94
96
97
3.27
81
3.81
Appendix
Table 11: Task-wise comparison of LingBotVA and AnyStep-LingBotVA across 50 tasks. We report success rates (SR, %) under clean and randomized (Rand.) settings, and average denoising steps (Steps) for AnyStep-LingBotVA. The final row reports the aggregate results. Bold indicates the highest SR within each row and setting, including ties.
World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off-the-shelf methods break down in the joint video-action setting because video and action streams use different SNR-shifted noise schedules and reach training with substantially different marginal noise distributions, an asymmetry that single-modality distillation methods cannot accommodate. We introduce \textbf{Flash-WAM}, a modality-aware step-distillation framework inspired by consistency distillation that selects the consistency function for each modality to match its noise regime: a linear-gradient-scaling parametrization for the action stream's low-noise regime, paired with a variance-preserving parametrization for the video stream's high-noise regime, grounded in a structural analysis of the consistency-function family that characterizes the achievable gradient scaling under the consistency boundary condition. Instantiated on LingBot-VA, Flash-WAM compresses inference to a single step in each modality. On RoboTwin 2.0, this reduces per-chunk latency from 8.1 seconds to 348 ms on NVIDIA L40S, a 23× speedup that enables real-time inference. Flash-WAM preserves task success on simulation benchmarks (85.5% RoboTwin 2.0, 95.7% LIBERO) and substantially recovers real-world performance (60% average on a Unitree G1 humanoid robot), while naive consistency distillation drops to 24% at the same step budget.
Arman Akbari, Ci Zhang, Arash Akbari +6
1Northeastern University · University of Georgia · 3EmbodyX Inc.
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference costly, full video prediction spends capacity on action-irrelevant temporal and appearance details, and long-horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better-matched prior: it only needs to model a target-frame transformation, focuses on action-relevant current-to-target visual differences, and grounds task instructions to localized visual changes through edit pretraining. In practice, ImageWAM does not decode the target frame at inference time; instead, it conditions a flow-matching action expert on the KV caches produced by image-editing denoising, using them as a compact world-action context. ImageWAM outperforms standard VLA baselines and matching competitive WAMs without additional policy pretraining across different simulator and real-world experiments. It also reduces FLOPs to 1/6 and latency to 1/4 of video-based WAMs. Attention analysis further shows that editing caches focus on task-relevant change regions, supporting image editing as an effective alternative to video-based world-action modeling.
Yuyang Zhang, Wenyao Zhang, Zekun Qi +7
Shanghai Jiao Tong University · Tsinghua University · Tencent Robotics X +1
World Action Models (WAMs) use video generation models to predict future visual dynamics for robotic manipulation, but iterative denoising introduces additional latency for closed-loop control. We empirically find that visual content converges at different rates during denoising. Static background structure forms early, whereas the gripper and manipulated object remain blurry after the first step, with their interaction dynamics emerging only through subsequent denoising. Consequently, naively truncating a multi-step video model to one step preserves scene structure but loses the interaction-centric dynamics most critical for manipulation. To address this issue, we propose DIDO, which distills the converged dynamics of a multi-step video model into a single denoising step. DIDO combines distribution matching distillation with interaction-centric representation guidance. Beyond compressing multi-step generation into one forward pass, DIDO explicitly models the gripper, manipulated object, and their interaction using supervised bounding-box visual reasoning tokens. Additionally, DIDO aligns the target object's representations across multiple model layers with features from a pretrained DINOv3 encoder. This interaction-centric guidance helps the distilled model preserve both the relevant entities and their future dynamics in a single step, while substantially reducing inference latency. DIDO achieves an average success rate of 99.0% on LIBERO, 76.6% on LIBERO-Plus, and 92.0% on RoboTwin, while also demonstrating effective transfer to long-horizon and generalization tasks in real-world robotic manipulation.
Jing Lyu, Shuanghao Bai, Runze Xiao +11
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · 3Beijing Academy of Artificial Intelligence (BAAI) +3