World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.
Figures & tables
Figure 1: Overview of AnyStep WAM. (a) Fixed budgets can waste computation on easy states or under-refine difficult ones. (b) Teacher-anchored interval supervision enables prediction across budgets. (c) Risk-benefit scheduling adapts denoising steps to scene difficulty and predicted fidelity.
Figure 2: Budget-aligned distillation and adaptive inference. (a) Teacher-trajectory displacements supervise finite-interval updates. (b) Shared LoRA adapters support multiple budgets, while teacher dynamics and student-teacher agreement provide offline risk and benefit targets. (c) Single-preview scheduling and recycling require exactly K⋆ denoising evaluations, without online teacher access or candidate rollouts.
Figure 3: Adaptive denoising rollouts. The left column shows representative observations. The middle column plots predicted video ( Cv ) and action ( Ca ) curvature risks, shaded regions marking high-risk stages. The right column shows student-teacher velocity similarities, combined by their minimum. Black boxes mark selected budgets.
Figure 4: Training stability and fixed-budget performance. (a) Per-sample target RMS distributions. (b) Gradient norms. (c) Average success rates at fixed denoising steps.
Figure 5: Real-world rollouts across six tasks on a Unitree G1D dual-arm robot. K denotes the denoising budget adaptively selected for each shown inference call.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Value
Optimizer
AdamW
Learning rate
2×10−5
Weight decay
0.01
Learning-rate schedule
Linear schedule
Training iterations
8,000
Training GPUs
4× NVIDIA H100
Appendix
Table 4: Hyperparameters and trainable parameter counts for LoRA-based AnyStep distillation. All original model parameters remain frozen during this stage.
Configuration
Value
Optimizer
AdamW
Learning rate
1×10−4
Weight decay
0.01
Learning-rate schedule
Cosine decay
Number of training GPUs
4
Batch size per GPU
64
Appendix
Table 5: Training hyperparameters of the adaptive scheduler. The teacher and AnyStep student remain frozen during scheduler training.
Figure 6: Internal architecture of the preview-based adaptive scheduler. Features and velocity predictions from a single-step AnyStep WAM preview are fused with instruction and robot-state context, projected into context tokens, and processed jointly with learned risk and budget queries by a shared Transformer. Query-specific heads predict modality-wise execution risk and the expected student-teacher similarity for each candidate denoising budget.
Task
Motus ( Bi et al., 2026 )
AnyStep-Motus
Clean
Rand.
Clean
Rand.
SR ↑
SR ↑
SR ↑
Steps ↓
SR ↑
Steps ↓
Adjust Bottle
89
93
89
5.52
95
4.53
Beat Block Hammer
95
88
97
3.49
93
4.03
Blocks Ranking RGB
99
97
100
2.82
95
3.11
Blocks Ranking Size
75
63
74
3.83
71
3.79
Appendix
Table 9: Task-wise comparison of Motus and AnyStep-Motus across 50 tasks. We report success rates (SR, %) under clean and randomized (Rand.) settings, and average denoising steps (Steps) for AnyStep-Motus. The final row reports the unweighted mean across tasks. Bold indicates the highest SR within each row and setting, including ties.
Task
FastWAM ( Yuan et al., 2026 )
AnyStep-FastWAM
Clean
Rand.
Clean
Rand.
SR ↑
SR ↑
SR ↑
Steps ↓
SR ↑
Steps ↓
Adjust Bottle
100
100
100
5.57
100
6.18
Beat Block Hammer
99
97
100
4.50
100
4.81
Blocks Ranking RGB
100
100
100
4.00
98
4.15
Blocks Ranking Size
94
98
93
4.55
92
4.81
Appendix
Table 10: Task-wise comparison of FastWAM and AnyStep-FastWAM across 50 tasks. We report success rates (SR, %) under clean and randomized (Rand.) settings, and average denoising steps (Steps) for AnyStep-FastWAM. The final row reports the aggregate results. Bold indicates the highest SR within each row and setting, including ties.
Task
LingBotVA ( Li et al., 2026b )
AnyStep-LingBotVA
Clean
Rand.
Clean
Rand.
SR ↑
SR ↑
SR ↑
Steps ↓
SR ↑
Steps ↓
Adjust Bottle
90
94
100
3.36
96
4.00
Beat Block Hammer
96
98
97
2.93
98
4.15
Blocks Ranking RGB
99
98
94
3.28
93
3.86
Blocks Ranking Size
94
96
97
3.27
81
3.81
Appendix
Table 11: Task-wise comparison of LingBotVA and AnyStep-LingBotVA across 50 tasks. We report success rates (SR, %) under clean and randomized (Rand.) settings, and average denoising steps (Steps) for AnyStep-LingBotVA. The final row reports the aggregate results. Bold indicates the highest SR within each row and setting, including ties.
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · 3Beijing Academy of Artificial Intelligence (BAAI) +3