World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, <1% drop) across these benchmarks while delivering significant end-to-end speedup (\eg, ∼25× on H100). Our code and checkpoints are available via this link.
Figures & tables
Figure 1: Inference efficiency of Fast-WAM ( left ) and Faster-WAM ( right ) across diverse hardware. Vertical bars compare inference latency between the original models and RealtimeWAM, while stacked bars show the runtime proportions of key components.
Figure 2: Overview of Teacher-Anchored Consistency Distillation (TACD). The frozen Video Expert provides shared KV conditioning, while the action student is trained with local consistency supervision and a teacher-anchored constraint derived from the frozen teacher’s multi-step rollout.
Figure 3: Training curves of two distillation methods: consistency loss ( Left ) and squared distance between the student’s velocity vθS and the teacher’s average velocity towards the clean endpoint uθT ( Right ).
Figure 4: Comparison of expert execution schedules. (a) Conventional execution waits for the complete video KV cache. (b) Cross-Expert Wavefront Pipelining releases block-wise KV cache entries after projection and synchronizes only before action attention consumes them, overlapping the Video and Action Experts.
Method
NFE
RoboTwin 2.0
LIBERO
Video
Action
Clean ↑
Random ↑
Overall ↑
Overall ↑
WAM and VLA Baselines
π0 ( Black et al., 2024 )
–
–
65.92
58.40
62.16
94.1
π0.5 ( Intelligence et al., 2025 )
–
–
82.74
76.76
79.75
96.9
X-VLA ( Zheng et al., 2026 )
–
–
72.90
72.80
72.85
98.1
Motus ( Bi et al., 2026 )
10
10
88.66
87.02
87.84
97.7
Table 1: Success rates (%) on RoboTwin 2.0 and LIBERO. NFE denotes the number of denoising steps for video and action generation. Methods marked with * and † are built on Fast-WAM and Faster-WAM, respectively.
Method
Camera ↑
Robot ↑
Lang. ↑
Light ↑
Backg. ↑
Noise ↑
Layout ↑
Overall ↑
UniVLA ( Bu et al., 2025 )
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
OpenVLA-OFT ( Kim et al., 2025 )
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
π0 ( Black et al., 2024 )
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
π0 -Fast ( Pertsch et al., 2025 )
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
WorldVLA ( Cen et al., 2025 )
0.1
27.9
41.6
43.7
17.1
10.9
38.0
25.0
Fast-WAM ( Yuan et al., 2026 )
18.8
45.7
70.1
83.2
45.7
29.8
62.7
49.1
Table 2: Success rates (%) on the seven perturbation subsets of LIBERO-Plus.
Table 7
Figure 5: Inference latency of RealtimeWAM built on Fast-WAM ( left ) and Faster-WAM ( right ) on a single NVIDIA H100 GPU.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
LCD
LTA
Video
Action
Training Time ↓
Peak Memory (GiB) ↓
✓
✗
✓
✓
8h 1min
65.98
✓
✗
✗
✓
5h 21min
24.73
✓
✓
✗
✓
8h 3min
24.73
Appendix
Table A.1: Training time and peak GPU memory for 30,000 training iterations on 16 NVIDIA H100 GPUs. LCD and LTA denote the consistency and teacher-anchored losses, respectively. Video/Action indicate whether each expert is fine-tuned.
Method
NFE
LIBERO
Video
Action
Spatial ↑
Object ↑
Goal ↑
LIBERO-10 ↑
Overall ↑
WAM and VLA Baselines
π0 ( Black et al., 2024 )
–
–
96.8
98.8
95.8
85.2
94.1
π0.5 ( Intelligence et al., 2025 )
–
–
98.8
98.2
98.0
92.4
96.9
X-VLA ( Zheng et al., 2026 )
–
–
98.2
98.6
97.8
97.6
98.1
Motus ( Bi et al., 2026 )
10
10
96.8
99.8
96.6
97.6
97.7
Appendix
Table F.1: Detailed success rates (%) on the four LIBERO task suites.
Rollout Interval
Clean ↑
Random ↑
Overall ↑
1→0
91.26
89.90
90.63
t→r
90.96
89.46
90.23
t→0
91.96
89.72
90.84
Appendix
Table F.2: Effect of teacher rollout intervals on RoboTwin 2.0. Results are success rates (%); best results are in bold.
TACD
CUDA Graph
CEWP
Efficient Kernels
Fast-WAM
Faster-WAM
RTX 5090
RTX 4090D
RTX 5090
RTX 4090D
✗
✗
✗
✗
251.5 ( 1.0× )
480.8 ( 1.0× )
214.1 ( 1.0× )
362.4 ( 1.0× )
✗
✓
✗
✗
97.8 ( 2.6× )
106.1 ( 4.5× )
81.5 ( 2.6× )
96.0 ( 3.8× )
✓
✗
✗
✗
54.3 ( 4.6× )
96.8 ( 5.0× )
60.7 ( 3.5× )
91.7 ( 4.0× )
✓
✓
✗
✗
30.8 ( 8.2× )
35.3 ( 13.6× )
38.5 ( 5.6× )
54.9 ( 6.6× )
✓
✓
✓
✗
24.3 ( 10.3× )
31.2 ( 15.4× )
34.7 ( 6.2× )
51.1 ( 7.1× )
Appendix
Table F.3: Inference latency (ms) on additional GPUs. Each entry reports latency followed by the cumulative speedup in parentheses, relative to the unoptimized baseline for the same model and GPU.
Figure F.1: Real-world garment folding with AgileX PiPER. Four keyframes from the recorded rollout show the initial unfolded T-shirt, the configuration after the first fold, a narrow configuration after further lengthwise folding, and the final folded configuration. The rollout was executed with the one-step RealtimeWAM * model. Consecutive actions proceed without perceptible pauses, demonstrating the smooth, low-latency inference that sustains long-horizon deformable-object manipulation.
World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural design. More general caching strategies exploit feature redundancy, but redundancy alone does not capture the changing computational demands of closed-loop control. To address these challenges, we present RealtimeWAM, a general, training-free framework that coordinates parallel execution with adaptive computation for low-latency inference across diverse WAM architectures. We exploit layerwise dependencies to overlap observation processing with prediction. However, concurrent branches still compete for GPU resources, limiting the benefit of parallel execution. We therefore adapt computation throughout the pipeline through selective reuse, caching observation features in visually stable regions and reusing Transformer residuals while reserving additional refinement for small predicted adjustments. We evaluate RealtimeWAM on FastWAM and OpenWAM across RoboTwin, LIBERO, and LIBERO-Plus. On an RTX 4090, measured mean inference latencies are 24.09 and 63.09 ms, corresponding to average speedups of 8.90× and 10.67×. Average success rates are 82.75% and 87.41%, respectively, within 0.02 and 0.53 percentage points of native inference. Across five real-world tasks, RealtimeWAM improves average success rates over native inference by 17.2 and 37.2 percentage points on FastWAM and OpenWAM, respectively.
Huanan Liu, Ye Li, Kangye Ji +8
Tsinghua University · YuanxingGuangnian Robotics · Nanjing University
World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off-the-shelf methods break down in the joint video-action setting because video and action streams use different SNR-shifted noise schedules and reach training with substantially different marginal noise distributions, an asymmetry that single-modality distillation methods cannot accommodate. We introduce \textbf{Flash-WAM}, a modality-aware step-distillation framework inspired by consistency distillation that selects the consistency function for each modality to match its noise regime: a linear-gradient-scaling parametrization for the action stream's low-noise regime, paired with a variance-preserving parametrization for the video stream's high-noise regime, grounded in a structural analysis of the consistency-function family that characterizes the achievable gradient scaling under the consistency boundary condition. Instantiated on LingBot-VA, Flash-WAM compresses inference to a single step in each modality. On RoboTwin 2.0, this reduces per-chunk latency from 8.1 seconds to 348 ms on NVIDIA L40S, a 23× speedup that enables real-time inference. Flash-WAM preserves task success on simulation benchmarks (85.5% RoboTwin 2.0, 95.7% LIBERO) and substantially recovers real-world performance (60% average on a Unitree G1 humanoid robot), while naive consistency distillation drops to 24% at the same step budget.
Arman Akbari, Ci Zhang, Arash Akbari +6
1Northeastern University · University of Georgia · 3EmbodyX Inc.
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference costly, full video prediction spends capacity on action-irrelevant temporal and appearance details, and long-horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better-matched prior: it only needs to model a target-frame transformation, focuses on action-relevant current-to-target visual differences, and grounds task instructions to localized visual changes through edit pretraining. In practice, ImageWAM does not decode the target frame at inference time; instead, it conditions a flow-matching action expert on the KV caches produced by image-editing denoising, using them as a compact world-action context. ImageWAM outperforms standard VLA baselines and matching competitive WAMs without additional policy pretraining across different simulator and real-world experiments. It also reduces FLOPs to 1/6 and latency to 1/4 of video-based WAMs. Attention analysis further shows that editing caches focus on task-relevant change regions, supporting image editing as an effective alternative to video-based world-action modeling.
Yuyang Zhang, Wenyao Zhang, Zekun Qi +7
Shanghai Jiao Tong University · Tsinghua University · Tencent Robotics X +1