World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, <1% drop) across these benchmarks while delivering significant end-to-end speedup (\eg, ∼25× on H100). Our code and checkpoints are available via this link.
Figures & tables
Figure 1: Inference efficiency of Fast-WAM ( left ) and Faster-WAM ( right ) across diverse hardware. Vertical bars compare inference latency between the original models and RealtimeWAM, while stacked bars show the runtime proportions of key components.
Figure 2: Overview of Teacher-Anchored Consistency Distillation (TACD). The frozen Video Expert provides shared KV conditioning, while the action student is trained with local consistency supervision and a teacher-anchored constraint derived from the frozen teacher’s multi-step rollout.
Figure 3: Training curves of two distillation methods: consistency loss ( Left ) and squared distance between the student’s velocity vθS and the teacher’s average velocity towards the clean endpoint uθT ( Right ).
Figure 4: Comparison of expert execution schedules. (a) Conventional execution waits for the complete video KV cache. (b) Cross-Expert Wavefront Pipelining releases block-wise KV cache entries after projection and synchronizes only before action attention consumes them, overlapping the Video and Action Experts.
Method
NFE
RoboTwin 2.0
LIBERO
Video
Action
Clean ↑
Random ↑
Overall ↑
Overall ↑
WAM and VLA Baselines
π0 ( Black et al., 2024 )
–
–
65.92
58.40
62.16
94.1
π0.5 ( Intelligence et al., 2025 )
–
–
82.74
76.76
79.75
96.9
X-VLA ( Zheng et al., 2026 )
–
–
72.90
72.80
72.85
98.1
Motus ( Bi et al., 2026 )
10
10
88.66
87.02
87.84
97.7
Table 1: Success rates (%) on RoboTwin 2.0 and LIBERO. NFE denotes the number of denoising steps for video and action generation. Methods marked with * and † are built on Fast-WAM and Faster-WAM, respectively.
Method
Camera ↑
Robot ↑
Lang. ↑
Light ↑
Backg. ↑
Noise ↑
Layout ↑
Overall ↑
UniVLA ( Bu et al., 2025 )
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
OpenVLA-OFT ( Kim et al., 2025 )
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
π0 ( Black et al., 2024 )
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
π0 -Fast ( Pertsch et al., 2025 )
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
WorldVLA ( Cen et al., 2025 )
0.1
27.9
41.6
43.7
17.1
10.9
38.0
25.0
Fast-WAM ( Yuan et al., 2026 )
18.8
45.7
70.1
83.2
45.7
29.8
62.7
49.1
Table 2: Success rates (%) on the seven perturbation subsets of LIBERO-Plus.
Table 7
Figure 5: Inference latency of RealtimeWAM built on Fast-WAM ( left ) and Faster-WAM ( right ) on a single NVIDIA H100 GPU.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
LCD
LTA
Video
Action
Training Time ↓
Peak Memory (GiB) ↓
✓
✗
✓
✓
8h 1min
65.98
✓
✗
✗
✓
5h 21min
24.73
✓
✓
✗
✓
8h 3min
24.73
Appendix
Table A.1: Training time and peak GPU memory for 30,000 training iterations on 16 NVIDIA H100 GPUs. LCD and LTA denote the consistency and teacher-anchored losses, respectively. Video/Action indicate whether each expert is fine-tuned.
Method
NFE
LIBERO
Video
Action
Spatial ↑
Object ↑
Goal ↑
LIBERO-10 ↑
Overall ↑
WAM and VLA Baselines
π0 ( Black et al., 2024 )
–
–
96.8
98.8
95.8
85.2
94.1
π0.5 ( Intelligence et al., 2025 )
–
–
98.8
98.2
98.0
92.4
96.9
X-VLA ( Zheng et al., 2026 )
–
–
98.2
98.6
97.8
97.6
98.1
Motus ( Bi et al., 2026 )
10
10
96.8
99.8
96.6
97.6
97.7
Appendix
Table F.1: Detailed success rates (%) on the four LIBERO task suites.
Rollout Interval
Clean ↑
Random ↑
Overall ↑
1→0
91.26
89.90
90.63
t→r
90.96
89.46
90.23
t→0
91.96
89.72
90.84
Appendix
Table F.2: Effect of teacher rollout intervals on RoboTwin 2.0. Results are success rates (%); best results are in bold.
TACD
CUDA Graph
CEWP
Efficient Kernels
Fast-WAM
Faster-WAM
RTX 5090
RTX 4090D
RTX 5090
RTX 4090D
✗
✗
✗
✗
251.5 ( 1.0× )
480.8 ( 1.0× )
214.1 ( 1.0× )
362.4 ( 1.0× )
✗
✓
✗
✗
97.8 ( 2.6× )
106.1 ( 4.5× )
81.5 ( 2.6× )
96.0 ( 3.8× )
✓
✗
✗
✗
54.3 ( 4.6× )
96.8 ( 5.0× )
60.7 ( 3.5× )
91.7 ( 4.0× )
✓
✓
✗
✗
30.8 ( 8.2× )
35.3 ( 13.6× )
38.5 ( 5.6× )
54.9 ( 6.6× )
✓
✓
✓
✗
24.3 ( 10.3× )
31.2 ( 15.4× )
34.7 ( 6.2× )
51.1 ( 7.1× )
Appendix
Table F.3: Inference latency (ms) on additional GPUs. Each entry reports latency followed by the cumulative speedup in parentheses, relative to the unoptimized baseline for the same model and GPU.
Figure F.1: Real-world garment folding with AgileX PiPER. Four keyframes from the recorded rollout show the initial unfolded T-shirt, the configuration after the first fold, a narrow configuration after further lengthwise folding, and the final folded configuration. The rollout was executed with the one-step RealtimeWAM * model. Consecutive actions proceed without perceptible pauses, demonstrating the smooth, low-latency inference that sustains long-horizon deformable-object manipulation.