Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.
Figures & tables
Figure 1 : Overview of Long-WAM. (1) AR video pretraining. LongLive2.0-Robot learns predictive dynamics from approximately 10,000 window-equivalent hours of robot and egocentric videos. (2) Causal-to-causal adaptation. We preserve causal video dependencies while conditioning action-chunk denoising on observed history. (3) Context scaling. Relative to current-observation-only control, success increases from 63.3% to 78.7% on RoboCasa GR-1 and from 94.5% to 99.5% on LIBERO-Long, peaking at 19.2 and 2.4 seconds of history, respectively. (4) Efficient deployment. Asynchronous execution, streaming VAE encoding, decision-time prefill, within-call KV reuse, and hardware-specific acceleration enable real-time control on RTX 5090, DGX Spark, and Jetson AGX Thor.
Figure 2 : Attention patterns and video–action inference strategies. (a) Action generation from the current observation only; (b) history-conditioned action generation without future prediction; (c) joint video–action co-denoising (CoD); (d) video prediction followed by action denoising (IDM). IDM conditions actions on both observed history and predicted future latents, reusing visual KV across action-denoising steps.
Figure 3 : Asynchronous model–robot execution with streaming VAE. Both schedules use the same nominal trigger stride S and overlap O=R−S . Streaming VAE (top) encodes observation (OBS) chunks as they arrive and meets the handoff deadline Tready≤OΔt . Full-window encoding (bottom) delays handoffs and subsequent triggers, accumulating robot idle time. Time is shown in units of Δt .
Figure 4 : Optimizations for efficient edge deployment. Shared optimizations and device-specific tuning accelerate edge inference. “Base” denotes the Quant/GEMM implementation; shape-specific dispatch and autotuning remain enabled.
Figure 5 : Simulation environments and real-world tasks. Our evaluation spans four simulation benchmarks—LIBERO, RoboTwin 2.0, DOMINO, and RoboCasa—and eight real-world task configurations on G1 and YAM. These include cup pickup at four conveyor speeds, dynamic cup stacking, bowl stacking, brick sorting by color, and dumpling placement into a pan.
Table 5 : Success rate (SR, %) on RoboCasa365. Overall averages 50 tasks: 18 Atomic-Seen, 16 Composite-Seen, and 16 Composite-Unseen. Baseline sources: Appendix D .
Figure 7 : Dynamic manipulation on Unitree G1. Long-WAM maintains 90–100% grasping success across conveyor speeds and achieves 95% success on dynamic cup stacking. Top: policy and human teleoperation comparisons. Bottom: Long-WAM and Fast-WAM rollouts.
Figure 8 : Long-horizon execution on YAM: 20 trials per task.
Optimization
RTX 5090
DGX Spark
Thor
E2E (ms) ↓
Speedup ↑
E2E (ms) ↓
Speedup ↑
E2E (ms) ↓
Speedup ↑
Shared optimizations
BF16 eager
356.0
1.0 ×
1342.8
1.0 ×
1215.2
1.0 ×
+ NVFP4 quantization + CUDA Graph + compile
172.4
2.1 ×
686.4
2.0 ×
762.4
1.6 ×
+ Denoising-invariant reuse
144.9
2.5 ×
464.5
2.9 ×
524.5
2.3 ×
+ Shared input quantization
130.3
2.7 ×
427.8
3.1 ×
476.5
2.6 ×
Table 8: End-to-end latency (E2E) of Long-WAM under cumulative optimizations across devices, including the full VAE computation. Each row keeps all preceding optimizations on a fixed input; speedups are relative to BF16 eager. The first optimized stage combines NVFP4 quantization, CUDA Graph replay, and compilation.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Source subset
Samples
RoVid-X [ 14 ]
1,898,562
AgiBot World: full146 [ 1 ]
69,834
AgiBot World: 5.2–32 s supplement
40,497
EgoDex [ 18 ]
207,044
EgoVerse [ 39 ]
59,742
VITRA [ 27 ]
19,210
Appendix
Table 9 : Pretraining data composition. Counts describe the full corpus before the train/validation/test split, not the published size of each source dataset.
Figure 9 : Task-conditioned video prediction with LongLive2.0-Robot. Five selected text-and-image-to-video (TI2V) examples, with the task prompt shown above each row. f denotes the zero-based frame index. All frames retain the original field of view.
S
Overlap O
SR (%) ↑
RMSE ↓
Jerk ↓
12
12
94.2
0.0246
0.0436
16
8
95.0
0.0063
0.0156
20
4
94.3
0.0058
0.0143
Appendix
Table 10 : Effect of asynchronous overlap length. With R=24 , increasing the trigger stride S shortens the overlap O . RMSE and jerk decrease throughout the tested range, whereas the highest observed success rate (SR) occurs at S=16 . Stride and overlap are measured in control steps.
V4/A4
V2/A2
Method
Latency (ms) ↓
Rate (Hz) ↑
Latency (ms) ↓
Rate (Hz) ↑
Long-WAM (IDM)
107.4
9.3
81.8
12.2
Long-WAM (CoD)
90.9
11.0
65.7
15.2
Appendix
Table 11 : Optimized IDM and CoD inference on RTX 5090. V4/A4 and V2/A2 use four and two steps per expert, respectively; CoD shares the denoising schedule across experts. Latency includes observation VAE computation. Rates are the reciprocals of inference latencies, not robot control frequencies.
Context P
History (s)
Latency (ms) ↓
0
0.0
74.6
48
2.4
107.4
96
4.8
138.3
192
9.6
204.5
384
19.2
341.0
Appendix
Table 12 : Context-dependent latency on RTX 5090. End-to-end inference time per action chunk with the Long-WAM infrastructure. Context P counts preceding control intervals, corresponding to P/20 seconds of history; P=0 retains the current observation.
Figure 10 : Dynamic cup stacking with π0.5 . Time progresses left to right within each four-column block, then continues in the next block below; each overview frame is accompanied by two wrist views. The policy grasps the blue cup but misses the moving green cup (red circle), leaving the composite task incomplete.
Figure 11 : Dynamic cup stacking with Fast-WAM. The sequence follows the same reading order as Figure 10 . After grasping the blue cup, the policy reaches toward the green cup but misses it as it moves along the conveyor. The subsequent frames show that the nesting stage is not completed.
Figure 12 : Dynamic cup stacking with Long-WAM. Long-WAM grasps the blue cup, intercepts the moving green cup, and places the green cup inside the blue cup. The overview and wrist views reveal the transition from interception to alignment and insertion, illustrating coordinated execution across the stages of this dynamic task.
Figure 13 : Moving-object grasping with π0.5 . Columns show 3.0, 4.5, 6.0, and 7.5 cm/s; time advances downward through paired overview and wrist views. The selected rollout succeeds at 3.0 cm/s, misses the cup at 4.5 and 7.5 cm/s, and exhibits the annotated gripper-stuck failure at 6.0 cm/s.
Figure 14 : Moving-object grasping with Fast-WAM. Columns and temporal ordering match Figure 13 . The selected rollouts succeed at 3.0 and 4.5 cm/s but miss the cup at 6.0 and 7.5 cm/s. The wrist views show the target moving beyond the gripper before a secure grasp is established.
Figure 15 : Moving-object grasping with Long-WAM. The illustrated rollouts complete the grasp at all four conveyor speeds, including 6.0 and 7.5 cm/s. Successive wrist views show the moving cup entering the gripper and being retained after closure, illustrating responsive interception under progressively tighter timing constraints.
Figure 16 : Long-horizon manipulation with Long-WAM on YAM. Top to bottom: brick sorting by color, placing dumplings in a pan, and stacking bowls. Each task contains six chronological snapshots, read left to right, with paired wrist views below each overview frame. The final snapshots show successful task configurations across all three tasks.