WAMJET: A Harness for World Action Model Acceleration
Authors: Le Chen, Lixin Liu, Jan Schneider, Zeju Qiu, Simon Guist, Bernhard Schölkopf, Dieter Büchler
Organizations: Max Planck Institute for Intelligent Systems, Tubingen, Germany · The Chinese University of Hong Kong, Hong Kong SAR, China · Johannes Kepler University Linz, Austria
World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tackle this bottleneck, we present WAMJET, an agentic harness that accelerates WAM inference by equipping coding agents with reusable optimization guidance and measurement and validation tools. WAMJET follows a bottleneck-driven workflow where the agent profiles inference, modifies targeted code, validates effects, and iteratively refines the acceleration stack as bottlenecks shift, while preserving action quality. Experiments span six WAMs, three coding agents, and two GPU architectures. WAMJET achieves up to 9.95x lossless speedup over upstream implementations. Approximation and hardware-aware optimization yield additional latency reductions, with comparable success rates. The results show that WAMJET can produce effective acceleration stacks for WAM deployment.
Figures & tables
Fig. 1 : Overview of WAMJET. (A) Manual optimization requires diverse domain expertise, with substantial engineering for each deployment. (B) A coding agent without guidance may only apply generic optimizations while leaving deeper optimization opportunities unexplored. (C) WAMJET equips the agent with reusable guidance and tools for faster startup, bottleneck analysis, hardware-aware optimization, and iterative validation, supporting acceleration strategies tailored to different WAMs and GPU architectures.
Fig. 2 : WAMJET’s bottleneck-driven optimization workflow. The agent reduces startup costs, profiles inference to identify bottlenecks, explores lossless acceleration followed by approximate techniques when permitted, and validates candidates. The agent retains accepted changes, reassesses the remaining bottlenecks, and continues searching within the specified budget.
Component
Function
Agent guidance
Run configuration
Defines policy, checkpoint, workload, quality requirements, and search budget limits.
Optimization skill
Provides procedures for bottleneck analysis, code changes, accept/revert decisions, and validation.
External memory
Records validated notes for reuse across optimization runs.
Execution and Analysis
Policy runner
Runs the policy with reference checks, timing measurement, and run records.
TABLE I : Functional components of the WAMJET harness.
LLM
WAM
Baseline
w/o WAMJET
w/ WAMJET Lossless (Ours)
Latency (ms)
Latency (ms)
Speedup
R1 Latency (ms) ∗
R2 Latency (ms) ∗
R1 / R2†
Speedup ‡
GPT-5.6-sol
Fast-WAM
84.7
59.2
1.43x
50.4
37.4
1.35x
2.27x
Cosmos-Policy
384.2
369.2
1.04x
236.2
163.1
1.45x
2.36x
DreamZero
2780.7
2687.7
1.03x
2373.9
2030.6
1.17x
1.37x
LingBot-VA
5323.6
2670.2
1.99x
755.8
686.5
1.10x
7.75x
Claude-opus-5
Fast-WAM
84.7
54.1
1.57x
47.5
43.1
1.10x
1.96x
TABLE II : Latency on H100 across baseline, w/o WAMJET, and w/ WAMJET Lossless across LLMs and WAMs.
WAM
GPU
Baseline
w/ WAMJET Lossless
w/ WAMJET Approx.
Latency (ms)
Latency (ms)
Speedup
Latency (ms)
Speedup
Approx. Method
DreamZero
H100
2780.7
1785.2
1.56x
1297.0
2.14x
FP8
DreamZero
B200
1598.5
823.9
1.94x
722.6
2.21x
NVFP4 + FP8
Cosmos3-Nano-Policy-DROID
H100
814.8
768.8
1.06x
651.2
1.25x
FP8
Cosmos3-Nano-Policy-DROID
B200
446.3
368.3
1.21x
304.7
1.47x
MXFP8
TABLE III : Latency of two WAMs under baseline and different WAMJET acceleration levels on two GPU architectures.
Fig. 3 : Optimization progress over search time for DreamZero (top) and OpenWAM (bottom) on B200.
WAM
Configuration
Approx. Method
Success Rate
DreamZero
Baseline
-
25.0%
w/ WAMJET Lossless
-
24.0%
w/ WAMJET Approx.
FP8
26.0%
w/ WAMJET Approx.
NVFP4 + FP8
25.0%
Cosmos3-Nano-Policy
Baseline
-
43.0%
w/ WAMJET Lossless
-
47.0%
TABLE IV : RoboLab subset success rates of Cosmos3-Nano-Policy-DROID and DreamZero under baseline and different WAMJET acceleration levels.
World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural design. More general caching strategies exploit feature redundancy, but redundancy alone does not capture the changing computational demands of closed-loop control. To address these challenges, we present RealtimeWAM, a general, training-free framework that coordinates parallel execution with adaptive computation for low-latency inference across diverse WAM architectures. We exploit layerwise dependencies to overlap observation processing with prediction. However, concurrent branches still compete for GPU resources, limiting the benefit of parallel execution. We therefore adapt computation throughout the pipeline through selective reuse, caching observation features in visually stable regions and reusing Transformer residuals while reserving additional refinement for small predicted adjustments. We evaluate RealtimeWAM on FastWAM and OpenWAM across RoboTwin, LIBERO, and LIBERO-Plus. On an RTX 4090, measured mean inference latencies are 24.09 and 63.09 ms, corresponding to average speedups of 8.90× and 10.67×. Average success rates are 82.75% and 87.41%, respectively, within 0.02 and 0.53 percentage points of native inference. Across five real-world tasks, RealtimeWAM improves average success rates over native inference by 17.2 and 37.2 percentage points on FastWAM and OpenWAM, respectively.
Huanan Liu, Ye Li, Kangye Ji +8
Tsinghua University · YuanxingGuangnian Robotics · Nanjing University
World Action Models (WAMs) build on pretrained video models, whose representations are grounded in physical dynamics and provide a natural basis for action prediction. Despite this natural foundation, many WAMs still rely on deep, parameter-heavy action-prediction modules that incur high inference latency and may overfit to limited robot demonstrations, restricting their real-world applicability. In this paper, we advocate a world-model-centric principle that concentrates capacity and computation in the video world model, while a lightweight action expert translates the backbone's representations into executable robot actions. We realize this principle through three key choices: Dock of Transformers (DoT) with Lite KV-Fusion to give the shallow, lightweight action expert access to representations from all video layers; world-model-only conditioning of the action expert; and retracted 1D-RoPE for positional alignment between video keys and action queries. We test this principle using Faster-WAM, a world-model-centric WAM with only a single-layer action expert. Despite this restriction on action-specific computation, Faster-WAM achieves competitive control performance on LIBERO and RoboTwin~2.0 without additional embodied pretraining. It provides approximately 3.7× and 1.3× inference speedups over Fast-WAM and π0.5, respectively. Consistent with its world-model-centric design, Faster-WAM demonstrates stronger generalizability under distribution shifts: the same LIBERO-trained policy achieves 78.3% success on LIBERO-Plus, exceeding Fast-WAM and LingBot-VA by 26.8 and 8.8 percentage points, respectively. Finally, real-robot experiments demonstrate success rates comparable to Fast-WAM, with substantially lower inference latency and shorter task-completion times.
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present WAMACHINE, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that WAMACHINE achieves 1.47-3.05× speedups in observation-to-action latency and 2.23-3.27× speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.
Zhinnan Liu, Haozhi Han, Ruge Zhang +8
Xiamen University · Institute for AI Industry Research, Tsinghua University · School of Computer Science, Peking University +2