WAMJET: A Harness for World Action Model Acceleration
Authors: Le Chen, Lixin Liu, Jan Schneider, Zeju Qiu, Simon Guist, Bernhard Schölkopf, Dieter Büchler
Organizations: Max Planck Institute for Intelligent Systems, Tubingen, Germany · The Chinese University of Hong Kong, Hong Kong SAR, China · Johannes Kepler University Linz, Austria
World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tackle this bottleneck, we present WAMJET, an agentic harness that accelerates WAM inference by equipping coding agents with reusable optimization guidance and measurement and validation tools. WAMJET follows a bottleneck-driven workflow where the agent profiles inference, modifies targeted code, validates effects, and iteratively refines the acceleration stack as bottlenecks shift, while preserving action quality. Experiments span six WAMs, three coding agents, and two GPU architectures. WAMJET achieves up to 9.95x lossless speedup over upstream implementations. Approximation and hardware-aware optimization yield additional latency reductions, with comparable success rates. The results show that WAMJET can produce effective acceleration stacks for WAM deployment.
Figures & tables
Fig. 1 : Overview of WAMJET. (A) Manual optimization requires diverse domain expertise, with substantial engineering for each deployment. (B) A coding agent without guidance may only apply generic optimizations while leaving deeper optimization opportunities unexplored. (C) WAMJET equips the agent with reusable guidance and tools for faster startup, bottleneck analysis, hardware-aware optimization, and iterative validation, supporting acceleration strategies tailored to different WAMs and GPU architectures.
Fig. 2 : WAMJET’s bottleneck-driven optimization workflow. The agent reduces startup costs, profiles inference to identify bottlenecks, explores lossless acceleration followed by approximate techniques when permitted, and validates candidates. The agent retains accepted changes, reassesses the remaining bottlenecks, and continues searching within the specified budget.
Component
Function
Agent guidance
Run configuration
Defines policy, checkpoint, workload, quality requirements, and search budget limits.
Optimization skill
Provides procedures for bottleneck analysis, code changes, accept/revert decisions, and validation.
External memory
Records validated notes for reuse across optimization runs.
Execution and Analysis
Policy runner
Runs the policy with reference checks, timing measurement, and run records.
TABLE I : Functional components of the WAMJET harness.
LLM
WAM
Baseline
w/o WAMJET
w/ WAMJET Lossless (Ours)
Latency (ms)
Latency (ms)
Speedup
R1 Latency (ms) ∗
R2 Latency (ms) ∗
R1 / R2†
Speedup ‡
GPT-5.6-sol
Fast-WAM
84.7
59.2
1.43x
50.4
37.4
1.35x
2.27x
Cosmos-Policy
384.2
369.2
1.04x
236.2
163.1
1.45x
2.36x
DreamZero
2780.7
2687.7
1.03x
2373.9
2030.6
1.17x
1.37x
LingBot-VA
5323.6
2670.2
1.99x
755.8
686.5
1.10x
7.75x
Claude-opus-5
Fast-WAM
84.7
54.1
1.57x
47.5
43.1
1.10x
1.96x
TABLE II : Latency on H100 across baseline, w/o WAMJET, and w/ WAMJET Lossless across LLMs and WAMs.
WAM
GPU
Baseline
w/ WAMJET Lossless
w/ WAMJET Approx.
Latency (ms)
Latency (ms)
Speedup
Latency (ms)
Speedup
Approx. Method
DreamZero
H100
2780.7
1785.2
1.56x
1297.0
2.14x
FP8
DreamZero
B200
1598.5
823.9
1.94x
722.6
2.21x
NVFP4 + FP8
Cosmos3-Nano-Policy-DROID
H100
814.8
768.8
1.06x
651.2
1.25x
FP8
Cosmos3-Nano-Policy-DROID
B200
446.3
368.3
1.21x
304.7
1.47x
MXFP8
TABLE III : Latency of two WAMs under baseline and different WAMJET acceleration levels on two GPU architectures.
Fig. 3 : Optimization progress over search time for DreamZero (top) and OpenWAM (bottom) on B200.
WAM
Configuration
Approx. Method
Success Rate
DreamZero
Baseline
-
25.0%
w/ WAMJET Lossless
-
24.0%
w/ WAMJET Approx.
FP8
26.0%
w/ WAMJET Approx.
NVFP4 + FP8
25.0%
Cosmos3-Nano-Policy
Baseline
-
43.0%
w/ WAMJET Lossless
-
47.0%
TABLE IV : RoboLab subset success rates of Cosmos3-Nano-Policy-DROID and DreamZero under baseline and different WAMJET acceleration levels.