World action models (WAMs) are large embodied policies that jointly predict future video and the actions to execute, emitting a fixed-length action chunk per inference call. Such a policy allocates its computational budget uniformly in time, unable to execute for longer over free-space motion or to spend more inference on contact-rich manipulation, which limits the throughput a WAM can reach when served in the cloud. We present SplineWAM, which adaptively compresses the action trajectory into a fixed-size window of cubic B-spline parameters, fitting the knot times to the characteristics of the motion. One parameter budget then decodes into chunks of varying temporal resolution and duration, and both the executed span and the interval until the next policy call follow from the prediction itself. Aligning the video supervision to the fitted knot times of the demonstration rather than to a uniform grid concentrates the supervised frames where the action trajectory is complex. For asynchronous deployment we introduce Jacobian-Pullback Real-Time Chunking (JP-RTC), which imposes chunk continuity on the decoded raw actions the robot executes rather than on the spline parameters, and corrects the parameters through the decoder so that the executed prefix agrees with the actions already committed. On LIBERO-Plus and RoboCasa, SplineWAM improves success rate over an action chunking WAM by 8.2 and 4.4 points while cutting policy calls per episode by 22% and 26%. On three bimanual real-robot tasks under asynchronous execution, it leads or matches the baseline while decoding 1.2 to 1.6 times as much executed motion per call.
Figures & tables
Figure 1: (a) A WAM predicts a dense, uniform chunk of discrete actions, whereas SplineWAM predicts a continuous trajectory from a few adaptively spaced knots and their control points. (b) Those knots also set the replanning schedule, so chunks stay long through free space and shorten on contact. (c) The result is fewer policy calls per episode at a higher success rate.
Figure 2: Three ways to sample the video stream against one action window. The time axis above marks every action step, with the predicted knots among them.
Figure 3: Where the continuity constraint is imposed; arrows show what each rule pulls together. (a) Naive RTC must align parameter rows, including those outside the committed steps whose cubic support reaches into them. (b) JP-RTC constrains the decoded actions instead, so the two windows need neither matching rows nor equal knot counts.
LIBERO-Plus ( ε=0.01 )
RoboCasa ( ε=0.08 )
Action representation
Success
Calls/ep.
Steps/call
Success
Calls/ep.
Steps/call
Action Chunking
66.3
33.6
8.00
69.1
21.9
16.00
BEAST
66.6
35.5
8.00
65.0
24.6
16.00
Naive B-Spline w/o video
60.0
38.7
7.92
55.0
32.7
14.26
Naive B-Spline
70.7
27.3
8.36
69.5
16.9
17.98
SplineWAM-Window
73.3
27.1
8.39
71.9
17.5
18.10
Table 1: Success rate and inference efficiency by action representation, under synchronous execution at the default tolerance of each suite. Success is episode-weighted, Calls/ep. counts policy-model invocations, and Steps/call denotes the mean control steps executed per invocation. SplineWAM-Window and SplineWAM-Knot differ only in video–action alignment.
Figure 4: Ablations on (a) RoboCasa and (b) LIBERO-Plus: success rate against the fitting tolerance ε (left), and against the two costs of the executed length (middle, right). Dashed line is the action chunking baseline; the circled point is the tolerance we deploy.
Figure 5: The three real-robot tasks, each stressing a different temporal profile: (a) long-horizon mobile manipulation, (b) high-precision alignment, and (c) repetitive motion under a visual decision.
Arrange Bookshelf
Charge Earphone
Clean Whiteboard
Model
SR (%)
PG (%)
Chunk
SR (%)
PG (%)
Chunk
SR (%)
PG (%)
Chunk
Action Chunking
20.0
52.2
32.0
26.7
73.3
32.0
50.0
67.5
32.0
SplineWAM, Naive RTC
10.0
35.6
39.8
13.3
43.3
46.5
46.7
68.3
53.3
SplineWAM, JP-RTC
23.3
54.4
38.6
26.7
57.5
45.7
60.0
74.2
52.4
Table 2: Real-robot results over 30 trials per task and method, all under asynchronous execution. Success rate (SR) requires every scored node of Table 6 ; progress (PG) is the fraction reached, averaged over trials. Chunk is the mean control steps executed per policy call.
Figure 6: Horizon length at every policy call of one successful Charge Earphone trial under SplineWAM with JP-RTC. Shaded bands mark the stages we identify from the recorded observations, with a keyframe above each stage’s extremum.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Task Performance
Efficiency
Action representation
BT
CV
LI
LC
OL
IS
SN
Average
Calls/episode
Steps/call
Action Chunking
61.1
43.5
70.8
87.4
75.9
66.6
63.8
66.3
33.6
8.00
BEAST
65.1
43.1
65.3
89.8
75.8
61.8
71.6
66.6
35.5
8.00
Naive B-Spline w/o video
55.6
32.9
71.0
85.1
66.7
53.7
61.0
60.0
38.7
7.92
Naive B-Spline
58.7
51.4
82.4
89.5
77.8
73.2
64.4
70.7
27.3
8.36
SplineWAM-Window
54.7
57.0
87.8
84.2
79.1
73.7
74.3
73.3
27.1
8.39
Appendix
Table 3: LIBERO-Plus success rate (%) by action representation, at ε=0.01 under synchronous execution, expanding Table 1 . Perturbation factors are background textures (BT), camera viewpoints (CV), language instructions (LI), light conditions (LC), objects layout (OL), initial states (IS) and sensor noise (SN); averages are episode-weighted.
Task Performance
Efficiency
Action representation
Pick-and-place
Open/close
Turn
Coffee
Average
Calls/episode
Steps/call
Action Chunking
53.1
89.7
72.7
62.0
69.1
21.9
16.00
BEAST
44.1
88.7
71.0
59.7
65.0
24.6
16.00
Naive B-Spline w/o video
34.2
66.7
65.7
62.3
55.0
32.7
14.26
Naive B-Spline
52.8
91.8
69.6
69.3
69.5
16.9
17.98
SplineWAM-Window
55.2
95.3
72.1
69.0
71.9
17.5
18.10
Appendix
Table 4: RoboCasa success rate (%) by action representation at ε=0.08 under synchronous execution, expanding Table 1 under the same protocol and efficiency measures as Table 3 . Columns are task families; averages are episode-weighted.
Task Performance
Efficiency
Setting
Spatial
Object
Goal
Long
Average
Calls/episode
Steps/call
Action representation (at ε=0.01 )
Action Chunking
97.4
99.4
95.0
93.2
96.3
19.8
8.00
BEAST
95.4
98.4
94.4
90.8
94.8
20.2
8.00
Naive B-Spline w/o video
86.6
96.8
94.2
84.2
90.5
21.6
7.15
Naive B-Spline
98.2
98.4
95.2
91.0
95.7
17.6
8.76
Appendix
Table 5: Original LIBERO success rate (%) across the synchronous experiments of Section 4 , with blocks matching Table 3 and Figure 4 . Efficiency was instrumented for a subset of the representations only; blank entries were not measured.
Task
Node 1
Node 2
Node 3
Node 4
Arrange Bookshelf
first book shelved
second book shelved
third book shelved
—
Charge Earphone
grasp plug
plug into socket
grasp USB-C
insert USB-C
Clean Whiteboard
erase two thirds
nearly all erased
fully erased
eraser returned
Appendix
Table 6: The scored nodes of each task, in order. Progress credits an equal share per node reached: 33.3% for Arrange Bookshelf and 25% for the other two.
Model & inference setting
Pick-and-place
Open/close
Turn
Coffee
Average
Action Chunking, sync
53.1
89.7
72.7
62.0
69.1
SplineWAM, sync
56.1
94.7
76.3
71.3
73.5
Action Chunking, async RTC
42.6
88.8
70.4
65.3
65.1
SplineWAM, async naive RTC
40.5
88.3
70.3
61.3
63.8
SplineWAM, async JP-RTC
43.5
87.7
73.3
63.3
65.7
Appendix
Table 7: RoboCasa success rate (%) under synchronous and asynchronous execution with a controlled inference delay. The asynchronous rows sit below the synchronous references by construction, so the comparison of interest is among the three of them, which differ in where the continuity residual is defined.
Model & inference setting
BT
CV
LI
LC
OL
IS
SN
Average
Action Chunking, sync
61.1
43.5
70.8
87.4
75.9
66.6
63.8
66.3
SplineWAM, sync
62.2
54.5
90.2
87.3
80.1
73.5
74.3
74.5
Action Chunking, async RTC
30.9
15.6
50.5
78.7
56.4
42.4
44.8
44.8
SplineWAM, async naive RTC
20.1
17.1
53.7
51.6
53.3
35.0
43.0
39.4
SplineWAM, async JP-RTC
25.7
27.9
65.5
61.7
59.7
39.3
50.5
47.5
Appendix
Table 8: LIBERO-Plus success rate (%) under synchronous and asynchronous execution, under the same protocol as Table 7 . Columns follow Table 3 .
World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, 3.62× the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.
Guoheng Sun, Chen Chen, Jin Wang +2
University of Maryland, College Park · Independent Researcher · Oxford Robotics Institute, University of Oxford
World action models (WAMs) that use future visual prediction at inference time incur substantial generation costs. Asynchronous execution reduces waiting by overlapping inference with robot motion, but visual predictions used for subsequent action generation must anticipate the effects of actions already scheduled for execution during inference. We introduce Streaming-WAM, which couples action-conditioned world modeling with asynchronous robot control to account for committed actions in future visual prediction. At each streaming update, the model conditions future visual prediction on the latest observation and the committed actions, which form the fixed prefix of the next action chunk. The resulting action-conditioned visual features guide generation of the remaining actions within the same joint update, so the continuation is informed by the scene changes expected during execution of the fixed prefix. On LIBERO, Streaming-WAM achieves an average success rate of 98.35% and reduces mean episode time by a factor of 2.93 relative to Fast-WAM. On the real-world Stamp Paper task, mean episode time falls from 90 s with synchronous Joint-WAM to 38 s with Streaming-WAM. These results show that Streaming-WAM supports efficient asynchronous control while maintaining high task success rates.
World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
Yinghua Zhou, Junjie Ye, Yiqi Zhao +8
University of Southern California · Brown University · Fudan University +1