World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, 3.62× the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.
Figures & tables
Figure 1: Success rate versus measured inference throughput on LIBERO. The measurement protocol, per-method results, and methods omitted from the figure are provided in Appendix C .
Figure 2: Overview of Staircase Policy . Here ti denotes a time step and ti∗ the predicted future.
Figure 3: Real-robot results on nine tasks. Task settings and details are provided in Appendix F .
Method
Background
Robot
Camera
Language
Noise
Layout
Light
Avg.
ST-WAM ( Wang et al., 2026b )
74.2
60.1
55.4
79.3
79.5
74.3
93.0
73.7
DreamWAM ( Yuan et al., 2026a )
71.5
63.6
53.7
94.8
67.1
80.7
96.6
75.4
VLA-JEPA ( Sun et al., 2026c )
93.6
67.1
63.3
85.4
66.3
85.1
95.6
79.5
ROCKET-VLA ( Sun et al., 2026a )
91.8
41.8
91.8
78.0
92.5
81.2
94.7
81.7
VLANeXt ( Wu et al., 2026 )
82.5
65.7
90.4
81.8
94.1
80.8
95.9
84.5
WorldPilot ( Lin et al., 2026c )
96.4
60.6
82.8
87.2
93.6
80.5
98.6
85.7
Table 1: Robustness on LIBERO-Plus across seven perturbation dimensions.
Figure 4: Success rate versus training steps across five global batch sizes.
Figure 6
Figure 6: Success rate versus execution horizon, averaged over four suites.
Figure 7: Comparison with compact VLA models on LIBERO, LIBERO-Plus, DOMINO.
Figure 8: SR against throughput for fixed and adaptive commitment on LIBERO-Long.
Success rate
Speed
Backbone
Policy
Spatial
Object
Goal
Long
Avg.
Act/s ↑
TTFA ↓
LaWAM
S-WAM
98.6
100.0
97.0
95.0
97.65
292.7
73.3
( Chen et al., 2026a )
Vanilla, Hexec=10
96.2
98.4
95.6
90.6
95.20
80.9
123.6
(2.3B)
Vanilla, Hexec=50
87.2
80.6
88.0
73.8
82.40
402.5
124.2
π0.5
S-WAM
98.0
98.6
95.2
94.2
96.50
243.2
80.6
( Intelligence et al., 2025 )
Vanilla, Hexec=10
94.6
99.0
91.6
87.4
93.15
75.7
132.2
Table 3: Staircase Policy transfers across different backbones. Act/s is executed actions per second; TTFA is time to first action (ms).
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
LIBERO
LIBERO-Plus
DOMINO
Cosine horizon
25k
25k
50k
S-WAM , steps trained
10k
10k
20k
S-WAM , reported checkpoint
5k
8k
20k
Vanilla, steps trained
25k
25k
50k
Vanilla, reported checkpoint
20k
25k
50k
Appendix
Table 4: Training-step budgets for the LaWAM host.
Training FLOPs (EFLOPs)
GPU-hours
S-WAM
Vanilla
Ratio
S-WAM
Vanilla
Ratio
LIBERO
3.75
5.02
0.75×
11.4
12.9
0.89×
LIBERO-Plus
3.75
5.02
0.75×
11.6
12.8
0.90×
DOMINO
8.74
12.99
0.67×
37.0
47.4
0.78×
Appendix
Table 5: Total training cost on the LaWAM backbone.
π0.5
FLOWER
Initialisation
released pi05_base
Florence-2-large with the official 360 k-step pretrained weights
Global batch
8
32
Steps
30 k
30 k ( 30 epochs of 1000 )
Reported checkpoint
ours 10 k, vanilla 30 k
ours epoch 10 ( ≈10 k), vanilla epoch 30
Learning rate
5×10−5→5×10−6 cosine, 1500 warmup
2×10−5 , the host’s three-stage schedule, weight decay 0.05
Optimiser
AdamW, gradient clipping 1.0
AdamW, betas (0.9,0.99)
Appendix
Table 6: Training settings for the two transfer hosts. Unlisted settings follow the published defaults.
Method
Spatial
Object
Goal
Long
Avg.
Hexec
Act./s
VRAM (GB)
S-WAM (LaWAM)
98.6
100.0
97.0
95.0
97.7
50
292.7
5.2
MiniCPM-RobotManip ( OpenBMB, 2026 )
–
–
–
–
97.5
30
166.8
3.7
JEPA-WAM ( Lin et al., 2026b )
95.6
99.4
97.2
94.6
96.7
20
150.3
4.3
Evo-Depth ( Lin et al., 2026a )
95.6
99.2
95.6
91.3
95.4
50
145.5
3.0
π0.5 ( Intelligence et al., 2025 ) †
98.8
98.2
98.0
92.4
96.9
5
132.1
9.5
OpenVLA-OFT ( Kim et al., 2025 )
97.6
98.4
97.9
94.5
97.1
8
114.7
16.1
Appendix
Table 7: LIBERO success rate and measured throughput across methods.
Figure 9: Per-suite breakdown of Fig. 6 .
Method
Spatial
Object
Goal
Long
Avg.
TurboVLA
97.4
99.4
96.2
93.2
96.5
VLA-Adapter
96.6
99.8
95.8
84.0
94.0
Evo-1
92.8
98.2
92.4
86.4
92.5
Vanilla, Hexec=10
96.2
98.4
95.6
90.6
95.2
Vanilla, full chunk ( H=50 )
87.2
80.6
88.0
73.8
82.4
Ours, full chunk ( H=50 )
98.6
100.0
97.0
95.0
97.7
Appendix
Table 8: Detailed results on LIBERO. Avg. is the unweighted mean over the four suites.
Method
Background
Robot
Camera
Language
Noise
Layout
Light
Avg.
TurboVLA
77.1
30.3
76.0
72.2
71.2
60.0
81.5
66.9
VLA-Adapter
89.0
38.4
89.3
66.6
91.9
73.1
87.5
76.5
Evo-1
92.6
37.7
87.2
61.1
89.7
60.9
91.7
74.4
Vanilla, Hexec=50
83.0
49.7
77.1
53.9
80.0
64.8
85.7
70.6
Vanilla, Hexec=10
95.2
73.4
91.5
77.7
94.3
82.1
96.0
87.2
S-WAM (ours), Hexec=50
96.6
76.0
89.5
85.2
90.4
79.3
98.6
87.9
Appendix
Table 9: Per-category LIBERO-Plus results averaged over the four LIBERO suites. Avg. is the unweighted mean of the seven categories.
Method
Background
Robot
Camera
Language
Noise
Layout
Light
Avg.
TurboVLA
91.5
28.3
69.9
82.8
68.9
54.3
92.5
69.7
VLA-Adapter
98.8
50.0
95.7
75.4
98.9
93.0
98.3
87.2
Evo-1
92.2
36.9
87.8
68.5
91.5
63.6
89.7
75.7
Vanilla, Hexec=50
88.0
49.4
83.0
53.3
80.1
78.2
94.2
75.2
Vanilla, Hexec=10
98.4
73.1
95.7
81.5
97.2
93.0
98.3
91.0
S-WAM (ours), Hexec=50
98.1
75.7
93.1
87.2
94.3
85.2
98.6
90.3
Appendix
Table 10: Per-category LIBERO-Plus results on LIBERO-Spatial . Avg. is the unweighted mean of the seven categories.
Method
Background
Robot
Camera
Language
Noise
Layout
Light
Avg.
TurboVLA
99.6
36.7
100.0
99.4
99.8
78.4
99.7
87.7
VLA-Adapter
96.0
27.1
97.2
84.5
96.9
75.7
95.3
81.8
Evo-1
96.8
27.1
95.2
77.7
93.4
72.2
98.7
80.1
Vanilla, Hexec=50
84.3
35.7
78.3
58.5
84.1
65.0
89.2
70.7
Vanilla, Hexec=10
99.6
72.1
98.5
81.1
98.6
91.1
99.7
91.5
S-WAM (ours), Hexec=50
96.8
71.9
94.7
88.1
97.4
83.4
100.0
90.3
Appendix
Table 11: Per-category LIBERO-Plus results on LIBERO-Object . Avg. is the unweighted mean of the seven categories.
Method
Background
Robot
Camera
Language
Noise
Layout
Light
Avg.
TurboVLA
45.9
10.5
54.2
28.3
40.6
35.5
47.7
37.5
VLA-Adapter
92.5
42.5
91.2
53.7
93.4
59.1
82.1
73.5
Evo-1
91.5
39.6
85.7
44.9
85.8
50.4
92.1
70.0
Vanilla, Hexec=50
87.5
61.9
81.9
48.3
83.9
58.8
82.1
72.1
Vanilla, Hexec=10
94.0
78.7
90.0
67.8
92.9
63.1
90.7
82.4
S-WAM (ours), Hexec=50
95.4
80.0
89.5
78.5
90.5
67.1
97.8
85.5
Appendix
Table 12: Per-category LIBERO-Plus results on LIBERO-Goal . Avg. is the unweighted mean of the seven categories.
Method
Background
Robot
Camera
Language
Noise
Layout
Light
Avg.
TurboVLA
71.3
45.5
79.7
78.3
75.5
71.8
86.1
72.6
VLA-Adapter
68.5
33.8
73.0
53.0
78.4
64.7
74.5
63.7
Evo-1
90.0
47.3
80.1
53.3
88.4
57.4
86.1
71.8
Vanilla, Hexec=50
72.3
51.7
65.4
55.4
71.7
57.1
77.4
64.4
Vanilla, Hexec=10
88.9
69.5
81.6
80.4
88.6
81.4
95.3
83.7
S-WAM (ours), Hexec=50
96.2
76.3
80.7
86.9
79.5
81.7
97.8
85.6
Appendix
Table 13: Per-category LIBERO-Plus results on LIBERO-Long . Avg. is the unweighted mean of the seven categories.
Method
adjust bottle
beat block
click alarm
click bell
grab roller
move can
move card
press stapler
rotate QR
Avg.
TurboVLA
0
0
8
0
0
0
0
2
0
1.11
VLA-Adapter
10
0
2
0
24
6
0
6
0
5.33
Evo-1
0
0
6
0
0
0
0
2
0
0.89
Vanilla, Hexec=10
0
0
0
0
12
0
0
8
0
2.22
Vanilla, full chunk ( H=75 )
66
6
4
0
30
14
6
14
6
16.22
Ours, full chunk ( H=75 )
60
24
12
2
38
2
12
20
4
19.33
Appendix
Table 14: Task-level success rates (%) on DOMINO. Avg. is the macro-average over the nine tasks.
Method
Act/s (p50)
TTFA (ms)
VRAM (GB)
TurboVLA
340.3
35.3
0.5
VLA-Adapter
91.3
87.7
3.6
Evo-1
41.9
334.1
2.4
Vanilla, Hexec=10
80.9
123.6
5.2
Vanilla, Hexec=50
402.5
124.2
5.2
Ours, Hexec=50
292.7
73.3
5.2
Appendix
Table 15: Measured inference efficiency across methods.
Task
Demos
Frames
Length
Span
Prompt
Stack bowl (static)
30
27,944
19.1 s
90%
stack the red bowl on the blue bowl
Stack bowl (dynamic)
30
23,073
14.6 s
69%
stack the red bowl on the blue bowl
Hang cup (static)
30
25,958
17.0 s
52%
hang the yellow cup on the cup rack
Hang cup (dynamic)
30
32,247
21.2 s
49%
hang the yellow cup on the cup rack
Place corn in bowl (static)
30
30,969
18.1 s
99%
place the corn in the bowl
Place corn in bowl (dynamic)
30
24,838
16.1 s
78%
place the corn in the bowl
Appendix
Table 16: The nine real-robot datasets. Length denotes the median episode duration. Span denotes the range over which manipulated objects are re-placed between demonstrations, measured as a percentage of image width. All data is recorded at 50 Hz.
Task
Counted as a success when
Stack bowl
the red bowl rests inside the blue bowl after release without tipping or falling outside
Hang cup
the cup remains on the rack arm after release
Place corn in bowl
the yellow corn is placed inside the bowl and remains there; grasping another object is a failure
Pour water
the water is poured into the blue cup without dropping the red cup or pouring outside the target
Put lid on the cup
the lid rests on the cup rim and covers the opening without falling onto the table
Place bowl in drawer
the drawer is opened and the bowl is placed inside; both stages are required
Appendix
Table 17: Success criteria for the real-robot evaluation. The static and dynamic versions of a task share the same criterion.
Figure 10: Stack bowl, static. The robot picks up the red bowl and places it inside the blue bowl.
Figure 14: Place corn in bowl, static. The robot picks up the yellow corn and places it in the bowl in the presence of distractor objects.
Task
Setting
Vanilla, Hexec=20
Vanilla, Hexec=50
S-WAM
Stack bowl
Static
85.0
72.5
90.0
Hang cup
Static
57.5
52.5
70.0
Place corn in bowl
Static
77.5
60.0
75.0
Put lid on the cup
Static
75.0
67.5
80.0
Pour water
Static
62.5
50.0
67.5
Place bowl in drawer
Static
47.5
40.0
52.5
Appendix
Table 18: Real-robot success rate (%) per task, each cell over 40 trials. S-WAM and vanilla at Hexec=50 are compute-matched; vanilla at Hexec=20 replans 2.5× as often.
World action models (WAMs) that use future visual prediction at inference time incur substantial generation costs. Asynchronous execution reduces waiting by overlapping inference with robot motion, but visual predictions used for subsequent action generation must anticipate the effects of actions already scheduled for execution during inference. We introduce Streaming-WAM, which couples action-conditioned world modeling with asynchronous robot control to account for committed actions in future visual prediction. At each streaming update, the model conditions future visual prediction on the latest observation and the committed actions, which form the fixed prefix of the next action chunk. The resulting action-conditioned visual features guide generation of the remaining actions within the same joint update, so the continuation is informed by the scene changes expected during execution of the fixed prefix. On LIBERO, Streaming-WAM achieves an average success rate of 98.35% and reduces mean episode time by a factor of 2.93 relative to Fast-WAM. On the real-world Stamp Paper task, mean episode time falls from 90 s with synchronous Joint-WAM to 38 s with Streaming-WAM. These results show that Streaming-WAM supports efficient asynchronous control while maintaining high task success rates.
World Action Models (WAMs) extend robot policy learning by incorporating future prediction as an additional training objective, encouraging the policy to encode task-relevant temporal structure in its representations. Current WAMs often rely on large-scale generative architectures that incur high training costs and inference latency, making them difficult to deploy as efficient closed-loop policies. We propose Light-WAM, a lightweight World Action Model for efficient robot manipulation. Specifically, it is built with a compact video backbone and performs future-video supervision in a downsampled latent space, reducing the cost of video co-training while retaining its benefits for representation learning. For action prediction, Light-WAM introduces the StateFusionActionExpert, which reads adapted states from multiple backbone layers, fuses them through learned-query pooling, and directly predicts action chunks in a single forward pass. This design provides an efficient interface between video backbone representations and robot actions, avoiding the need for heavy generative action experts. Experiments demonstrate that Light-WAM maintains strong performance on LIBERO and achieves usable multi-task performance on RoboTwin 2.0, while using only 0.44B trainable parameters. It also achieves 72.03ms inference latency with 4.1GiB peak GPU memory and improved training throughput.
Ziang Li, Dongzhou Cheng, Yibin Wang +5
Wuhan University · Shanghai Innovation Institute · Southeast University +2
World action models (WAMs) are large embodied policies that jointly predict future video and the actions to execute, emitting a fixed-length action chunk per inference call. Such a policy allocates its computational budget uniformly in time, unable to execute for longer over free-space motion or to spend more inference on contact-rich manipulation, which limits the throughput a WAM can reach when served in the cloud. We present SplineWAM, which adaptively compresses the action trajectory into a fixed-size window of cubic B-spline parameters, fitting the knot times to the characteristics of the motion. One parameter budget then decodes into chunks of varying temporal resolution and duration, and both the executed span and the interval until the next policy call follow from the prediction itself. Aligning the video supervision to the fitted knot times of the demonstration rather than to a uniform grid concentrates the supervised frames where the action trajectory is complex. For asynchronous deployment we introduce Jacobian-Pullback Real-Time Chunking (JP-RTC), which imposes chunk continuity on the decoded raw actions the robot executes rather than on the spline parameters, and corrects the parameters through the decoder so that the executed prefix agrees with the actions already committed. On LIBERO-Plus and RoboCasa, SplineWAM improves success rate over an action chunking WAM by 8.2 and 4.4 points while cutting policy calls per episode by 22% and 26%. On three bimanual real-robot tasks under asynchronous execution, it leads or matches the baseline while decoding 1.2 to 1.6 times as much executed motion per call.
Jun Guo, Xiaoshen Han, Qiwei Li +7
Tsinghua University · Xiaomi Robotics · Shanghai Jiao Tong University +2