Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks
Organizations: University of Maryland, College Park · Independent Researcher · Oxford Robotics Institute, University of Oxford
Abstract
World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.
Figures & tables
| Method | Background | Robot | Camera | Language | Noise | Layout | Light | Avg. |
|---|---|---|---|---|---|---|---|---|
| ST-WAM ( Wang et al., 2026b ) | 74.2 | 60.1 | 55.4 | 79.3 | 79.5 | 74.3 | 93.0 | 73.7 |
| DreamWAM ( Yuan et al., 2026a ) | 71.5 | 63.6 | 53.7 | 94.8 | 67.1 | 80.7 | 96.6 | 75.4 |
| VLA-JEPA ( Sun et al., 2026c ) | 93.6 | 67.1 | 63.3 | 85.4 | 66.3 | 85.1 | 95.6 | 79.5 |
| ROCKET-VLA ( Sun et al., 2026a ) | 91.8 | 41.8 | 91.8 | 78.0 | 92.5 | 81.2 | 94.7 | 81.7 |
| VLANeXt ( Wu et al., 2026 ) | 82.5 | 65.7 | 90.4 | 81.8 | 94.1 | 80.8 | 95.9 | 84.5 |
| WorldPilot ( Lin et al., 2026c ) | 96.4 | 60.6 | 82.8 | 87.2 | 93.6 | 80.5 | 98.6 | 85.7 |
| Success rate | Speed | |||||||
|---|---|---|---|---|---|---|---|---|
| Backbone | Policy | Spatial | Object | Goal | Long | Avg. | Act/s | TTFA |
| LaWAM | S-WAM | 98.6 | 100.0 | 97.0 | 95.0 | 97.65 | 292.7 | 73.3 |
| ( Chen et al., 2026a ) | Vanilla, | 96.2 | 98.4 | 95.6 | 90.6 | 95.20 | 80.9 | 123.6 |
| (2.3B) | Vanilla, | 87.2 | 80.6 | 88.0 | 73.8 | 82.40 | 402.5 | 124.2 |
| S-WAM | 98.0 | 98.6 | 95.2 | 94.2 | 96.50 | 243.2 | 80.6 | |
| ( Intelligence et al., 2025 ) | Vanilla, | 94.6 | 99.0 | 91.6 | 87.4 | 93.15 | 75.7 | 132.2 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| LIBERO | LIBERO-Plus | DOMINO | |
| Cosine horizon | 25k | 25k | 50k |
| S-WAM , steps trained | 10k | 10k | 20k |
| S-WAM , reported checkpoint | 5k | 8k | 20k |
| Vanilla, steps trained | 25k | 25k | 50k |
| Vanilla, reported checkpoint | 20k | 25k | 50k |
| Training FLOPs (EFLOPs) | GPU-hours | |||||
|---|---|---|---|---|---|---|
| S-WAM | Vanilla | Ratio | S-WAM | Vanilla | Ratio | |
| LIBERO | 3.75 | 5.02 | 11.4 | 12.9 | ||
| LIBERO-Plus | 3.75 | 5.02 | 11.6 | 12.8 | ||
| DOMINO | 8.74 | 12.99 | 37.0 | 47.4 | ||
| FLOWER | ||
|---|---|---|
| Initialisation | released pi05_base | Florence-2-large with the official k-step pretrained weights |
| Global batch | ||
| Steps | k | k ( epochs of ) |
| Reported checkpoint | ours k, vanilla k | ours epoch ( k), vanilla epoch |
| Learning rate | cosine, warmup | , the host’s three-stage schedule, weight decay |
| Optimiser | AdamW, gradient clipping | AdamW, betas |
| Method | Spatial | Object | Goal | Long | Avg. | Act./s | VRAM (GB) | |
| S-WAM (LaWAM) | 98.6 | 100.0 | 97.0 | 95.0 | 97.7 | 50 | 292.7 | 5.2 |
| MiniCPM-RobotManip ( OpenBMB, 2026 ) | – | – | – | – | 97.5 | 30 | 166.8 | 3.7 |
| JEPA-WAM ( Lin et al., 2026b ) | 95.6 | 99.4 | 97.2 | 94.6 | 96.7 | 20 | 150.3 | 4.3 |
| Evo-Depth ( Lin et al., 2026a ) | 95.6 | 99.2 | 95.6 | 91.3 | 95.4 | 50 | 145.5 | 3.0 |
| ( Intelligence et al., 2025 ) † | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 | 5 | 132.1 | 9.5 |
| OpenVLA-OFT ( Kim et al., 2025 ) | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 | 8 | 114.7 | 16.1 |
| Method | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|
| TurboVLA | 97.4 | 99.4 | 96.2 | 93.2 | 96.5 |
| VLA-Adapter | 96.6 | 99.8 | 95.8 | 84.0 | 94.0 |
| Evo-1 | 92.8 | 98.2 | 92.4 | 86.4 | 92.5 |
| Vanilla, | 96.2 | 98.4 | 95.6 | 90.6 | 95.2 |
| Vanilla, full chunk ( ) | 87.2 | 80.6 | 88.0 | 73.8 | 82.4 |
| Ours, full chunk ( ) | 98.6 | 100.0 | 97.0 | 95.0 | 97.7 |
| Method | Background | Robot | Camera | Language | Noise | Layout | Light | Avg. |
|---|---|---|---|---|---|---|---|---|
| TurboVLA | 77.1 | 30.3 | 76.0 | 72.2 | 71.2 | 60.0 | 81.5 | 66.9 |
| VLA-Adapter | 89.0 | 38.4 | 89.3 | 66.6 | 91.9 | 73.1 | 87.5 | 76.5 |
| Evo-1 | 92.6 | 37.7 | 87.2 | 61.1 | 89.7 | 60.9 | 91.7 | 74.4 |
| Vanilla, | 83.0 | 49.7 | 77.1 | 53.9 | 80.0 | 64.8 | 85.7 | 70.6 |
| Vanilla, | 95.2 | 73.4 | 91.5 | 77.7 | 94.3 | 82.1 | 96.0 | 87.2 |
| S-WAM (ours), | 96.6 | 76.0 | 89.5 | 85.2 | 90.4 | 79.3 | 98.6 | 87.9 |
| Method | Background | Robot | Camera | Language | Noise | Layout | Light | Avg. |
|---|---|---|---|---|---|---|---|---|
| TurboVLA | 91.5 | 28.3 | 69.9 | 82.8 | 68.9 | 54.3 | 92.5 | 69.7 |
| VLA-Adapter | 98.8 | 50.0 | 95.7 | 75.4 | 98.9 | 93.0 | 98.3 | 87.2 |
| Evo-1 | 92.2 | 36.9 | 87.8 | 68.5 | 91.5 | 63.6 | 89.7 | 75.7 |
| Vanilla, | 88.0 | 49.4 | 83.0 | 53.3 | 80.1 | 78.2 | 94.2 | 75.2 |
| Vanilla, | 98.4 | 73.1 | 95.7 | 81.5 | 97.2 | 93.0 | 98.3 | 91.0 |
| S-WAM (ours), | 98.1 | 75.7 | 93.1 | 87.2 | 94.3 | 85.2 | 98.6 | 90.3 |
| Method | Background | Robot | Camera | Language | Noise | Layout | Light | Avg. |
|---|---|---|---|---|---|---|---|---|
| TurboVLA | 99.6 | 36.7 | 100.0 | 99.4 | 99.8 | 78.4 | 99.7 | 87.7 |
| VLA-Adapter | 96.0 | 27.1 | 97.2 | 84.5 | 96.9 | 75.7 | 95.3 | 81.8 |
| Evo-1 | 96.8 | 27.1 | 95.2 | 77.7 | 93.4 | 72.2 | 98.7 | 80.1 |
| Vanilla, | 84.3 | 35.7 | 78.3 | 58.5 | 84.1 | 65.0 | 89.2 | 70.7 |
| Vanilla, | 99.6 | 72.1 | 98.5 | 81.1 | 98.6 | 91.1 | 99.7 | 91.5 |
| S-WAM (ours), | 96.8 | 71.9 | 94.7 | 88.1 | 97.4 | 83.4 | 100.0 | 90.3 |
| Method | Background | Robot | Camera | Language | Noise | Layout | Light | Avg. |
|---|---|---|---|---|---|---|---|---|
| TurboVLA | 45.9 | 10.5 | 54.2 | 28.3 | 40.6 | 35.5 | 47.7 | 37.5 |
| VLA-Adapter | 92.5 | 42.5 | 91.2 | 53.7 | 93.4 | 59.1 | 82.1 | 73.5 |
| Evo-1 | 91.5 | 39.6 | 85.7 | 44.9 | 85.8 | 50.4 | 92.1 | 70.0 |
| Vanilla, | 87.5 | 61.9 | 81.9 | 48.3 | 83.9 | 58.8 | 82.1 | 72.1 |
| Vanilla, | 94.0 | 78.7 | 90.0 | 67.8 | 92.9 | 63.1 | 90.7 | 82.4 |
| S-WAM (ours), | 95.4 | 80.0 | 89.5 | 78.5 | 90.5 | 67.1 | 97.8 | 85.5 |
| Method | Background | Robot | Camera | Language | Noise | Layout | Light | Avg. |
|---|---|---|---|---|---|---|---|---|
| TurboVLA | 71.3 | 45.5 | 79.7 | 78.3 | 75.5 | 71.8 | 86.1 | 72.6 |
| VLA-Adapter | 68.5 | 33.8 | 73.0 | 53.0 | 78.4 | 64.7 | 74.5 | 63.7 |
| Evo-1 | 90.0 | 47.3 | 80.1 | 53.3 | 88.4 | 57.4 | 86.1 | 71.8 |
| Vanilla, | 72.3 | 51.7 | 65.4 | 55.4 | 71.7 | 57.1 | 77.4 | 64.4 |
| Vanilla, | 88.9 | 69.5 | 81.6 | 80.4 | 88.6 | 81.4 | 95.3 | 83.7 |
| S-WAM (ours), | 96.2 | 76.3 | 80.7 | 86.9 | 79.5 | 81.7 | 97.8 | 85.6 |
| Method | adjust bottle | beat block | click alarm | click bell | grab roller | move can | move card | press stapler | rotate QR | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|
| TurboVLA | 0 | 0 | 8 | 0 | 0 | 0 | 0 | 2 | 0 | 1.11 |
| VLA-Adapter | 10 | 0 | 2 | 0 | 24 | 6 | 0 | 6 | 0 | 5.33 |
| Evo-1 | 0 | 0 | 6 | 0 | 0 | 0 | 0 | 2 | 0 | 0.89 |
| Vanilla, | 0 | 0 | 0 | 0 | 12 | 0 | 0 | 8 | 0 | 2.22 |
| Vanilla, full chunk ( ) | 66 | 6 | 4 | 0 | 30 | 14 | 6 | 14 | 6 | 16.22 |
| Ours, full chunk ( ) | 60 | 24 | 12 | 2 | 38 | 2 | 12 | 20 | 4 | 19.33 |
| Method | Act/s (p50) | TTFA (ms) | VRAM (GB) |
|---|---|---|---|
| TurboVLA | 340.3 | 35.3 | 0.5 |
| VLA-Adapter | 91.3 | 87.7 | 3.6 |
| Evo-1 | 41.9 | 334.1 | 2.4 |
| Vanilla, | 80.9 | 123.6 | 5.2 |
| Vanilla, | 402.5 | 124.2 | 5.2 |
| Ours, | 292.7 | 73.3 | 5.2 |
| Task | Demos | Frames | Length | Span | Prompt |
|---|---|---|---|---|---|
| Stack bowl (static) | 30 | 19.1 s | 90% | stack the red bowl on the blue bowl | |
| Stack bowl (dynamic) | 30 | 14.6 s | 69% | stack the red bowl on the blue bowl | |
| Hang cup (static) | 30 | 17.0 s | 52% | hang the yellow cup on the cup rack | |
| Hang cup (dynamic) | 30 | 21.2 s | 49% | hang the yellow cup on the cup rack | |
| Place corn in bowl (static) | 30 | 18.1 s | 99% | place the corn in the bowl | |
| Place corn in bowl (dynamic) | 30 | 16.1 s | 78% | place the corn in the bowl |
| Task | Counted as a success when |
|---|---|
| Stack bowl | the red bowl rests inside the blue bowl after release without tipping or falling outside |
| Hang cup | the cup remains on the rack arm after release |
| Place corn in bowl | the yellow corn is placed inside the bowl and remains there; grasping another object is a failure |
| Pour water | the water is poured into the blue cup without dropping the red cup or pouring outside the target |
| Put lid on the cup | the lid rests on the cup rim and covers the opening without falling onto the table |
| Place bowl in drawer | the drawer is opened and the bowl is placed inside; both stages are required |
| Task | Setting | Vanilla, | Vanilla, | S-WAM |
|---|---|---|---|---|
| Stack bowl | Static | 85.0 | 72.5 | 90.0 |
| Hang cup | Static | 57.5 | 52.5 | 70.0 |
| Place corn in bowl | Static | 77.5 | 60.0 | 75.0 |
| Put lid on the cup | Static | 75.0 | 67.5 | 80.0 |
| Pour water | Static | 62.5 | 50.0 | 67.5 |
| Place bowl in drawer | Static | 47.5 | 40.0 | 52.5 |