World action models (WAMs) combine robot action generation with future state prediction. Existing WAMs typically predict videos or learned visual latents, which represent interaction geometry only implicitly and may retain appearance information unrelated to control. We introduce SkeleWAM, a compact WAM that represents a manipulation scene as a sparse 3D skeleton composed of robot joints, object centers, and interaction points. Constructed online from current RGB-D observations and robot proprioception, the skeleton provides a unified geometric state for action generation and future skeleton prediction. Future skeleton prediction provides additional geometric supervision for action learning without requiring visual reconstruction. At inference, SkeleWAM generates actions directly from the current skeleton and language instruction, while Medoid Action Consensus (MAC) serves as an auxiliary consensus strategy for stochastic action samples. On LIBERO-Plus, SkeleWAM achieves an overall success rate of 85.9% with 57.1M parameters, outperforming Cosmos-Policy by 3.7 percentage points. These results demonstrate that sparse 3D robot--object structure provides an effective state space for robust and parameter-efficient world action learning. The project is available at https://skelewam-project.github.io/.
Figures & tables
Figure 1 : Sparse skeleton states enable efficient world action modeling. Existing world action models represent future states as videos, visual latents, or dense 3D dynamics. In contrast, SkeleWAM converts RGB-D observations and robot states into a sparse 3D skeleton comprising robot joints, object centers, and interaction points. During training, it jointly learns robot action generation and future skeleton prediction. This explicit geometric representation achieves a favorable trade-off among task success, inference speed, model size, and computational cost.
Figure 2 : Overview of SkeleWAM. Frozen visual perception and forward kinematics construct the current skeleton. Action generation and future skeleton prediction are jointly trained with shared current-skeleton context. At inference, future skeleton prediction is omitted, and Medoid Action Consensus (MAC) selects a representative sampled action chunk. The robot executes its first He actions before observing and replanning.
Figure 3 : Qualitative real-world results. Representative rollouts for opening a drawer, placing a block in a drawer, and stacking bowls. Each row shows the early, interaction, and late stages of a successful rollout, together with the corresponding sparse 3D skeleton. Success rates are computed over 20 trials per task.
Method
Params
Camera
Robot
Language
Light
Background
Noise
Layout
Overall
OpenVLA [ 11 ]
≈7.5 B
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
OpenVLA-OFT [ 9 ]
≈7.7 B
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
NORA [ 7 ]
≈3.8 B
2.2
37.0
65.1
45.7
58.6
12.8
62.1
39.0
UniVLA [ 18 ]
≈7 B
1.8
46.2
69.6
69.0
81.0
21.2
31.9
43.9
π0 [ 1 ]
≈3.5 B
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
π0 -Fast [ 16 ]
≈3 B
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
Table 1 : Zero-shot success rates (%) on LIBERO-Plus. Parameter counts are approximate. Best observation-based results are bolded; § denotes privileged simulator input.
Figure 4 : Real-world experimental setup. An ARX R5 robot with external and wrist-mounted RealSense cameras operates in a tabletop workspace containing a drawer unit, blocks, and bowls.
Method
Open Drawer
Close Drawer
Stack Blocks
Stack Bowls
Put Block in Drawer
Average ↑
Fast-WAM
80
85
80
85
85
83
π0.5
80
80
70
75
75
76
Cosmos-Policy
85
85
90
90
85
87
SkeleWAM (ours)
80
90
90
95
90
89
Table 2 : Success rates (%) on five real-world manipulation tasks. Each method is evaluated over 20 trials per task. Average is the unweighted mean across tasks. Best results in each column, including ties, are bolded.
Configuration
Success (%)
(a) Object skeleton composition
Centers only
82.3
Interaction points only
77.7
Centers + interaction points (default)
85.9
(b) Training objective
Action only
80.1
Table 3 : Model ablations on the full LIBERO-Plus benchmark in the RGB-D setting. Each panel changes one design choice. Arrows indicate information flow between action tokens ( A ) and future-skeleton tokens ( S+ ). Inference uses Ha=32 , He=16 , K=3 , and h=10 .
Configuration
Success (%)
(a) Action execution horizon
He=5
80.7
He=16 (default)
85.9
He=32
81.8
(b) Number of MAC candidates
K=1 (without MAC)
84.2
Table 4 : Inference ablations on LIBERO-Plus in the RGB-D setting. We fix Ha=32 . Panel (a) varies He with K=3 and h=10 ; panel (b) varies K with He=16 and h=10 ; panel (c) varies h with He=16 and K=3 .