Robotic foundation models offer a promising path toward general-purpose humanoid robot control, often through hierarchical architectures. However, their effectiveness depends on the command interface between the planner and the controller, which must support accurate execution while remaining easy to predict, and ideally allow new behaviors to be composed from prior ones. In this work, we introduce spectral skills, a latent representation of this interface that meets these requirements through predictive representation learning. By design, spectral skills compactly encode short motion segments and are learned by predicting subsequent motion rather than reconstructing the encoder input. On a 29-DoF humanoid, a controller conditioned on spectral skills reduces global tracking error by 62% relative to the state of the art. The same frozen controller chains independently encoded skills without a separate transition policy. It also composes new behaviors by adding orthogonal directions to any compatible base skill, producing combinations unseen in the training data. We demonstrate tracking, chaining, and composition, as well as control through a language-conditioned planner, on Unitree G1 hardware. Project page: https://spectral-skill.github.io
Figures & tables
Figure 1: Skill chaining and composition with one controller. One frozen controller chains skills and adds or removes steering, executing a loop in physics simulation.
Figure 2: Overview. (a) An encoder Eψ maps a state history to a skill z that linearly parameterizes a low-rank transition model. (b) One controller πθlo takes this skill from reference tracking, language-conditioned planning, or skill composition.
Figure 3: Real-robot deployment of the frozen controller on a Unitree G1. Top: tracking diverse whole-body motions. Middle: skill chaining, a side step, a 180∘ turn and another side step, by switching skill codes. Bottom: knobs composed on one base skill. From left to right: base; right arm up; right arm up with increasing left turns (three panels) and increasing right turns (two panels) with each instance being a separate trajectory.
Method
SR ↑
MPJPE-L ↓
MPJPE-G ↓
Any2Track 1
69.4
≈ 60
–
BeyondMimic 1
85.4
39.1
–
SONIC 1
99.2
23.8
–
124-motion capability set
SONIC
100.00
23.79
173.92
Ours
100.00
18.22
65.06
Table 1: Whole-body tracking. 1 Reported by SONIC on its own evaluation sets, not on these clips.
Direction
Measured effect
Direction
Measured effect
Left arm raise
shoulder pitch ±0.86 rad
Arms out
roll +0.45 rad/side, +0.43 (carry)
Both arms
shoulder pitch ±0.67 rad
High steps
swing apex 15→36 cm (L), 14→32 (R)
Right arm raise
shoulder pitch ±0.90 rad
Turn
±130∘ in 4 s
Right arm yaw
shoulder yaw ±0.43 – 0.51 rad †
Speed
+0.33 m/s, 6∘ heading change ( a=3 )
Table 2: Effects on the walk base at a=2 in Isaac PhysX.
Figure 4: One controller, many composed skills. Rows: base skills (walk, jog, squat, backward walk); columns: knobs added to the same frozen controller.
Figure 5: Steering composes skills rarely seen in the data. (a) Map of all 129,785 training clips (t-SNE), colored by motion kind; (b) walking corner. (c)–(d) One walk steered along the arm, turn and high-steps directions: (c) steered skill on the arm and turn directions, with contours of corpus walks that raise an arm or turn; (d) distance of the robot’s motion to the corpus (dashed: unsteered) and corpus windows showing each combination.
Interface
N
L ↓
G ↓
SR ↑
Skill tracker: 256-D z , hold 1
Oracle
—
17.19
51.7
1.000
Skills ( z→z )
30
38.41
251.2
0.914
Skills ( z→z )
10
42.83
230.2
0.911
Re-encode ( o→z )
10
60.19
425.1
0.771
Re-encode ( o→z )
21
88.64
796.8
0.546
Table 3: Language-conditioned planning interfaces. Arrows denote planner output → tracker command; z is skill and o is explicit root-and-joint frame. L/G: local/global MPJPE (mm).
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Reference data
129,785 retargeted BONES-SEED clips
Robot and simulator
29-DoF G1; Isaac Lab with the Newton/MuJoCo-Warp backend
Control frequency
50 Hz
Offline pretraining
50,000 updates; batch size 8,192
Offline optimizer
AdamW; weight decay 0; gradient norm clip 1
Encoder learning rate
3×10−4
Appendix
Table 4: Representation-learning settings.
Setting
Value
Tracker optimizer
AdamW; initial actor/critic LR 10−3
Actor LR adaptation
Per-iteration KL; target 0.02; limits [10−5,10−3]
Ours: critic LR
Linear decay 10−3→10−5 over the 50B budget
Ours: tracker weight decay
10−2
PPO clipping; GAE; discount
0.2 ; 0.95 ; 0.97
Entropy coefficient; gradient norm clip
0; 1
Appendix
Table 5: Tracker-training settings.
Prediction target
SR ↑
MPJPE-L ↓
MPJPE-G ↓
Ours
92.85
24.47
99.51
End-point
91.48
26.24
128.32
End-point-det
91.92
26.25
110.77
Next-skill
92.46
25.32
102.42
Appendix
Table 6: Prediction-target ablation. All arms, including Ours , share one interface and are scored at the same 2.0 B-frame checkpoint on the same 4096 -motion set. Ours here is not the 50B tracker of Table 1 .
Representation
Training
SR ↑
MPJPE-L ↓
MPJPE-G ↓
Ours
offline, cont. 64
92.85
24.47
99.51
Ours
offline, cont. 256
92.77
24.68
104.02
Ours
offline, FSQ
89.28
31.19
108.87
Ours
offline, cat.
68.48
46.97
273.29
Recon.
offline, cont.
93.53
22.93
278.98
Recon.
offline, FSQ
93.68
22.50
238.12
Appendix
Table 7: Representation-learning ablation.
Design axis
Variant
SR ↑
MPJPE-L ↓
MPJPE-G ↓
Ours
92.85
24.47
99.51
Skill width
Cont. 128-D
92.90
24.79
99.81
Cont. 256-D
92.77
24.68
104.02
FSQ
64×32
89.28
31.19
108.87
Codebook
Gumbel 64×32
69.60
44.03
180.60
Cat. 64×32
68.48
46.97
273.29
Appendix
Table 8: Additional interface design choices.
Planner output
SR (%) ↑
MPJPE-L ↓
MPJPE-G ↓
Oracle
–
20.47
90.5
FSQ latent, temporal ensembling
90.4
45.49
220.5
FSQ latent, no ensembling
87.7
52.30
246.5
Explicit re-encode
90.2
44.75
252.0
Appendix
Table 9: Language-conditioned planning with an FSQ tracker ( 64 -D, hold 10 ). Every planner row re-plans every N=10 control steps (one held FSQ skill, or 10 frames of the explicit chunk). Compare Table 3 .
Figure 6: Language-conditioned rollouts. The planner produces skills that the low-level policy executes: walking in a circle, lifting and carrying a box, opening a door while turning, and reaching from low to high.
Set
Method
SR
MPJPE-L
MPJPE-G
Root drift d
d2+L2
Final
124
SONIC
100.00
23.60
176.21
174.95
176.54
136
Ours
100.00
18.00
65.05
58.66
61.35
15
4096
SONIC
98.88
26.52
190.82
189.27
191.12
177
Ours
98.29
20.55
70.13
63.20
66.46
17
4096∗
SONIC
–
26.35
187.7
186.1
–
–
Ours
–
20.48
70.0
63.1
–
–
Appendix
Table 10: World-frame error decomposition, re-measured with matched frame indices (one seed, one evaluation per method). MPJPE and root drift d are frame-weighted means over successful clips. Final is the median, over successful clips, of the root error at the last frame. All errors in mm. ∗ The 4,013 clips of the 4096 set that both methods complete.
Figure 7: Global root tracking, ours vs. SONIC. (a, b) Root position error over the first 8 s of jointly completed clips lasting at least 8 s: mean (line) and interquartile range (band). In (a), a few SONIC clips that drift by over a meter lift its mean above the upper quartile. (c) MPJPE-L, root drift and MPJPE-G over successful clips. The tick marks drift2+L2 . (d) Mean root error per jointly completed clip of the 4096 set. The dashed line is equality.
Figure 8: Four long clips from the 4096 set, seen from above. Left: rollouts, one fixed camera per clip. Root paths: reference dashed, ours blue, SONIC orange. Robots are drawn at the numbered times, ours and SONIC shaded, the reference as an outline. The scale bar is 1 m at the depth of the nearest path point. Right: root position error, with the pose times marked. In (b) and (c) the poses span the part of the clip in which the reference moves; (d) shows one pose, at the time of our largest error.
Mean root error
Final root error
Clip
Reference motion
Ours
SONIC
Ours
SONIC
(a) Walk, random directions
28.3 s, 12.6 m path
3.9
43.8
0.7
58.3
(b) Jog forward
8.9 m in 4.8 s, then stands
2.7
62.1
1.6
76.2
(c) Walk sideways
8.3 m
2.2
55.4
1.3
89.5
(d) Balance on one foot
24.3 s
29.9
17.2
91.8
40.8
Appendix
Table 11: Root error of the four clips in Figure 8 , averaged over the clip and at the last frame, in cm.
Interface
Falls
Pose
Speed
Turn
Steering
0/66
1.00
0.95
279∘
Offset on action
17/66
1.30
0.78
162∘
Offset on servo target
5/66
0.76
0.91
68∘
Appendix
Table 12: Steering vs. matched joint offsets over 66 triples (5 bases, 5 directions, 3 amplitudes, and one wave gesture). Pose is the achieved pose relative to the steered run, so it is 1.00 for steering by definition; steered runs have 0 falls by construction. Speed is the median ratio to base for arm directions on walk, jog and backward walk. Turn is the turn direction’s median heading change. The speed ratios are not robust across seeds (Appendix G.1 ).
Method
Right arm
Spread
Speed
Turn
Ours, target-directed
9/12
9/10
4/12
6/11
Ours, spectral directions only
3/5
0/5
1/5
2/5
Joint offset (oracle)
2/10
8/9
0/10
5/10
SONIC, token PCA
4/5
2/5
0/5
0/5
SONIC, ridge directions
0/5
0/5
0/5
4/5
Appendix
Table 13: Steering pass rates at a≥2 on five bases (walk, jog, squat, carry, backward walk). Target-directed directions are computed from chosen output rows of Jξ ; spectral directions are response-Gram eigenvectors. Joint offset (oracle) adds the matched steered run’s mean joint deviation to the servo target. Denominators count runs: 10 Isaac runs ( a=2 and 3 per base, one joint-offset spread run lost), plus 2 , 2 and 1 walk runs of our right-arm, speed and turn steering in the MuJoCo deployment simulator. Other rows have one run per base at a=2 (SONIC in the MuJoCo simulator). Spectral-direction and token-PCA rows use directions chosen on walk; single seed.
Prompt
speed (m/s)
yaw ( ∘ )
R. shoulder pitch (rad)
other
Plain walk
0.60
22
+0.21
Raise right arm
0.69
81
−1.20
R. shoulder roll −0.61
Spread arms
0.64
49
+0.29
roll L +0.55 / R −0.50
Speed up
0.92
21
+0.27
Turn
0.61
88
+0.22
Appendix
Table 14: BFM-Zero reward prompts applied to a walk (hold window, single seed). Yaw is the heading change over the window.
Figure 9: A walk under our steering (top row) and under BFM-Zero’s reward prompts (middle row). Tiles show poses 1.2 – 3.6 s after onset. The bottom row shows ground paths from above (left turns bend upward). Numbers are changes against the same method’s plain walk over its hold window (ours 0.5 – 5.9 s, BFM-Zero 1.0 – 6.0 s after onset); plain-walk tiles instead give that walk’s own forward speed and heading change. Spread is the mean outward shoulder-roll change of both arms, and “speed × ” is the ratio of mean ground-path speeds. The “front” insets show the newest pose from the robot’s front right.
Line
λ or a
Speed
Heading
R. shoulder
L. shoulder
L. foot apex
(m/s)
( ∘ )
(rad)
(rad)
(cm)
Unsteered walk
–
0.71
−2
−0.05
−0.04
15.6
Blend toward standing, arm raised
0.5
0.26
−50
−1.09
+0.10
8.0
1
0.00
−11
−2.38
+0.27
0.0
Arm-raise clip − standing idle
0.5
0.71
−42
−0.96
+0.17
20.6
0.75
0.66
−55
−1.64
+0.33
23.7
Appendix
Table 15: Raising the right arm during the walk (Isaac PhysX, walk base, 400 control steps, one seed; means over steps 65 – 335 ). Blend: (1−λ)zwalk+λzarm ; skill differences: zwalk+λ(zB−zA) ; steering: zwalk+av , with v the right-arm spectral direction. Shoulder pitch is negative when the arm rises.
Stage
Speed (m/s)
Yaw ( ∘ /s)
Shoulder (rad)
Apex L/R (cm)
Distance
Support
Forward walk
Base
0.98±0.03
−8±2
0.05±0.01
15±0 / 16±0
2.73±0.09 [2.76]
90009, 92259, 93794 [90252]
+ right arm
1.07±0.01
3±1
−0.76±0.02
15±0 / 16±0
3.32±0.04 [3.37]
2334, 1988, 2124 [2180]
+ turn
1.06±0.03
35±2
−0.77±0.03
14±1 / 15±0
3.63±0.08 [3.60]
112, 102, 100 [119]
+ clearance
1.11±0.06
36±3
−0.80±0.04
19±2 / 24±1
3.93±0.03 [3.96]
19, 35, 95 [21]
Released
1.02±0.01
−1±2
0.01±0.02
16±0 / 15±1
2.87±0.15 [2.69]
–
Appendix
Table 16: Composition ladder over three randomized seeds (mean ± SD; brackets: deterministic run). Steering along the right-arm, turn and high-steps directions starts at 4 , 8 and 12 s ( a=2 , 1.2 and 1.2 ), and all three are released together at 17 s. Each row averages the steady control steps of one stage. Speed is horizontal path speed, yaw the mean yaw rate, shoulder the mean right shoulder pitch, and apex the swing-foot height range over 1 s windows (left/right). Distance is the median over the stage of the Euclidean distance from the re-encoded executed skill to the nearest of 263,395 skills from 4,000 corpus clips (corpus skills to the nearest skill of another clip: median 2.80 , p90 4.01 ). Support counts corpus 1 s windows (of 1,697,723 ) that move like the base and show every attribute moved so far, one entry per seed. No run fell.
Interface
Falls
Pose
Speed
Drift ( ∘ )
Turn ( ∘ )
Steering
0, 0, 0 [0]
1.00
0.95±0.02 [0.90]
16±3 [20]
295±10 [279]
Offset on action
4, 5, 5 [5]
1.37±0.03 [1.38]
0.98±0.10 [0.89]
36±8 [29]
192±4 [177]
Offset on servo target
1, 1, 2 [1]
0.73±0.00 [0.76]
0.92±0.01 [0.91]
23±3 [19]
92±13 [80]
Appendix
Table 17: Steering vs. matched joint offsets at a=2 over three randomized seeds (mean ± SD; brackets: deterministic run). The 22 triples cover five bases × five directions, minus three steered clearance runs that fell deterministically. Falls are per seed ( 0 , 1 , 2 ), out of 22 . Speed is the median ratio to base of arm directions on walk, jog and backward walk, drift the median ∣ heading change ∣ of non-turn directions, and turn the median heading change of the turn direction, all over runs that stayed up. Pose is relative to the seed’s steered run.
Setting
Frames
Success
MPJPE-L
MPJPE-G
Terminating, whole clip
6.0B
10/40
22.5†
85.2†
27.0B
17/40
22.1†
77.7†
No termination, whole clip
6.0B
23/40 fall-free
110.9 (42.7)
829.9 (398.4)
27.0B
24/40 fall-free
98.7 (27.8)
578.3 (142.0)
20×10 s windows per clip
6.0B
665/800 fall-free ‡
41.0 (31.0)
157.1 (127.6)
6.0B
594/800 complete
–
–
Appendix
Table 18: Our method trained and evaluated on the 40 LAFAN1 clips. MPJPE in mm, mean over clips with the median in parentheses. No-termination rows include steps after a fall. Window rows average the pre-fall steps. † Successful clips only, frame-weighted. ‡ Torso-height fall test: torso below 0.4 m while the reference pelvis is at or above 0.4 m. With the relative fall test, 688/800 windows are fall-free at 27.0 B.
BFM-Zero
Ours, 27.0B
Whole clip, terminating
Success
0/40
17/40
Mean steps before termination
1,028
7,904
First failure: orientation / end effector / pelvis height
24 / 15 / 1
2 / 19 / 2
MPJPE-L / G, successful clips
–
22.1 / 77.7
Fall-and-get-up clips completed
0/6
0/6
Appendix
Table 19: BFM-Zero (released checkpoint, its MuJoCo deployment simulator and retarget) and our 27.0B LAFAN1 model (Isaac Lab) on the 40 LAFAN1 clips. MPJPE in mm, frame-weighted over the scored steps of the stated clips. The torso-height rows exclude every clip in which the torso reaches the floor: five fall-and-get-up clips for BFM-Zero, all six for ours. ∗ From a separate run in which a fall ends the clip.
Learning from Demonstration (LfD) enables robots to learn complex behaviors from expert examples, yet existing approaches often fail to generalize to new compositions of known skills without retraining. Modern generative policies model distributions over action trajectories alone, thus are unable to reason about the symbolic outcomes required for robust composition. We propose that skills should jointly model action trajectories and the symbolic outcomes they induce. To address this gap, we introduce Predicate Action Skills (PACTS), a class of closed-loop visuomotor policies that model skills as a joint generative process over action and predicate belief trajectories, producing coherent action-outcome rollouts within a single model. Jointly generating actions and predicates enables PACTS to learn internal representations that improve both action generation and predicate classification. Furthermore, we demonstrate zero-shot composition of learned skills via planning by leveraging online predicate predictions from PACTS as a symbolic interface for sequencing and monitoring execution. Project website: https://planpacts.github.io/
Cross-task generalization is a core challenge in open-world robotic manipulation, and the key lies in extracting transferable manipulation knowledge from seen tasks. Recent in-context learning approaches leverage seen task demonstrations to generate actions for unseen tasks without parameter updates. However, existing methods provide only low-level continuous action sequences as context, failing to capture composable skill knowledge and causing models to degenerate into superficial trajectory imitation. We propose Decompose and Recompose, a skill reasoning framework using atomic skill-action pairs as intermediate representations. Our approach decomposes seen demonstrations into interpretable skill--action alignments, enabling the model to recompose these skills for unseen tasks through compositional reasoning. Specifically, we construct a task-adaptive dynamic demonstration library via visual-semantic retrieval combined with skill sequences from a planning agent, complemented by a coverage-aware static library to fill missing skill patterns. Together, these yield skill-comprehensive demonstrations that explicitly elicit compositional reasoning for skill composition and execution ordering. Experiments on the AGNOSTOS benchmark and real-world environments validate our method's zero-shot cross-task generalization capability.
Xitie Zhang, Aming Wu, Yahong Han
School of Artificial Intelligence, College of Intelligence and Computing, Tianjin University, China · School of Computer Science and Information Engineering, Hefei University of Technology, China
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcity of large-scale embodied trajectory datasets, leading to insufficient compositional generalization in out-of-distribution (OOD) scenarios with limited capability to capture reusable skill structures. To address this limitation, we propose Skill-Based Memory (SkillMemo) framework that implicitly decomposes long-horizon demonstrations into latent atomic skills and integrates skill-level features into a dynamic episodic memory bank for solving compositional tasks. Specifically, we first introduce an expert-guided trajectory segmentation module built upon a Mixture-of-Experts (MoE) architecture, which implicitly partitions trajectories into distinct skill primitives represented by learned gating coefficients. We further design a skill-level episodic memory architecture that stores compact skill representations as retrievable key-value pairs. During inference, the memory bank retrieves the most relevant skill primitives which are subsequently fused with the model's current gating distribution, providing a robust contextual prior to refine action predictions. Extensive experiments on the simulation benchmark and real-world manipulation tasks demonstrate that SkillMemo consistently enhances both DP and VLA backbones, achieving state-of-the-art performance and outperforming π0.5, while exhibiting strong compositional generalization to unseen task configurations.
Changyuan Wang, Chubin Zhang, Zhenyu Wu +8
Shenzhen International Graduate School, Tsinghua University · Department of Automation, Tsinghua University · Nanyang Technological University +1