Behavior Foundation Models (BFMs) give humanoids a promptable policy over a latent behavior space, enabling one single vector to represent a motion to imitate, a pose to reach, or a reward to maximize. Forward-Backward representations successfully produce such spaces, but at the cost of hundreds of GPU-hours for a single robot. Moreover, when the training process is repeated for a second robot, it produces a second space unrelated to the first, resulting in embodiment-specific latents that do not unify or transfer. We address these problems with CrossBFM, treating the latent space as the transferable asset for various embodiments. As retargeting provides frame-level cross-embodiment correspondence, we propose a unified encoder architecture with no robot-specific parameters for distilling the behavior space to address all training embodiments simultaneously in less than a GPU-hour. Following this encoder, latent-conditioned trackers turn the distilled latent into whole-body control in a conventional PPO training manner in just 10 more GPU-hours. On three distilled humanoids, all three prompting modes transfer: motion tracking with latent-conditioned policy losing only 0.025 rad to its joint-conditioned counterpart, smooth goal reaching between poses with no falls, and reward optimization for all 41 reward prompts. Our experiments further reveal that 1) regressing the encoder on a quarter of the motion corpus costs only 5% of tracking performance and 2) training the encoder on a subset of robots and evaluating on an unseen one recovers up to 89% of the tracking performance of seen robots, demonstrating cross-embodiment generalization to morphologically similar robots. We also verify the pipeline on real robots across all three prompting modes and with flow-based generated latents. Project website: https://dotandung.github.io/crossbfm/
Figures & tables
Figure 1: CrossBFM leverages a unified encoder to distill a frozen source BFM latent onto new robots. The distilled latent can process various prompts, including the three original modes – motion tracking, goal reaching, and reward optimization – and also a new mode for flow-based generated latents.
Figure 2: CrossBFM overview. Top: a unified encoder encoder mapping all new robots’ proprioception onto the frozen source latent using frame-level correspondences from retargeted data. Bottom: a z -conditioned policy and a flow-based latent generator on top of the distilled latent space.
Robot
DoF / bodies
Tracking MAE (rad) ↓
Goal MAE (rad) ↓
Joint (TWIST2)
Latent (ours)
Δcond
Latent (ours)
Δgoal
M3
27/30
0.1777±0.0033
0.2024±0.0067
+0.0247
0.2345±0.0011
+0.0321
T1
23/24
0.1844±0.0043
0.1901±0.0032
+0.0057
0.1967±0.0012
+0.0066
N1
23/29
0.1362±0.0032
0.1391±0.0038
+0.0029
0.1512±0.0023
+0.0121
Table 1: Latent-conditioned vs. joint-conditioned trackerby joint mean absolute error.
Robot
Source
Target (ours)
Gain
M3
0.145
0.508
3.50×
T1
0.228
0.261
1.14×
N1
0.177
0.523
2.95×
Mean
0.183
0.431
2.35×
Table 2: Reward optimization prompt results by normalized returns. Source-side computes zrew on the G1 with the true backward map BS ; target-side (ours) computes it via Equation 8 .
Encoder split
M3
T1
N1
oracle
infer.
oracle
infer.
oracle
infer.
Training ( 30 )
8.14
8.02
5.70
5.43
5.58
5.57
Validation ( 10 )
6.40
7.07
5.26
5.60
5.46
5.60
Table 3: Closed-loop key-body error mpjpe_local ↓ ( ×10−2 m) for the tracking task using the oracle latent z⋆ versus the encoder’s inference, split by training/validation set.
Figure 3: t-SNE projection of source (G1, dark) and distilled (M3, bright) latent trajectories, colored by behavior type. Dashed outlines represent validation clips ( unseen clips of similar behaviors).
Figure 4: Data scaling for the M3 robot, evaluated by closed-loop joint MAE as we shrink the distillation corpus from 100% to 5% to only noise.
Unseen robot
Validation cosine ↑
Closed-loop joint MAE (rad) ↓
Signal gap closed ↑
2-R E.
3-R E.
2-R E.
3-R E.
random z
M3
0.6251±0.0196
0.8658
0.1829±0.0054
0.1645
0.3371
89.3%±3.5
N1
0.5861±0.0100
0.8640
0.2237±0.0048
0.1370
0.3320
55.5%±3.2
T1
0.2533±0.0168
0.8300
0.2886±0.0138
0.1679
0.3063
12.8%±10.3
Table 4: Generalization to unseen robots. 2-R E. represents the encoder trained on two robots and evaluated on the third, with fixed pretrained trackers. 3-R E. is the unified encoders trained on all 3 robots.
Figure 5: Real-world rollouts of T1 and M3 across all 4 prompt modes. Videos in our website .
Mode
Prompt
Real
Sim
T (joint ref.)
5 clips
0.1985
0.1777
T (latent, ours)
5 clips
0.2206
0.2024
G
9 poses
0.2675
0.2345
R
3 tasks
0.562
0.508
Table 5: Sim2real results on 3 prompt modes: motion tracking (T), goal reaching (G) and reward optimization (R). Metrics: T and G are joint MAE in radians, R is normalized return.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Robot
Role
DoF
Bodies
Index filled
Morphology
Unitree G1
source
29
32
—
reference topology; 3 -DoF waist, 3 -DoF wrists
M3
target
27
30
27/33
nearest to the source; waist yaw only
Booster T1
target
23
24
21/33
wrist-less, 4 -DoF arms, 2 -DoF head
Fourier N1
target
23
29
23/33
one wrist DoF per arm; waist yaw only
Appendix
Table 6: The four humanoid configurations.
Range
Contents
Dims
[0:24)
key-body positions, root-relative, /\textscscale
8×3
[24:48)
key-body velocities (finite difference × fps)
8×3
[48:56)
key-body presence mask
8
[56:57)
root height /\textscscale
1
[57:61)
root rotation quaternion
4
[61:64)
root linear velocity /\textscscale
3
Appendix
Table 7: The unified cross-embodiment encoder canonical input.
Figure 6: Unified encoder architecture.
Component
Value
Input
Input width
163 (Table 7 )
Window length T
64 frames =2.13 s at 30 fps
Input projection
Linear(163→H)
Positional encoding
sinusoidal, max length 1024 , non-persistent buffer
Trunk
Appendix
Table 8: Encoder configuration.
Parameter
Value
Objective
L=N1∑t(1−cos(z^t,zt⋆))
Optimizer
AdamW
Learning rate
2×10−4
Weight decay
10−4
Gradient clipping
global norm 1.0
Epochs
400
Appendix
Table 9: Encoder training protocol across experiments.
Parameter
Value
Environment
Simulator
mjlab (MuJoCo), flat terrain
Physics rate
200 Hz (timestep 0.005 s)
Control rate
50 Hz (decimation 4 )
Episode length
10 s =500 control steps
Parallel environments
4096
Appendix
Table 10: Tracker configuration and PPO hyperparameters.
move-arms- θ - v -{l,m}-{l,m} for (θ,v)∈{(0,0.7),(90,0.7),(180,0.4),(−90,0.7)}
Rotation + arms
6
spin-arms- ±5 -{l-l, l-m, m-l}
Appendix
Table 13: The 41 source task rewards by group.
Robot
Low band
Medium band
M3
(0.497,0.824,0.276)
(1.157,∞,0.138)
N1
(0.416,0.627,0.240)
(0.925,∞,0.120)
T1
(0.509,0.675,0.172)
(0.820,∞,0.086)
Appendix
Table 14: Arm-height bands after quantile mapping, given as (lower, upper, margin) in metres.
Corpus
Encoder
Alignment ↑
Joint MAE (rad) ↓
Δ vs. oracle
100%
bidirectional
0.8695
0.2068
+0.0025
causal
0.8420
0.2043
+0.0000
25%
bidirectional
0.6825
0.2129
+0.0086
causal
0.6829
0.2045
+0.0002
—
oracle z⋆
—
0.2043
—
Appendix
Table 15: Causal versus bidirectional attention.
Per-robot
Unified
Δ
Alignment ↑
M3
0.9026±0.0017
0.8899±0.0040
−0.0127
T1
0.8307±0.0012
0.8265±0.0058
−0.0042
N1
0.8879±0.0011
0.8698±0.0043
−0.0181
Agreement ↑
all pairs
0.867 – 0.913
0.938 – 0.965
+0.07
Appendix
Table 16: Unified versus per-robot encoders.
H
M3
T1
N1
Mean
512
0.8788
0.8254
0.8603
0.8548
256 (ours)
0.8928
0.8304
0.8722
0.8652
128
0.6968
0.6722
0.6841
0.6844
Appendix
Table 17: Trunk width of the unified encoder, measured against z⋆ .
Corruption
Align ↑
Phase
cos@phase ↑↑
MAE (rad) ↓
Gap closed ↑
clean
0.8916
0
0.8916
0.1655
99.4%
bias ×1
0.8897
0
0.8897
0.1626
101.1%
bias ×2
0.8833
0
0.8833
0.1611
101.9%
bias ×4
0.8580
0
0.8580
0.1647
99.8%
bias ×8
0.7647
0
0.7647
0.1913
84.4%
bias ×16
0.5210
0
0.5210
0.2849
30.2%
Appendix
Table 18: Deployment regime. Pretrained encoder with corrupted inference input. Phase : per-clip offset that best match a clean clip; cos@phase : recovered cosine upon back-shifting.
Corruption
Align ↑
Phase
cos@phase ↑
MAE (rad) ↓
Gap closed ↑
clean
0.8632
0
0.8644
0.1669
98.6%
bias ×2
0.8525
0
0.8537
0.1740
94.4%
bias ×4
0.8242
0
0.8254
0.1848
88.2%
Gaussian σ=0.05
0.8427
0
0.8440
0.1675
98.2%
shift 2 fr
0.8191
+2
0.8673
0.1755
93.6%
shift −2 fr
0.8033
−2
0.8599
0.1673
98.3%
Appendix
Table 19: Training regime. The unified encoder retrained with M3 corrupted input.
Timing ( δ , k )
Geometry (bias ×N , σ )
Alignment
collapses ( 0.89→0.54 )
degrades smoothly
Closed-loop tracking
unchanged once re-phased
flat to ×8 , breaks by ×16
Recoverable
yes, by a one-dimensional search
not shown to be recoverable
Appendix
Table 20: Data corruption behavior across the two regimes.
Figure 7: Motion tracking and Goal reaching task for LAFAN motions.
Figure 8: Reward optimization for 3 tasks in the 41-task suite and 3 prompt conditions.