Behavior Foundation Models (BFMs) give humanoids a promptable policy over a latent behavior space, enabling one single vector to represent a motion to imitate, a pose to reach, or a reward to maximize. Forward-Backward representations successfully produce such spaces, but at the cost of hundreds of GPU-hours for a single robot. Moreover, when the training process is repeated for a second robot, it produces a second space unrelated to the first, resulting in embodiment-specific latents that do not unify or transfer. We address these problems with CrossBFM, treating the latent space as the transferable asset for various embodiments. As retargeting provides frame-level cross-embodiment correspondence, we propose a unified encoder architecture with no robot-specific parameters for distilling the behavior space to address all training embodiments simultaneously in less than a GPU-hour. Following this encoder, latent-conditioned trackers turn the distilled latent into whole-body control in a conventional PPO training manner in just 10 more GPU-hours. On three distilled humanoids, all three prompting modes transfer: motion tracking with latent-conditioned policy losing only 0.025 rad to its joint-conditioned counterpart, smooth goal reaching between poses with no falls, and reward optimization for all 41 reward prompts. Our experiments further reveal that 1) regressing the encoder on a quarter of the motion corpus costs only 5% of tracking performance and 2) training the encoder on a subset of robots and evaluating on an unseen one recovers up to 89% of the tracking performance of seen robots, demonstrating cross-embodiment generalization to morphologically similar robots. We also verify the pipeline on real robots across all three prompting modes and with flow-based generated latents. Project website: https://dotandung.github.io/crossbfm/
Figures & tables
Figure 1: CrossBFM leverages a unified encoder to distill a frozen source BFM latent onto new robots. The distilled latent can process various prompts, including the three original modes – motion tracking, goal reaching, and reward optimization – and also a new mode for flow-based generated latents.
Figure 2: CrossBFM overview. Top: a unified encoder encoder mapping all new robots’ proprioception onto the frozen source latent using frame-level correspondences from retargeted data. Bottom: a z -conditioned policy and a flow-based latent generator on top of the distilled latent space.
Robot
DoF / bodies
Tracking MAE (rad) ↓
Goal MAE (rad) ↓
Joint (TWIST2)
Latent (ours)
Δcond
Latent (ours)
Δgoal
M3
27/30
0.1777±0.0033
0.2024±0.0067
+0.0247
0.2345±0.0011
+0.0321
T1
23/24
0.1844±0.0043
0.1901±0.0032
+0.0057
0.1967±0.0012
+0.0066
N1
23/29
0.1362±0.0032
0.1391±0.0038
+0.0029
0.1512±0.0023
+0.0121
Table 1: Latent-conditioned vs. joint-conditioned trackerby joint mean absolute error.
Robot
Source
Target (ours)
Gain
M3
0.145
0.508
3.50×
T1
0.228
0.261
1.14×
N1
0.177
0.523
2.95×
Mean
0.183
0.431
2.35×
Table 2: Reward optimization prompt results by normalized returns. Source-side computes zrew on the G1 with the true backward map BS ; target-side (ours) computes it via Equation 8 .
Encoder split
M3
T1
N1
oracle
infer.
oracle
infer.
oracle
infer.
Training ( 30 )
8.14
8.02
5.70
5.43
5.58
5.57
Validation ( 10 )
6.40
7.07
5.26
5.60
5.46
5.60
Table 3: Closed-loop key-body error mpjpe_local ↓ ( ×10−2 m) for the tracking task using the oracle latent z⋆ versus the encoder’s inference, split by training/validation set.
Figure 3: t-SNE projection of source (G1, dark) and distilled (M3, bright) latent trajectories, colored by behavior type. Dashed outlines represent validation clips ( unseen clips of similar behaviors).
Figure 4: Data scaling for the M3 robot, evaluated by closed-loop joint MAE as we shrink the distillation corpus from 100% to 5% to only noise.
Unseen robot
Validation cosine ↑
Closed-loop joint MAE (rad) ↓
Signal gap closed ↑
2-R E.
3-R E.
2-R E.
3-R E.
random z
M3
0.6251±0.0196
0.8658
0.1829±0.0054
0.1645
0.3371
89.3%±3.5
N1
0.5861±0.0100
0.8640
0.2237±0.0048
0.1370
0.3320
55.5%±3.2
T1
0.2533±0.0168
0.8300
0.2886±0.0138
0.1679
0.3063
12.8%±10.3
Table 4: Generalization to unseen robots. 2-R E. represents the encoder trained on two robots and evaluated on the third, with fixed pretrained trackers. 3-R E. is the unified encoders trained on all 3 robots.
Figure 5: Real-world rollouts of T1 and M3 across all 4 prompt modes. Videos in our website .
Mode
Prompt
Real
Sim
T (joint ref.)
5 clips
0.1985
0.1777
T (latent, ours)
5 clips
0.2206
0.2024
G
9 poses
0.2675
0.2345
R
3 tasks
0.562
0.508
Table 5: Sim2real results on 3 prompt modes: motion tracking (T), goal reaching (G) and reward optimization (R). Metrics: T and G are joint MAE in radians, R is normalized return.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Robot
Role
DoF
Bodies
Index filled
Morphology
Unitree G1
source
29
32
—
reference topology; 3 -DoF waist, 3 -DoF wrists
M3
target
27
30
27/33
nearest to the source; waist yaw only
Booster T1
target
23
24
21/33
wrist-less, 4 -DoF arms, 2 -DoF head
Fourier N1
target
23
29
23/33
one wrist DoF per arm; waist yaw only
Appendix
Table 6: The four humanoid configurations.
Range
Contents
Dims
[0:24)
key-body positions, root-relative, /\textscscale
8×3
[24:48)
key-body velocities (finite difference × fps)
8×3
[48:56)
key-body presence mask
8
[56:57)
root height /\textscscale
1
[57:61)
root rotation quaternion
4
[61:64)
root linear velocity /\textscscale
3
Appendix
Table 7: The unified cross-embodiment encoder canonical input.
Figure 6: Unified encoder architecture.
Component
Value
Input
Input width
163 (Table 7 )
Window length T
64 frames =2.13 s at 30 fps
Input projection
Linear(163→H)
Positional encoding
sinusoidal, max length 1024 , non-persistent buffer
Trunk
Appendix
Table 8: Encoder configuration.
Parameter
Value
Objective
L=N1∑t(1−cos(z^t,zt⋆))
Optimizer
AdamW
Learning rate
2×10−4
Weight decay
10−4
Gradient clipping
global norm 1.0
Epochs
400
Appendix
Table 9: Encoder training protocol across experiments.
Parameter
Value
Environment
Simulator
mjlab (MuJoCo), flat terrain
Physics rate
200 Hz (timestep 0.005 s)
Control rate
50 Hz (decimation 4 )
Episode length
10 s =500 control steps
Parallel environments
4096
Appendix
Table 10: Tracker configuration and PPO hyperparameters.
move-arms- θ - v -{l,m}-{l,m} for (θ,v)∈{(0,0.7),(90,0.7),(180,0.4),(−90,0.7)}
Rotation + arms
6
spin-arms- ±5 -{l-l, l-m, m-l}
Appendix
Table 13: The 41 source task rewards by group.
Robot
Low band
Medium band
M3
(0.497,0.824,0.276)
(1.157,∞,0.138)
N1
(0.416,0.627,0.240)
(0.925,∞,0.120)
T1
(0.509,0.675,0.172)
(0.820,∞,0.086)
Appendix
Table 14: Arm-height bands after quantile mapping, given as (lower, upper, margin) in metres.
Corpus
Encoder
Alignment ↑
Joint MAE (rad) ↓
Δ vs. oracle
100%
bidirectional
0.8695
0.2068
+0.0025
causal
0.8420
0.2043
+0.0000
25%
bidirectional
0.6825
0.2129
+0.0086
causal
0.6829
0.2045
+0.0002
—
oracle z⋆
—
0.2043
—
Appendix
Table 15: Causal versus bidirectional attention.
Per-robot
Unified
Δ
Alignment ↑
M3
0.9026±0.0017
0.8899±0.0040
−0.0127
T1
0.8307±0.0012
0.8265±0.0058
−0.0042
N1
0.8879±0.0011
0.8698±0.0043
−0.0181
Agreement ↑
all pairs
0.867 – 0.913
0.938 – 0.965
+0.07
Appendix
Table 16: Unified versus per-robot encoders.
H
M3
T1
N1
Mean
512
0.8788
0.8254
0.8603
0.8548
256 (ours)
0.8928
0.8304
0.8722
0.8652
128
0.6968
0.6722
0.6841
0.6844
Appendix
Table 17: Trunk width of the unified encoder, measured against z⋆ .
Corruption
Align ↑
Phase
cos@phase ↑↑
MAE (rad) ↓
Gap closed ↑
clean
0.8916
0
0.8916
0.1655
99.4%
bias ×1
0.8897
0
0.8897
0.1626
101.1%
bias ×2
0.8833
0
0.8833
0.1611
101.9%
bias ×4
0.8580
0
0.8580
0.1647
99.8%
bias ×8
0.7647
0
0.7647
0.1913
84.4%
bias ×16
0.5210
0
0.5210
0.2849
30.2%
Appendix
Table 18: Deployment regime. Pretrained encoder with corrupted inference input. Phase : per-clip offset that best match a clean clip; cos@phase : recovered cosine upon back-shifting.
Corruption
Align ↑
Phase
cos@phase ↑
MAE (rad) ↓
Gap closed ↑
clean
0.8632
0
0.8644
0.1669
98.6%
bias ×2
0.8525
0
0.8537
0.1740
94.4%
bias ×4
0.8242
0
0.8254
0.1848
88.2%
Gaussian σ=0.05
0.8427
0
0.8440
0.1675
98.2%
shift 2 fr
0.8191
+2
0.8673
0.1755
93.6%
shift −2 fr
0.8033
−2
0.8599
0.1673
98.3%
Appendix
Table 19: Training regime. The unified encoder retrained with M3 corrupted input.
Timing ( δ , k )
Geometry (bias ×N , σ )
Alignment
collapses ( 0.89→0.54 )
degrades smoothly
Closed-loop tracking
unchanged once re-phased
flat to ×8 , breaks by ×16
Recoverable
yes, by a one-dimensional search
not shown to be recoverable
Appendix
Table 20: Data corruption behavior across the two regimes.
Figure 7: Motion tracking and Goal reaching task for LAFAN motions.
Figure 8: Reward optimization for 3 tasks in the 41-task suite and 3 prompt conditions.
Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to further improve their capabilities, it remains unclear how key factors, including the learning paradigm, behavioral data and model architecture should be coordinated to enable effective scaling. In this work, we revisit the scaling recipe for BFMs and demonstrate that substantial performance gains can be achieved through the coordination of three core components: 1) the learning paradigm of motion tracking that reformulates diverse humanoid control problems as the reproduction of integrated whole-body behaviors in the global frame; 2) the strategic synergy between on-policy rollout quantity and reference motion diversity; and 3) the expressive and scalable model architecture termed Humanoid Transformer that facilitates the natural emergence of structured behavioral representations. Through extensive experiments in both simulation and real-world deployment, we demonstrate that our approach yields significant improvements in control fidelity and task generalization, reducing Mean Per-Keypoint Position Error (MPKPE) on the test set by over 10% in local mode and 82% in global mode compared with existing humanoid controllers. These results establish BFM as a principled and effective foundation for scalable and general-purpose humanoid control.
Weishuai Zeng, Kangning Yin, Xiaojie Niu +15
The Chinese University of Hong Kong · Galbot · Shanghai Artificial Intelligence Laboratory +4
As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAMs are sensitive to background visual noise, and the same motion from two different robots may be encoded with different latents. One solution to the background visual noise is to add an auxiliary loss predicting the robot action from the latent action, further associating the latent action space to the embodiment specific robot action space. We study a different use of the same labels, through action-similarity supervision. The similarity between any two latent actions is trained to match the similarity of the two ground-truth robot action sequences. The ground-truth actions are never predicted by the LAM, so the latent action does not need to encode embodiment specifics. We evaluate cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated. With the policy architecture and its hyperparameters, the dataset, and the evaluation protocol fixed, predicting latent actions instead of ground-truth actions more than doubles cross-embodiment success. Given the same ground-truth actions, similarity supervision transfers better than an auxiliary loss that predicts the ground-truth action during the LAM training. Computing the similarities on end-effector motion rather than joint-space motion, and letting the loss compare latent actions across the two robots, gives the best approach of the study.
Maxime Alvarez, Renzo Caballero, Tatsuya Matsushima +2
Graduate School of Engineering, The University of Tokyo
Behavioral Foundation Models (BFMs) offer a promising path toward universal physics-based character control by organizing a rich repertoire of physically plausible behaviors into a latent space, guided by a large-scale motion dataset. While these models excel at time-invariant tasks, such as goal-reaching and state-based reward optimization, their latent space does not directly support time-varying objectives, such as tracking a motion sequence. For tracking, existing heuristics rely on moving-window-averaging that fails to capture the nuances of highly dynamic motions. In this work, we propose a novel Latent Sequence Optimization (LSO) to address these shortcomings. Our approach combines simulation rollouts with a policy gradient update to optimize over a sequence of latents, extending the capabilities of BFMs toward precise motion tracking without requiring reward engineering and tuning. To guide the optimization toward smooth, coherent latent trajectories, we model the latent sequence using temporally correlated noise. We validate our approach across dense tracking, sparse keyframing, and direct deployment onto a real humanoid robot.