Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art models struggle with reliability in zero-shot settings, i.e. skeletons with different topologies which were unseen during training, and recent Transformer-based attempts have failed to outperform specialized geometric methods. We bridge this gap with a Transformer Autoencoder that learns a topology- and translation-invariant latent space. Our core contribution is a learnable flattening of skeletal graphs that captures both local dependencies and global structure. Unlike the standard transformer architecture, which adds positional information to token content, we integrate graph-based positional encodings multiplicatively, a design choice that follows directly from our flattening formulation. The resulting model handles diverse skeletal topologies within a single unified architecture and trains in a fully unsupervised manner, requiring no paired retargeting data. Ablation studies show, that the graph encodings, multiplicative formulation, and Transformer backbone is critical for the performance. In zero-shot evaluations, our method reduces global joint position error by 43−47% over current benchmarks. A user study (n=37), including expert animators, further ranks our approach highest in motion alignment and physical plausibility (p<0.05). These results demonstrate that our model design is key to making transformer architectures effective for motion retargeting, outperforming existing approaches.
Figures & tables
Figure 1: Our method retargets a source motion (left) to characters with vastly different skeletal topologies via learnable flattening of skeletal graphs in a fully zero-shot setting.
Figure 2: Schematic of the our model architecture. The Skeleton Position Encoding uses GraphSAGE to produce a spatial mask from the rest pose S . The Encoder ”flattens” skeletal data into latent pose and trajectory tokens, while the Decoder ”unflattens” them to recover joint rotations.
Figure 3: Schematic of the training framework.
Method
JP
JR
RT
FS
GP
[cm]
[rad]
[cm]
[cm]
Bandai-Namco (Unseen)
SAME
12.40
0.34
8.58
0.08
-0.20
Ours
2.63
0.19
1.65
0.05
-0.01
Mixamo (Unseen)
SAME
10.07
0.31
4.47
0.02
-0.02
Table 1: Reconstruction Results on unseen testing datasets. We compare Joint Position (JP), Joint Rotation (JR), Root Trajectory (RT), Foot Sliding (FS), and Ground Penetration (GP) errors. Bold indicates best.
Figure 4: Skeleton invariance latent spaces of the model: 2D Principal Component Analysis projected latent space of the model for different skeletons performing semantically identical motions for different motions. The PCA space is shared across all sequences.
Figure 5: Translation invariance of our pose latent space. Left shows distributions of cosine similarities between latents of untranslated and translated motions. SAME is invariant by definition, so all the values are 1. Right shows the mean cosine similarities across translations in the x-, z- and xz-axes.
Intra
Cross
Method
GJP
Jerk
PART
GJP
Jerk
PART
Copy rotations
8.86
-
-
N/A
N/A
N/A
Villegas ’18
6.24
-
-
243
-
-
Lim ’19
5.72
-
-
N/A
N/A
N/A
Aberman
2.76
-
-
2.25
1.15
0.08
SAME †
2.91
1.87
0.18
2.47
1.94
0.17
Table 2: Animation retargeting on Mixamo. Bold : best; underline : second best. GT Jerk: 1.28 (Intra), 1.32 (Cross). † : not trained on Mixamo motions. ‡ : not trained on Mixamo skeletons.
Figure 6: Retargeting across diverse skeletons: The green skeletons are the source motions, red are ground truth retargets from the Mixamo dataset, the orange are retargets from our method, blue are from SAME and purple from Aberman et al. (2020) .
Method
Align. ↑
Qual. ↑
Aberman
4.00±0.05
3.49±0.06∗
SAME
3.42±0.06∗
3.69±0.06∗
Ours
4.06±0.05
3.86±0.05
Ground truth
4.42±0.04∗
4.39±0.04∗
Table 3: User study. Bold : best; * significant vs. ours ( p<0.05 ).
Intra
Cross
Method
GJP
Jerk
PART
GJP
Jerk
PART
Ours Full
1.45
0.72
0.03
1.28
0.72
0.05
with GAT
891.96
30.07
0.38
927.38
31.43
0.36
w/o PE
62.98
0.01
0.57
62.98
0.01
0.60
additive PE
2.06
0.70
0.05
2.01
0.72
0.06
w/o aug.
3.92
0.63
0.04
2.66
0.64
0.05
Table 4: Ablation Study on the Mixamo dataset. We evaluate the impact of removing key components on Intra- and Cross-retargeting performance. Bold indicates best; underline indicates second best.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Specification
Encoder
Linear Projection
Embedding dimension: 128
Transformer Encoder
Layers: 4, Attention Heads: 8
FeedForward: 512
Positional Encoding
GraphSAGE 2 layers
Decoder
Appendix
Table 5: Model Architecture Specifications
Hyperparameter
Value
Batch Size
128×8
Weight Decay
1×10−4
Learning Rate
1×10−3
Learning Rate Scheduler
Cosine Annealing with Restarts
Optimizer
AdamW
Appendix
Table 6: Hyperparameter Settings
Loss
Weight
λpos
102
λchild
102
λrot
5
λjerk
10−5
λtraj
10
λvel
1
Appendix
Table 7: Loss Function Weighting Coefficients
Specification
Details
Hardware
NVIDIA Tesla A100
Training Time
Approximately 48 hours
Appendix
Table 8: Training Protocol Details
Intra
Cross
Method
GJP
Jerk
PART
GJP
Jerk
PART
Ours Full
1.45
0.72
0.03
1.28
0.72
0.05
w/o Lcyc. cons.
1.47
0.82
0.03
3.25
1.23
0.06
w/o Ladv .
1.65
0.67
0.03
1.47
0.69
0.05
w/o Lchild
3.21
0.77
0.05
2.51
0.75
0.06
w/o Lee
1.62
0.70
0.03
1.46
0.71
0.04
Appendix
Table 9: Ablation Study on the Mixamo dataset. We evaluate the impact of removing key components on Intra- and Cross-retargeting performance. Bold indicates best; underline indicates second best.
Figure 7: Error accumulation when using the delta representation for trajectories:. In contrast, predicting the global position (Ours, Aberman et al. (2020) ) does not result in as much error accumulation (SAME). It shows the global root trajectory projected on the xz plane (top) and the xz root trajectory error with increasing number of frames (bottom) for three different motions. The global root trajectory prediction method used by us and Aberman et al. (2020) does not suffer from error accumulation, unlike SAME who predicts only the difference between two frames.
Figure 8: Additional skeleton invariance latent spaces of the model: 2D Principal Component Analysis projected latent space of the model for different skeletons performing semantically identical motions for different motions. The PCA space is shared across all sequences.
Figure 9: Additional examples of retargeting across diverse skeletons: The green skeletons are the source motions, red are ground truth retargets from the Mixamo dataset, the orange are retargets from our method, blue are from SAME and purple from Aberman et al. (2020) . Each row represents the same motion, while each column frames 0, 50 and 100 from that clip.
The explosion of generative 3D assets has created a massive demand for animation, yet current motion capture methods remain brittle, restricted to species-specific templates (e.g., SMPL) or requiring labor-intensive manual rigging. We introduce TopoCap, the first unified framework capable of extracting motion from monocular video and retargeting it onto characters with arbitrary, unseen skeletal topologies, i.e., from bipeds to hexapods and inanimate objects, without test-time optimization. Our key insight is that while skeletal structures are combinatorial and discrete, the underlying physics of motion occupy a continuous, low-dimensional manifold. We materialize this insight via a two-stage generative pipeline. First, we learn a Universal Motion Manifold using a Graph CVAE that compresses heterogeneous kinematic chains into a shared, fixed-length latent code. By explicitly conditioning the decoder on a structural embedding of the target rig, we disentangle motion dynamics from skeletal topology. Second, we treat video-to-animation as a conditional flow matching problem, predicting these topology-agnostic codes from visual features. To learn this generalized prior, we introduce Mobjaverse, a massive-scale dataset curated from Objaverse-XL. Comprising over 5,000 unique skeletal topologies and 2 million frames, it exceeds the structural diversity of existing datasets by two orders of magnitude. Extensive experiments demonstrate that \MethodMotion outperforms specialist models on human and quadruped benchmarks while enabling zero-shot retargeting for the long tail of 3D creatures. Dataset is publicly available at https://huggingface.co/datasets/duckduckplz/Mobjaverse.
Retargeting human motion to humanoid robots is critical for teleoperation, imitation learning and human-robot interaction. However, it remains challenging because of substantial morphological discrepancies between humans and robots, including differences in skeletal topology, limb proportions and degrees of freedom, as well as the scarcity of paired motion data. This paper presents Human2Humanoid, an unsupervised motion retargeting framework that transfers human motions to humanoid robot behaviors with high fidelity. To bridge the domain gap under unpaired data, we adopt a CycleGAN-based architecture equipped with a skeleton-aware graph convolutional network to capture topology-dependent motion features. To address cross-domain scale mismatches, we introduce a morphology-invariant end-effector consistency loss that aligns normalized end-effector trajectories to preserve motion semantics across embodiments. To improve physical plausibility and reduce contact artifacts, we impose explicit physics-aware feasibility constraints to encourage reproduction of the contact patterns in the source motion. Experimental results show that the proposed method successfully retargets human motion to the Unitree G1 humanoid robot without paired data, and outperforms existing methods in both downstream controllability and physical feasibility.
Tianchen Huang, Feiyang Yuan, Junchi Gu +5
Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China
Retargeting motion across characters with varying body shapes while preserving interaction semantics, such as self-contact and near-body proximity, remains a challenging problem. While recent geometry-aware approaches address this by maintaining spatial relationships between predefined corresponding regions, their reliance on static correspondences often struggles when the target character exhibits exaggerated body proportions. In this paper, we present a geometry-aware motion retargeting framework that preserves interaction semantics by performing proximity matching over spatially adaptive anchors. Unlike prior methods with static anchor definitions, the proposed method dynamically repositions anchors to reachable regions on the target character. This is achieved via a Transformer-based anchor refinement strategy that predicts anchor displacements and constrains the translated anchors to remain on the target character geometry through differentiable soft projection. By incorporating pose-dependent spatial structures from the source character, the adapted anchors provide structurally coherent guidance for interaction-aware retargeting. Conditioned on these anchors, a graph-based autoencoder predicts target skeletal motion that preserves the spatial configuration of the source. To encourage task-aligned optimization between anchor adaptation and motion retargeting, we adopt an alternating training scheme in which each module is optimized in turn. Through extensive evaluations, we demonstrate that our method outperforms state-of-the-art approaches in preserving interaction fidelity across diverse character geometries.