Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models
Authors: Zhiyuan Li, Wenyan Yang, Pekka Marttinen, Joni Pajarinen
Organizations: Department of Electrical Engineering and Automation, Aalto University, Finland · Department of Computer Science, Aalto University, Finland
Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, \emph{conditional-mean degeneration}: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: https://cross-skeleton-retargeting.netlify.app/.
Figures & tables
Figure 1: Action-level success does not identify the source-conditioned map. Two source clips with the same action can lead either to distinct target motions under a source-preserving map or to a shared target-action habit under source-blind recovery. Standard action-level AUC accepts both explanations; Source-Instance Fidelity tests whether the source-instance relation survives the change of body.
Figure 2: Cell-paired squared-error regression collapses toward the target-cell prototype. Panel A illustrates random target-cell pairing: several admissible target clips define a cell, but the squared-error optimum is the cell mean prototype rather than a source-conditioned motion. Panel B reports the controlled ladder with known source-conditioned transport T⋆ ; from M=2 to M=50 pairs per cell, predicted output variance stays at about 2% of the oracle within-cell variance Var(T⋆(xa)) . The shaded band is the three-seed standard deviation.
Figure 3: Source-Instance Fidelity for the fourteen evaluated methods on the 1,891 Truebones triples. Filled circles use raw scoring and open squares length-controlled scoring; intervals are 95% bootstrap intervals that resample whole source skeletons, and n is the number of triples a method supports. The shaded band marks ρ\textscsif<0.10 . Eight methods do not stand out from shuffled pairings under either scoring, and Motion2Motion-BVH stays near zero.
Figure 4: Qualitative non-identifiability panorama. A single Bird → KingCobra attack query is shown across evaluated methods. All rows receive the same source clip, target skeleton, and action request, yet the generated target motions differ visibly across methods. The figure is qualitative evidence for the ambiguity isolated by SIF: action-consistent target motion can be produced without identifying a unique source-conditioned map.
Method
Overall n/600
Held-out n/200
Overall AUC
Held-out AUC
ANCHOR
600
200
0.757
0.681
random-same-cluster
570
177
0.769 (+0.012)
0.678 (-0.003)
random-same-exact-action
221
57
0.883 (+0.126)
0.923 (+0.242)
Parenthetical Δ values are method AUC minus ANCHOR. random-same-cluster matches ANCHOR within ±0.012 on both splits. random-same-exact-action exceeds ANCHOR by +0.126 overall and +0.242 held-out, but only on the 221/600 and 57/200 coverage where exact-action positives exist; the coverage qualifier travels with the gap.
Table 1: Label-only retrieval can match or exceed ANCHOR’s action-level AUC without solving the retargeting problem. AUC is measured under the same cluster-tier Procrustes protocol used for ANCHOR. The exact-action reference has restricted coverage, so its larger gaps are only defined on the rows where the target library contains an exact-action candidate.
Figure 5: Action-level AUC cannot distinguish source-conditioned transport from label-based retrieval. ANCHOR retrieves from the target library using action information and motion descriptors, while the two random references use only label constraints. Their AUC values remain in the same action-level regime, showing that the standard score can accept target-action recovery even when source-instance transfer has not been established. The exact-action reference is coverage-limited to 221/600 overall and 57/200 held-out queries.
SIF
Action AUC
Variation
Realistic
True retargeted counterpart
+0.960
0.980
0.52
97.2%
Random same-action clip
−0.008
0.978
0.41
97.5%
Unpaired objective (Theorem 1 class)
+0.203
0.499
0.001
100.0%
Averaging objective (Proposition 1.1 class)
+0.379
0.820
0.05
53.0%
Model trained on true pairs
+0.900
0.967
0.52
94.6%
Table 2: Human-to-G1 evaluation on identical test items (90 action groups, 355 clips). SIF uses raw scoring; variation is the median ratio of output spread to source spread; realistic is the share of outputs whose nearest real G1 pose lies within the 95th percentile of held-out real motion.
Appendix figures & tables47 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Value
Skeletons
70
Total clips
616
Coarse clusters
10
Exact actions
90
Skeleton-action cells
6,300 ( 70×90 )
Occupied cells
413 (6.6%)
Appendix
Table 3: Truebones zoo regime. SIF is evaluated on source groups with at least three clips; the main evaluation pairs each group with every target skeleton that performs the same action, and the original evaluation pairs it with a single target.
Method
Family
Train skeletons
AnyTop
source-conditioned diffusion
70
ACE-T
motion-space adversarial
70
ACE-I
motion-space adversarial
60
MoReFlow-T
flow matching
70
MoReFlow-I
flow matching
60
AL-Flow
label-conditional flow, no source motion
60
Appendix
Table 4: Method roster. Train-skeleton counts distinguish transductive evaluation on all skeletons from inductive evaluation with held-out target skeletons.
Method
Relation to the formal results
AnyTop
closest to Theorem 1 ; adds rank and view-consistency terms
AL-Flow, AL-Flow-Src, AL-Flow-Src-G
closest to Proposition 1.1 ; trained with flow matching rather than squared error
MoReFlow-T, MoReFlow-I
partly within Theorem 1 ; adds pairing by matched motion descriptors and a descriptor loss
DPG-SB-v3
partly within Theorem 1 ; adds adversarial and cycle terms in the latent space
ACE-T, ACE-I
outside both classes; adversarial and source-feature losses
retrieval rules with no trained objective (Corollary 1.1 )
Appendix
Table 5: Relation of each evaluated method to the formal results.
Figure 6: Decoder non-degeneracy check. Rotation-induced changes are compared with matched-magnitude noise in the latent space; ratios above one indicate that the decoder distinguishes at least some directions on the gauge orbit.
Method
Seed
Change under rotation
Change under noise
Ratio
Verdict
ACE-T (3 seeds)
42
0.240
0.213
1.93
non-degenerate
43
0.219
0.183
1.78
non-degenerate
44
0.209
0.171
2.04
non-degenerate
ACE-I (3 seeds)
42
0.121
0.089
2.35
non-degenerate
43
0.128
0.071
1.98
non-degenerate
44
0.141
0.087
1.92
non-degenerate
Appendix
Table 6: Rotation versus noise perturbations on the 130 queries of the 37-triple intersection. The ACE ablation with the motion-space adversarial loss removed is the boundary case near the noise floor, while the other evaluated methods make the orbit visible after decoding.
Group
Seed pair
Share explained by rotation
Cosine after rotation
ACE-T
42 vs 43
0.522
0.893
ACE-T
42 vs 44
0.703
0.900
ACE-T
43 vs 44
0.585
0.905
ACE-I
42 vs 43
0.689
0.935
ACE-I
42 vs 44
0.712
0.928
ACE-I
43 vs 44
0.775
0.941
Appendix
Table 7: Orthogonal Procrustes alignment across seeds. The table reports how much latent variation can be explained by an O(d) alignment across the SIF-intersection queries.
Figure 7: Cross-seed Procrustes alignment. High explained fractions indicate that latent spaces differ partly by approximately orthogonal transformations, matching the relative-gauge mechanism in Theorem 1 .
Method
Sample shape
Spectral flatness
Effective rank / 256
ACE without adversarial loss (seed 42)
3,376 × 256
0.0001
1.29
ACE-T (seed 42)
3,376 × 256
0.0001
2.13
ACE-I (seed 42)
3,376 × 256
0.0019
3.00
AL-Flow
1,040 × 256
0.0066
6.80
MoReFlow-I
3,823 × 256
0.0593
12.84
MoReFlow-T
3,823 × 256
0.0631
14.99
Appendix
Table 8: Effective-rank diagnostic on latent distributions. Higher effective rank indicates richer latent variation, but rank alone does not imply that the source-conditioned map has been identified.
Figure 8: Effective rank across latent representations. The diagnostic separates representation richness from source-instance fidelity.
Figure 9: Decoder non-degeneracy and SIF are distinct. A method may expose latent perturbations after decoding while still lying near the source-blind floor under SIF.
Method
n
Raw SIF [95% CI]
p
Length-controlled SIF [95% CI]
p
Variation
Reading
ACE-I
1,891
+ 0.420 [ + 0.31, + 0.53]
< 0.001
+ 0.248 [ + 0.13, + 0.38]
< 0.001
0.03
above floor
MoReFlow-T
1,891
+ 0.255 [ + 0.19, + 0.33]
< 0.001
+ 0.246 [ + 0.17, + 0.32]
< 0.001
0.20
above floor
MoReFlow-I
1,891
+ 0.257 [ + 0.19, + 0.33]
< 0.001
+ 0.244 [ + 0.18, + 0.32]
< 0.001
0.16
above floor
ACE-T
1,891
+ 0.336 [ + 0.24, + 0.44]
< 0.001
+ 0.180 [ + 0.05, + 0.31]
0.009
0.02
above floor
Motion2Motion-Direct
1,891
+ 0.143 [ + 0.08, + 0.22]
< 0.001
+ 0.129 [ + 0.06, + 0.20]
< 0.001
1.52
above floor
Motion2Motion-BVH
1,872
+ 0.032 [0.000, + 0.07]
0.024
+ 0.032 [ − 0.003, + 0.07]
0.016
0.79
near floor
Appendix
Table 9: SIF on the main evaluation (1,891 triples). Raw and length-controlled SIF with 95% intervals that resample whole source skeletons; p is the one-sided shuffle-test p -value; variation is the median ratio of output spread to source spread under raw scoring; n is the number of triples a method supports. The last row is the ACE ablation of Appendix I , not an evaluated method.
Method
n
Raw SIF
p
Length-controlled SIF
p
Variation
MoReFlow-T
49
+ 0.203
0.011
+ 0.220
0.005
0.25
ACE-T
49
+ 0.385
< 0.001
+ 0.205
0.013
0.02
ACE-I
49
+ 0.484
< 0.001
+ 0.178
0.022
0.02
MoReFlow-I
49
+ 0.137
0.061
+ 0.160
0.038
0.28
Motion2Motion-Direct
49
+ 0.053
0.290
+ 0.091
0.146
1.69
ANCHOR
49
+ 0.014
0.380
+ 0.035
0.226
0.00
Appendix
Table 10: SIF on the original evaluation (49 triples, one target per source group). Columns as in Table 9 ; here no source clip is shared between triples.
Method
each triple alone
by source skeleton
by target skeleton
by action
ACE-I
0.526 [0.351, 0.681]
0.526 [0.323, 0.709]
0.526 [0.365, 0.671]
0.526 [0.242, 0.629]
ACE-T
0.453 [0.293, 0.602]
0.453 [0.284, 0.618]
0.453 [0.274, 0.605]
0.453 [-0.009, 0.619]
MoReFlow-T
0.192 [-0.003, 0.380]
0.192 [-0.023, 0.386]
0.192 [-0.031, 0.405]
0.192 [-0.039, 0.357]
random-same-exact-action
0.185 [0.046, 0.326]
0.185 [0.069, 0.304]
0.185 [0.027, 0.354]
0.185 [0.000, 0.272]
Motion2Motion-Direct
0.131 [-0.061, 0.317]
0.131 [-0.098, 0.385]
0.131 [-0.075, 0.307]
0.131 [-0.720, 0.265]
random-same-cluster
0.104 [-0.095, 0.298]
0.104 [-0.057, 0.279]
0.104 [-0.115, 0.315]
0.104 [-0.272, 0.189]
Appendix
Table 11: Clustered bootstrap intervals for SIF. Clustering by source skeleton, target skeleton, or action tests whether the positive conclusions are driven by a narrow subset of triples.
Figure 10: Clustered confidence intervals for SIF. The figure visualizes the dependence-aware intervals reported in Table 11 .
Evaluation
Scoring
ACE-T minus MoReFlow-T
ACE-I minus MoReFlow-I
1,891 triples
raw
+ 0.081 [ − 0.02, + 0.19], p =0.139
+ 0.163 [ + 0.05, + 0.28], p =0.008
1,891 triples
length-controlled
− 0.066 [ − 0.19, + 0.06], p =0.325
+ 0.004 [ − 0.10, + 0.11], p =0.935
49 triples
raw
+ 0.182 [ − 0.06, + 0.44], p =0.156
+ 0.347 [ + 0.11, + 0.60], p =0.009
49 triples
length-controlled
− 0.015 [ − 0.21, + 0.19], p =0.888
+ 0.018 [ − 0.21, + 0.24], p =0.880
Appendix
Table 12: Paired SIF differences on identical triples, with 95% intervals that resample whole source skeletons and two-sided sign-flip p -values that flip all triples of a source skeleton together.
One triple
Average over k triples
Method
(three clips)
k=25
k=100
k=500
k=1,000
AnyTop
0.60
0.122
0.060
0.021
0.012
ACE-T
0.59
0.124
0.060
0.022
0.013
ACE-I
0.56
0.110
0.054
0.019
0.012
MoReFlow-T
0.57
0.123
0.059
0.023
0.013
random-same-exact-action
0.31
0.089
0.046
0.017
0.010
Appendix
Table 13: Noise of a single triple and stability of the average. The first column is the standard deviation of SIF across three-clip subsets of five-clip triples; the other columns are standard deviations of the average over k randomly drawn triples.
Independent noise, three seeds
Method
seed 1
seed 2
seed 3
Mean ± sd
Shared noise
AnyTop
− 0.033
+ 0.040
− 0.098
− 0.031 ± 0.069
+ 0.243
AL-Flow
− 0.066
− 0.043
+ 0.066
− 0.015 ± 0.070
+ 0.002
AL-Flow-Src
+ 0.070
+ 0.049
− 0.140
− 0.007 ± 0.115
+ 0.199
AL-Flow-Src-G
− 0.028
+ 0.048
+ 0.075
+ 0.031 ± 0.053
+ 0.239
DPG-SB-v3
− 0.058
+ 0.138
+ 0.052
+ 0.044 ± 0.098
+ 0.064
Appendix
Table 14: Generation noise control on the original 49 triples. Independent noise uses three seeds; shared noise reuses one noise sample within each triple.
Method
Min pairs
Median pairs
Max pairs
n triples
ACE-I
3
3
5
49
ACE-T
3
3
5
49
AL-Flow
3
3
5
45
AL-Flow-Src
3
3
5
38
AL-Flow-Src-G
3
3
5
49
ANCHOR
3
3
5
49
Appendix
Table 15: Pair counts per SIF triple. The minimum and median values are fixed by the eligibility rule; the few larger cells determine how much within-cell geometry is available.
Figure 11: Pair-count distribution for SIF-eligible triples.
Figure 12: SIF versus output variation (the diversity ratio R ). High output variation is not sufficient for source-instance preservation, because the geometry must align with the source-side geometry after the target skeleton and action are fixed.
Method
SIF ρ
L-SIF ρ
ACE trained with zsrc=0
0.423
0.037
ACE-I (3 seeds avg)
0.410
0.049
ACE-T (3 seeds avg)
0.408
0.088
ACE without adversarial loss (3 seeds avg)
0.214
0.119
MoReFlow-T
0.192
-0.200
MoReFlow-I
0.076
-0.205
Appendix
Table 16: SIF and L-SIF for each method. L-SIF is a latent-space diagnostic and is therefore reported only where the representation is comparable across queries.
Figure 13: SIF and L-SIF decoupling. Source-instance fidelity in decoded motion need not coincide with a positive latent-space correlation.
Figure 14: Synthetic 2×2 SIF calibration. The oracle uses the known source-conditioned transport T⋆(xa) ; the random-clip and Gaussian references deliberately ignore the source instance.
Output
SIF (95% CI)
p
Variation
Oracle (output =T∗(xa) )
+ 0.993 [ + 0.992, + 0.995]
< 0.001
0.94
Random clip from cell
+ 0.002 [ − 0.027, + 0.029]
0.456
0.96
Random Gaussian noise
+ 0.021 [ − 0.008, + 0.054]
0.164
1.3×105
Synthetic 2 × 2 paired-dense setting, 168 groups of three to six source clips. Another 168 groups are left out because their source clips are identical once rotation and scale are removed, so SIF is undefined there. The oracle saturates SIF; both source-blind references sit at zero. Noise has no motion structure, so its variation is very large.
Appendix
Table 17: Synthetic SIF calibration. The oracle lies near perfect source-instance preservation, while source-blind references remain near zero.
M pairs / cell
σpred2/σoracle2
MSE /σoracle2
nseeds
2
0.020 ± 0.006
1.45 ± 0.08
3
4
0.014 ± 0.002
1.29 ± 0.03
3
8
0.020 ± 0.006
1.34 ± 0.08
3
16
0.015 ± 0.002
1.31 ± 0.12
3
32
0.017 ± 0.001
1.22 ± 0.05
3
50
0.016 ± 0.003
1.26 ± 0.09
3
Appendix
Table 18: Conditional-mean degeneration in the controlled 2×2 setting. The predicted variance remains a small fraction of oracle variance even when the number of random target-cell pairs increases.
Domain
Cluster classifier accuracy
n correct / total
In-domain (train skeletons)
99.8%
543 / 544
Held-out (reserved cross-skeleton)
61.5%
59 / 96
Replacing the ground-truth cluster label with the predicted cluster pulls Procrustes performance below random on exact-tier and on a hand-curated query subset. The in-domain row is the accuracy on the clips the classifier was trained on.
Appendix
Table 19: ANCHOR action-cluster predictor. The held-out row shows that the cluster prediction itself is imperfect, which is why ANCHOR is interpreted as a comparator rather than an oracle.
Figure 15: Per-cluster ANCHOR matches. The comparator’s behavior varies across action clusters, reinforcing that it characterizes a floor rather than a universal retargeting rule.
Subset
n triples (cluster-eligible)
Cluster-tier AUC (95% CI)
All
30,029
0.824 [0.820, 0.829]
In-distribution (same train pair)
21,838
0.845 [0.841, 0.849]
Mixed (one held-out skeleton)
7,622
0.774 [0.762, 0.787]
Held-out (both cross-skeleton)
569
0.711 [0.669, 0.758]
ANCHOR on the 30,497-pair full enumeration. Cluster-eligibility filter removes triples whose target skeleton has no in-cluster positives. The mixed subset is the held-out case.
Appendix
Table 20: Full ANCHOR enumeration across cluster-eligible triples. The held-out subsets show that strong action-level ranking can arise from target-library label structure without identifying a source-conditioned map.
Figure 16: Enumeration breakdown for ANCHOR. The large enumeration separates in-distribution, mixed, and held-out skeleton regimes.
Variant
Conditioning channels
Cluster AUC
AnyTop
self-supervised source-conditioned diffusion
0.465
Precursor action-conditioned generator
action-conditioned generation
0.485
Precursor latent-bridge generator
latent bridge objective
0.483
DPG-SB-v3
latent bridge objective; no decoded-motion penalties
0.447
AL-Flow
cluster + exact-action labels
0.538
AL-Flow-Src
AL-Flow + source motion + skeleton identity
0.542
Appendix
Table 21: Action-level AUC for generated variants and supervision-matched comparators. These values are not interpreted as source-preserving transfer; they show how much action-level evidence can be explained without SIF.
Cluster-tier held-out AUC (average of folds 42 and 43)
Method
Procrustes
Z-DTW
Q-comp
Action oracle
0.873
0.846
0.894
Self-positive reference
0.860
0.847
0.909
ANCHOR
0.681
0.704
0.728
Cluster-Classifier Retrieval
0.646
0.669
0.693
Q-Retrieval (Q-feature only)
0.453
0.523
0.597
Appendix
Table 22: Master action-level AUC comparison on held-out queries. The table reports cluster-tier AUC under three distances: Procrustes trajectory distance; Z-DTW, dynamic time warping on z-normalized body-part trajectories; and Q-comp, a combination of centre-of-mass path, foot-contact timing, cadence, and limb-usage differences. Library-based references are listed separately from generative methods.
Figure 17: Action-level AUC band. Many methods occupy a narrow action-ranking band, which motivates SIF as the missing source-instance axis.
ACE-T
Model trained with zsrc=0
Source at generation
Output length
SIF
p
Variation
SIF
p
Variation
real
copied
+ 0.385
< 0.001
0.020
+ 0.419
< 0.001
0.023
real
fixed
+ 0.294
< 0.001
0.003
+ 0.213
0.006
0.004
removed
copied
+ 0.300
< 0.001
0.009
+ 0.299
< 0.001
0.007
removed
fixed
0.000
1.000
0.000
0.000
1.000
0.000
shuffled
copied
− 0.192
0.987
0.020
− 0.165
0.970
0.023
Appendix
Table 23: Removing or shuffling ACE’s source input at generation (49 original triples). Removed means zsrc set to zero at generation; shuffled means each output receives the source of another clip in the same triple; fixed means every output is generated with 64 frames. p is the shuffle-test p -value; variation as in Table 9 .
With adversarial loss
Without
Scoring
Triples
SIF
Variation
SIF
Variation
Difference [95% CI]
p
raw, length copied
49
+ 0.407
0.017–0.022
+ 0.228
0.271–0.906
+ 0.179 [ + 0.04, + 0.33]
0.022
length fixed at generation
49
+ 0.292
0.002–0.004
+ 0.266
0.073–0.107
+ 0.026 [ − 0.09, + 0.15]
0.677
length fixed at generation
1,891
+ 0.329
0.002–0.003
+ 0.329
0.078–0.111
0.000 [ − 0.05, + 0.06]
0.999
Appendix
Table 24: ACE-I with and without the adversarial loss, trained identically otherwise (three seeds each). SIF is averaged over seeds; variation gives the range of the per-seed medians; the difference is paired over triples, with a 95% interval that resamples whole source skeletons.
Actual outputs
Identical copies,
Method
Raw
Stretched to 64 frames
stretched to 64 frames
ACE-T
0.020
0.231
0.149
ACE-I
0.020
0.135
0.142
MoReFlow-T
0.247
0.217
0.093
MoReFlow-I
0.276
0.279
0.078
Appendix
Table 25: Output variation after stretching every output to 64 frames (49 original triples). Identical copies are one output cut to the lengths of its siblings, so any variation they show comes from stretching alone.
Method
Queries
Closer to the target pool
Mean distance to source
Mean distance to target pool
Ratio
ACE-I
130
103 (79.2%)
3.449
2.215
1.56
ACE-T
130
99 (76.2%)
3.697
2.650
1.39
Distance computed in a four-dimensional kinematic feature space (centre-of-mass displacement, centre-of-mass variance, mean velocity, foot-contact density). 76–79% of ACE outputs are closer to a target-pool training clip than to the corresponding source clip, weakening source-memorisation as an explanation of ACE’s SIF lift.
Appendix
Table 26: ACE nearest-neighbor leakage check. Distances are computed in a kinematic feature space; the majority of ACE outputs are closer to the target pool than to the source clip.
Figure 18: Human-to-robot renders. (A) One LAFAN1 dance routine (10 s) retargeted by GMR to six humanoid robots; five key poses per strip, earliest faded. Each strip is framed to its robot, so compare poses, not sizes; the base heights give the true scale. (B) Two performers dancing the same routine, and what each model produces for each of them on the Unitree G1. The true retargets and the model trained on true pairs differ between the performers. The averaging and unpaired objectives produce nearly the same motion for both, and that motion barely changes over time. Model outputs are predicted body positions, drawn as the nearest G1 pose; the fit error is listed under each row. The 5 to 6 cm errors of the two collapsed rows mean their raw predictions are less robot-like than drawn. The two performers are not aligned in time, so part of the difference in the top rows is timing rather than style.
Raw SIF
Length-controlled SIF
Action AUC
Variation
Realistic
True retargeted counterpart
+ 0.960
+ 0.959
0.980
0.518
97.2%
Random same-action clip
− 0.008
− 0.068
0.978
0.414
97.5%
Unpaired objective
+ 0.203
+ 0.154
0.499
0.001
100.0%
Averaging objective
+ 0.379
+ 0.343
0.820
0.048
53.0%
Model trained on true pairs
+ 0.900
+ 0.698
0.967
0.521
94.6%
Appendix
Table 27: Human-to-G1 study, all four measures on identical test items (90 groups, 355 clips), with SIF under both scorings. Variation is the median ratio of output spread to source spread; realistic as in Table 2 .
Prediction
Threshold
Result
The true retarget is recognized
SIF ≥0.80 , interval above 0
19 of 24; every interval above 0; misses 0.768–0.799
A random same-group clip sits at the floor
90% interval within ±0.20
12 of 24; 12 of 12 in the denser setting
The averaging objective loses variation
variation Q≤0.35 , upper bound <0.50
24 of 24
The unpaired objective loses variation
variation Q≤0.35 , upper bound <0.50
22 of 24; missed with variation Q=0.41
The true-pair model recovers both
SIF ≥0.60 and variation Q≥0.70
24 of 24; SIF 0.77–0.99
Appendix
Table 28: Six-robot predictions, fixed before any trained model was scored, and their outcomes over 24 evaluations (six robots, two settings, two scorings). Q is a model’s output variation divided by that of the true counterpart, matched group by group; intervals are group bootstraps.
True retarget
Random clip
Unpaired
Averaging
True-pair model
Robot
Setting
Scoring
SIF
SIF
Variation ( Q )
Variation ( Q )
SIF
Variation ( Q )
Unitree G1
three-clip
raw
+ 0.988
+ 0.066
0.059
0.286
+ 0.987
1.00
Unitree G1
three-clip
length-controlled
+ 0.987
+ 0.066
0.059
0.286
+ 0.987
1.00
Unitree G1
denser
raw
+ 0.966
− 0.061
0.012
0.035
+ 0.965
0.99
Unitree G1
denser
length-controlled
+ 0.966
− 0.058
0.012
0.035
+ 0.965
0.99
Booster T1
three-clip
raw
+ 0.799
+ 0.024
0.055
0.347
+ 0.845
0.99
Appendix
Table 29: Six-robot results for every robot, setting, and scoring. SIF columns give the correlation; Q columns give output variation relative to the true counterpart.
Raw SIF
Action AUC
Variation
Realistic
Planned run
+ 0.209
0.500
diverged
0.0%
Re-run, seed 1
+ 0.696
0.850
0.561
14.4%
Re-run, seed 2
+ 0.528
0.551
0.093
78.0%
Re-run, seed 3
+ 0.574
0.513
0.101
0.0%
Appendix
Table 30: An adversarial objective in the style of ACE on the human-to-G1 data, without pairs. The planned run diverged; the re-runs use ACE’s optimizer settings and are post hoc.
Method
Raw SIF
Action AUC
Variation
Realistic
ACE-I
+ 0.484
0.539
0.02
91.8%
ACE-T
+ 0.385
0.538
0.02
92.4%
MoReFlow-T
+ 0.203
0.596
0.25
99.4%
MoReFlow-I
+ 0.137
0.578
0.28
99.4%
random-same-exact-action
+ 0.096
0.799
0.00
100.0%
random-same-cluster
+ 0.073
0.813
0.25
100.0%
Appendix
Table 31: Truebones methods on the 49 original triples: SIF (raw), action AUC, output variation, and realism on identical test items. Realism is undefined for Motion2Motion-BVH, whose outputs use a different joint layout.
Raw SIF
Action AUC
Variation ( Q )
Realistic
Three clips per group
True retargeted counterpart
+0.87 (+0.77–+0.99)
0.61 (0.60–0.63)
1.00 (1.00–1.00)
99 (99–99)%
Random same-group clip
+0.08 (+0.02–+0.18)
0.62 (0.61–0.63)
0.73 (0.69–0.77)
99 (99–100)%
Unpaired objective
+0.24 (+0.11–+0.36)
0.50 (0.49–0.52)
0.15 (0.02–0.41)
98 (89–100)%
Averaging objective
+0.11 (+0.07–+0.20)
0.60 (0.58–0.63)
0.29 (0.26–0.35)
46 (33–83)%
Model trained on true pairs
+0.89 (+0.85–+0.99)
0.61 (0.60–0.63)
1.00 (0.99–1.00)
98 (97–99)%
Appendix
Table 32: Six-robot study, four measures on identical test items: mean over the six robots, with the range in parentheses. Variation is Q , a model’s output variation divided by that of the true counterpart. Action AUC and realism were computed after the pre-registration.
Figure 19: A source-clip substitution view on one Bird → KingCobra attack triple. The first three columns, labeled source 1 to source 3, are different Bird attack clips, while the target skeleton and action are fixed; the three columns on the right, labeled output 1 to output 3, show the corresponding KingCobra motions generated by each row. Each row is labeled by its source-instance failure mode: latent collapse for AL-Flow, where source identity is lost before decoding; decoder collapse for AnyTop, where latent variation is not expressed in the output; partial source-conditionality for MoReFlow-T, the clearest output-level source dependence among the displayed rows; and collapsed variation for ACE-I, whose outputs keep a positive correlation with their sources but remain nearly identical.
Figure 20: Full qualitative grid A. Three source clips are compared against target outputs from the evaluated methods on a fixed source-target-action triple.
Figure 21: Full qualitative grid B. The same layout is used to reveal whether output variation follows the source clips after target skeleton and action are fixed.
Figure 22: Full qualitative grid C. The grid complements Figure 19 by showing all evaluated methods on an additional held-out triple.
Cross-structural motion retargeting aims to transfer motion between different skeletal topologies. Despite recent progress, existing state-of-the-art models struggle with reliability in zero-shot settings, i.e. skeletons with different topologies which were unseen during training, and recent Transformer-based attempts have failed to outperform specialized geometric methods. We bridge this gap with a Transformer Autoencoder that learns a topology- and translation-invariant latent space. Our core contribution is a learnable flattening of skeletal graphs that captures both local dependencies and global structure. Unlike the standard transformer architecture, which adds positional information to token content, we integrate graph-based positional encodings multiplicatively, a design choice that follows directly from our flattening formulation. The resulting model handles diverse skeletal topologies within a single unified architecture and trains in a fully unsupervised manner, requiring no paired retargeting data. Ablation studies show, that the graph encodings, multiplicative formulation, and Transformer backbone is critical for the performance. In zero-shot evaluations, our method reduces global joint position error by 43−47% over current benchmarks. A user study (n=37), including expert animators, further ranks our approach highest in motion alignment and physical plausibility (p<0.05). These results demonstrate that our model design is key to making transformer architectures effective for motion retargeting, outperforming existing approaches.
Kia-Jüng Yang, Fabian H. Sinz, Paweł A. Pierzchlewicz
Institute of Computer Science, University of Göttingen · Campus Institute Data Science, University Göttingen · Pantomim P.S.A
Retargeting human motion to humanoid robots is critical for teleoperation, imitation learning and human-robot interaction. However, it remains challenging because of substantial morphological discrepancies between humans and robots, including differences in skeletal topology, limb proportions and degrees of freedom, as well as the scarcity of paired motion data. This paper presents Human2Humanoid, an unsupervised motion retargeting framework that transfers human motions to humanoid robot behaviors with high fidelity. To bridge the domain gap under unpaired data, we adopt a CycleGAN-based architecture equipped with a skeleton-aware graph convolutional network to capture topology-dependent motion features. To address cross-domain scale mismatches, we introduce a morphology-invariant end-effector consistency loss that aligns normalized end-effector trajectories to preserve motion semantics across embodiments. To improve physical plausibility and reduce contact artifacts, we impose explicit physics-aware feasibility constraints to encourage reproduction of the contact patterns in the source motion. Experimental results show that the proposed method successfully retargets human motion to the Unitree G1 humanoid robot without paired data, and outperforms existing methods in both downstream controllability and physical feasibility.
Tianchen Huang, Feiyang Yuan, Junchi Gu +5
Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China
Text-driven motion editing and intra-structural retargeting, where skeletons share topology but may differ in bone lengths and rest pose, are traditionally handled by fragmented pipelines with incompatible inputs and representations: editing relies on specialized generative steering, while retargeting is deferred to geometric post-processing. We present a unified conditional-flow framework that casts generation, semantic editing, and intra-structural retargeting as condition-modulated transport within one text- and skeleton-conditioned rectified-flow model. Under this formulation, editing changes the semantic condition while preserving skeletal structure, whereas retargeting changes the skeletal condition while preserving motion semantics. This makes FlowEdit-style transport a unified inference rule for motion manipulation rather than a task-specific editor. To instantiate this for articulated 3D motion, we develop a text- and skeleton-conditioned rectified-flow transformer. The model uses per-joint tokenization and explicit joint self-attention to capture spatial kinematic dependencies. We further inject text conditions at both joint and frame levels, while residual multi-condition classifier-free guidance balances text adherence and skeletal conformity. Experiments on SnapMoGen and a multi-character Mixamo subset show that one trained model supports text-to-motion generation, zero-shot editing, and zero-shot intra-structural retargeting without task-specific fine-tuning. This unified framework replaces separate pipelines with a single conditional motion transport model while keeping the same-topology retargeting scope explicit.