Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models
Organizations: Department of Electrical Engineering and Automation, Aalto University, Finland · Department of Computer Science, Aalto University, Finland
Abstract
Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, \emph{conditional-mean degeneration}: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: https://cross-skeleton-retargeting.netlify.app/.
Figures & tables
| Method | Overall | Held-out | Overall AUC | Held-out AUC |
| ANCHOR | 600 | 200 | 0.757 | 0.681 |
| random-same-cluster | 570 | 177 | 0.769 (+0.012) | 0.678 (-0.003) |
| random-same-exact-action | 221 | 57 | 0.883 (+0.126) | 0.923 (+0.242) |
| Parenthetical values are method AUC minus ANCHOR. random-same-cluster matches ANCHOR within on both splits. random-same-exact-action exceeds ANCHOR by overall and held-out, but only on the and coverage where exact-action positives exist; the coverage qualifier travels with the gap. | ||||
| SIF | Action AUC | Variation | Realistic | |
|---|---|---|---|---|
| True retargeted counterpart | 0.980 | 0.52 | 97.2% | |
| Random same-action clip | 0.978 | 0.41 | 97.5% | |
| Unpaired objective (Theorem 1 class) | 0.499 | 0.001 | 100.0% | |
| Averaging objective (Proposition 1.1 class) | 0.820 | 0.05 | 53.0% | |
| Model trained on true pairs | 0.967 | 0.52 | 94.6% |
Appendix figures & tables47 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Value |
|---|---|
| Skeletons | 70 |
| Total clips | 616 |
| Coarse clusters | 10 |
| Exact actions | 90 |
| Skeleton-action cells | 6,300 ( ) |
| Occupied cells | 413 (6.6%) |
| Method | Family | Train skeletons |
|---|---|---|
| AnyTop | source-conditioned diffusion | 70 |
| ACE-T | motion-space adversarial | 70 |
| ACE-I | motion-space adversarial | 60 |
| MoReFlow-T | flow matching | 70 |
| MoReFlow-I | flow matching | 60 |
| AL-Flow | label-conditional flow, no source motion | 60 |
| Method | Relation to the formal results |
|---|---|
| AnyTop | closest to Theorem 1 ; adds rank and view-consistency terms |
| AL-Flow, AL-Flow-Src, AL-Flow-Src-G | closest to Proposition 1.1 ; trained with flow matching rather than squared error |
| MoReFlow-T, MoReFlow-I | partly within Theorem 1 ; adds pairing by matched motion descriptors and a descriptor loss |
| DPG-SB-v3 | partly within Theorem 1 ; adds adversarial and cycle terms in the latent space |
| ACE-T, ACE-I | outside both classes; adversarial and source-feature losses |
| ANCHOR, random-same-cluster, random-same-exact-action | retrieval rules with no trained objective (Corollary 1.1 ) |
| Method | Seed | Change under rotation | Change under noise | Ratio | Verdict |
| ACE-T (3 seeds) | 42 | 0.240 | 0.213 | 1.93 | non-degenerate |
| 43 | 0.219 | 0.183 | 1.78 | non-degenerate | |
| 44 | 0.209 | 0.171 | 2.04 | non-degenerate | |
| ACE-I (3 seeds) | 42 | 0.121 | 0.089 | 2.35 | non-degenerate |
| 43 | 0.128 | 0.071 | 1.98 | non-degenerate | |
| 44 | 0.141 | 0.087 | 1.92 | non-degenerate |
| Group | Seed pair | Share explained by rotation | Cosine after rotation |
|---|---|---|---|
| ACE-T | 42 vs 43 | 0.522 | 0.893 |
| ACE-T | 42 vs 44 | 0.703 | 0.900 |
| ACE-T | 43 vs 44 | 0.585 | 0.905 |
| ACE-I | 42 vs 43 | 0.689 | 0.935 |
| ACE-I | 42 vs 44 | 0.712 | 0.928 |
| ACE-I | 43 vs 44 | 0.775 | 0.941 |
| Method | Sample shape | Spectral flatness | Effective rank 256 |
|---|---|---|---|
| ACE without adversarial loss (seed 42) | 3,376 256 | 0.0001 | 1.29 |
| ACE-T (seed 42) | 3,376 256 | 0.0001 | 2.13 |
| ACE-I (seed 42) | 3,376 256 | 0.0019 | 3.00 |
| AL-Flow | 1,040 256 | 0.0066 | 6.80 |
| MoReFlow-I | 3,823 256 | 0.0593 | 12.84 |
| MoReFlow-T | 3,823 256 | 0.0631 | 14.99 |
| Method | Raw SIF [95% CI] | Length-controlled SIF [95% CI] | Variation | Reading | |||
|---|---|---|---|---|---|---|---|
| ACE-I | 1,891 | 0.420 [ 0.31, 0.53] | 0.001 | 0.248 [ 0.13, 0.38] | 0.001 | 0.03 | above floor |
| MoReFlow-T | 1,891 | 0.255 [ 0.19, 0.33] | 0.001 | 0.246 [ 0.17, 0.32] | 0.001 | 0.20 | above floor |
| MoReFlow-I | 1,891 | 0.257 [ 0.19, 0.33] | 0.001 | 0.244 [ 0.18, 0.32] | 0.001 | 0.16 | above floor |
| ACE-T | 1,891 | 0.336 [ 0.24, 0.44] | 0.001 | 0.180 [ 0.05, 0.31] | 0.009 | 0.02 | above floor |
| Motion2Motion-Direct | 1,891 | 0.143 [ 0.08, 0.22] | 0.001 | 0.129 [ 0.06, 0.20] | 0.001 | 1.52 | above floor |
| Motion2Motion-BVH | 1,872 | 0.032 [0.000, 0.07] | 0.024 | 0.032 [ 0.003, 0.07] | 0.016 | 0.79 | near floor |
| Method | Raw SIF | Length-controlled SIF | Variation | |||
|---|---|---|---|---|---|---|
| MoReFlow-T | 49 | 0.203 | 0.011 | 0.220 | 0.005 | 0.25 |
| ACE-T | 49 | 0.385 | 0.001 | 0.205 | 0.013 | 0.02 |
| ACE-I | 49 | 0.484 | 0.001 | 0.178 | 0.022 | 0.02 |
| MoReFlow-I | 49 | 0.137 | 0.061 | 0.160 | 0.038 | 0.28 |
| Motion2Motion-Direct | 49 | 0.053 | 0.290 | 0.091 | 0.146 | 1.69 |
| ANCHOR | 49 | 0.014 | 0.380 | 0.035 | 0.226 | 0.00 |
| Method | each triple alone | by source skeleton | by target skeleton | by action |
|---|---|---|---|---|
| ACE-I | 0.526 [0.351, 0.681] | 0.526 [0.323, 0.709] | 0.526 [0.365, 0.671] | 0.526 [0.242, 0.629] |
| ACE-T | 0.453 [0.293, 0.602] | 0.453 [0.284, 0.618] | 0.453 [0.274, 0.605] | 0.453 [-0.009, 0.619] |
| MoReFlow-T | 0.192 [-0.003, 0.380] | 0.192 [-0.023, 0.386] | 0.192 [-0.031, 0.405] | 0.192 [-0.039, 0.357] |
| random-same-exact-action | 0.185 [0.046, 0.326] | 0.185 [0.069, 0.304] | 0.185 [0.027, 0.354] | 0.185 [0.000, 0.272] |
| Motion2Motion-Direct | 0.131 [-0.061, 0.317] | 0.131 [-0.098, 0.385] | 0.131 [-0.075, 0.307] | 0.131 [-0.720, 0.265] |
| random-same-cluster | 0.104 [-0.095, 0.298] | 0.104 [-0.057, 0.279] | 0.104 [-0.115, 0.315] | 0.104 [-0.272, 0.189] |
| Evaluation | Scoring | ACE-T minus MoReFlow-T | ACE-I minus MoReFlow-I |
|---|---|---|---|
| 1,891 triples | raw | 0.081 [ 0.02, 0.19], =0.139 | 0.163 [ 0.05, 0.28], =0.008 |
| 1,891 triples | length-controlled | 0.066 [ 0.19, 0.06], =0.325 | 0.004 [ 0.10, 0.11], =0.935 |
| 49 triples | raw | 0.182 [ 0.06, 0.44], =0.156 | 0.347 [ 0.11, 0.60], =0.009 |
| 49 triples | length-controlled | 0.015 [ 0.21, 0.19], =0.888 | 0.018 [ 0.21, 0.24], =0.880 |
| One triple | Average over triples | ||||
|---|---|---|---|---|---|
| Method | (three clips) | ||||
| AnyTop | 0.60 | 0.122 | 0.060 | 0.021 | 0.012 |
| ACE-T | 0.59 | 0.124 | 0.060 | 0.022 | 0.013 |
| ACE-I | 0.56 | 0.110 | 0.054 | 0.019 | 0.012 |
| MoReFlow-T | 0.57 | 0.123 | 0.059 | 0.023 | 0.013 |
| random-same-exact-action | 0.31 | 0.089 | 0.046 | 0.017 | 0.010 |
| Independent noise, three seeds | |||||
|---|---|---|---|---|---|
| Method | seed 1 | seed 2 | seed 3 | Mean sd | Shared noise |
| AnyTop | 0.033 | 0.040 | 0.098 | 0.031 0.069 | 0.243 |
| AL-Flow | 0.066 | 0.043 | 0.066 | 0.015 0.070 | 0.002 |
| AL-Flow-Src | 0.070 | 0.049 | 0.140 | 0.007 0.115 | 0.199 |
| AL-Flow-Src-G | 0.028 | 0.048 | 0.075 | 0.031 0.053 | 0.239 |
| DPG-SB-v3 | 0.058 | 0.138 | 0.052 | 0.044 0.098 | 0.064 |
| Method | Min pairs | Median pairs | Max pairs | triples |
|---|---|---|---|---|
| ACE-I | 3 | 3 | 5 | 49 |
| ACE-T | 3 | 3 | 5 | 49 |
| AL-Flow | 3 | 3 | 5 | 45 |
| AL-Flow-Src | 3 | 3 | 5 | 38 |
| AL-Flow-Src-G | 3 | 3 | 5 | 49 |
| ANCHOR | 3 | 3 | 5 | 49 |
| Method | SIF | L-SIF |
|---|---|---|
| ACE trained with | 0.423 | 0.037 |
| ACE-I (3 seeds avg) | 0.410 | 0.049 |
| ACE-T (3 seeds avg) | 0.408 | 0.088 |
| ACE without adversarial loss (3 seeds avg) | 0.214 | 0.119 |
| MoReFlow-T | 0.192 | -0.200 |
| MoReFlow-I | 0.076 | -0.205 |
| Output | SIF (95% CI) | Variation | |
|---|---|---|---|
| Oracle (output ) | 0.993 [ 0.992, 0.995] | 0.001 | 0.94 |
| Random clip from cell | 0.002 [ 0.027, 0.029] | 0.456 | 0.96 |
| Random Gaussian noise | 0.021 [ 0.008, 0.054] | 0.164 | |
| Synthetic 2 2 paired-dense setting, 168 groups of three to six source clips. Another 168 groups are left out because their source clips are identical once rotation and scale are removed, so SIF is undefined there. The oracle saturates SIF; both source-blind references sit at zero. Noise has no motion structure, so its variation is very large. | |||
| pairs / cell | MSE | ||
| 2 | 0.020 0.006 | 1.45 0.08 | 3 |
| 4 | 0.014 0.002 | 1.29 0.03 | 3 |
| 8 | 0.020 0.006 | 1.34 0.08 | 3 |
| 16 | 0.015 0.002 | 1.31 0.12 | 3 |
| 32 | 0.017 0.001 | 1.22 0.05 | 3 |
| 50 | 0.016 0.003 | 1.26 0.09 | 3 |
| Domain | Cluster classifier accuracy | correct total |
| In-domain (train skeletons) | 99.8% | 543 544 |
| Held-out (reserved cross-skeleton) | 61.5% | 59 96 |
| Replacing the ground-truth cluster label with the predicted cluster pulls Procrustes performance below random on exact-tier and on a hand-curated query subset. The in-domain row is the accuracy on the clips the classifier was trained on. | ||
| Subset | triples (cluster-eligible) | Cluster-tier AUC (95% CI) |
|---|---|---|
| All | 30,029 | 0.824 [0.820, 0.829] |
| In-distribution (same train pair) | 21,838 | 0.845 [0.841, 0.849] |
| Mixed (one held-out skeleton) | 7,622 | 0.774 [0.762, 0.787] |
| Held-out (both cross-skeleton) | 569 | 0.711 [0.669, 0.758] |
| ANCHOR on the 30,497-pair full enumeration. Cluster-eligibility filter removes triples whose target skeleton has no in-cluster positives. The mixed subset is the held-out case. | ||
| Variant | Conditioning channels | Cluster AUC |
| AnyTop | self-supervised source-conditioned diffusion | 0.465 |
| Precursor action-conditioned generator | action-conditioned generation | 0.485 |
| Precursor latent-bridge generator | latent bridge objective | 0.483 |
| DPG-SB-v3 | latent bridge objective; no decoded-motion penalties | 0.447 |
| AL-Flow | cluster + exact-action labels | 0.538 |
| AL-Flow-Src | AL-Flow + source motion + skeleton identity | 0.542 |
| Cluster-tier held-out AUC (average of folds 42 and 43) | |||
|---|---|---|---|
| Method | Procrustes | Z-DTW | Q-comp |
| Action oracle | 0.873 | 0.846 | 0.894 |
| Self-positive reference | 0.860 | 0.847 | 0.909 |
| ANCHOR | 0.681 | 0.704 | 0.728 |
| Cluster-Classifier Retrieval | 0.646 | 0.669 | 0.693 |
| Q-Retrieval (Q-feature only) | 0.453 | 0.523 | 0.597 |
| ACE-T | Model trained with | ||||||
|---|---|---|---|---|---|---|---|
| Source at generation | Output length | SIF | Variation | SIF | Variation | ||
| real | copied | 0.385 | 0.001 | 0.020 | 0.419 | 0.001 | 0.023 |
| real | fixed | 0.294 | 0.001 | 0.003 | 0.213 | 0.006 | 0.004 |
| removed | copied | 0.300 | 0.001 | 0.009 | 0.299 | 0.001 | 0.007 |
| removed | fixed | 0.000 | 1.000 | 0.000 | 0.000 | 1.000 | 0.000 |
| shuffled | copied | 0.192 | 0.987 | 0.020 | 0.165 | 0.970 | 0.023 |
| With adversarial loss | Without | ||||||
|---|---|---|---|---|---|---|---|
| Scoring | Triples | SIF | Variation | SIF | Variation | Difference [95% CI] | |
| raw, length copied | 49 | 0.407 | 0.017–0.022 | 0.228 | 0.271–0.906 | 0.179 [ 0.04, 0.33] | 0.022 |
| length fixed at generation | 49 | 0.292 | 0.002–0.004 | 0.266 | 0.073–0.107 | 0.026 [ 0.09, 0.15] | 0.677 |
| length fixed at generation | 1,891 | 0.329 | 0.002–0.003 | 0.329 | 0.078–0.111 | 0.000 [ 0.05, 0.06] | 0.999 |
| Actual outputs | Identical copies, | ||
|---|---|---|---|
| Method | Raw | Stretched to 64 frames | stretched to 64 frames |
| ACE-T | 0.020 | 0.231 | 0.149 |
| ACE-I | 0.020 | 0.135 | 0.142 |
| MoReFlow-T | 0.247 | 0.217 | 0.093 |
| MoReFlow-I | 0.276 | 0.279 | 0.078 |
| Method | Queries | Closer to the target pool | Mean distance to source | Mean distance to target pool | Ratio |
|---|---|---|---|---|---|
| ACE-I | 130 | 103 (79.2%) | 3.449 | 2.215 | 1.56 |
| ACE-T | 130 | 99 (76.2%) | 3.697 | 2.650 | 1.39 |
| Distance computed in a four-dimensional kinematic feature space (centre-of-mass displacement, centre-of-mass variance, mean velocity, foot-contact density). 76–79% of ACE outputs are closer to a target-pool training clip than to the corresponding source clip, weakening source-memorisation as an explanation of ACE’s SIF lift. | |||||
| Raw SIF | Length-controlled SIF | Action AUC | Variation | Realistic | |
|---|---|---|---|---|---|
| True retargeted counterpart | 0.960 | 0.959 | 0.980 | 0.518 | 97.2% |
| Random same-action clip | 0.008 | 0.068 | 0.978 | 0.414 | 97.5% |
| Unpaired objective | 0.203 | 0.154 | 0.499 | 0.001 | 100.0% |
| Averaging objective | 0.379 | 0.343 | 0.820 | 0.048 | 53.0% |
| Model trained on true pairs | 0.900 | 0.698 | 0.967 | 0.521 | 94.6% |
| Prediction | Threshold | Result |
|---|---|---|
| The true retarget is recognized | SIF , interval above 0 | 19 of 24; every interval above 0; misses 0.768–0.799 |
| A random same-group clip sits at the floor | 90% interval within | 12 of 24; 12 of 12 in the denser setting |
| The averaging objective loses variation | variation , upper bound | 24 of 24 |
| The unpaired objective loses variation | variation , upper bound | 22 of 24; missed with variation |
| The true-pair model recovers both | SIF and variation | 24 of 24; SIF 0.77–0.99 |
| True retarget | Random clip | Unpaired | Averaging | True-pair model | ||||
|---|---|---|---|---|---|---|---|---|
| Robot | Setting | Scoring | SIF | SIF | Variation ( ) | Variation ( ) | SIF | Variation ( ) |
| Unitree G1 | three-clip | raw | 0.988 | 0.066 | 0.059 | 0.286 | 0.987 | 1.00 |
| Unitree G1 | three-clip | length-controlled | 0.987 | 0.066 | 0.059 | 0.286 | 0.987 | 1.00 |
| Unitree G1 | denser | raw | 0.966 | 0.061 | 0.012 | 0.035 | 0.965 | 0.99 |
| Unitree G1 | denser | length-controlled | 0.966 | 0.058 | 0.012 | 0.035 | 0.965 | 0.99 |
| Booster T1 | three-clip | raw | 0.799 | 0.024 | 0.055 | 0.347 | 0.845 | 0.99 |
| Raw SIF | Action AUC | Variation | Realistic | |
|---|---|---|---|---|
| Planned run | 0.209 | 0.500 | diverged | 0.0% |
| Re-run, seed 1 | 0.696 | 0.850 | 0.561 | 14.4% |
| Re-run, seed 2 | 0.528 | 0.551 | 0.093 | 78.0% |
| Re-run, seed 3 | 0.574 | 0.513 | 0.101 | 0.0% |
| Method | Raw SIF | Action AUC | Variation | Realistic |
|---|---|---|---|---|
| ACE-I | 0.484 | 0.539 | 0.02 | 91.8% |
| ACE-T | 0.385 | 0.538 | 0.02 | 92.4% |
| MoReFlow-T | 0.203 | 0.596 | 0.25 | 99.4% |
| MoReFlow-I | 0.137 | 0.578 | 0.28 | 99.4% |
| random-same-exact-action | 0.096 | 0.799 | 0.00 | 100.0% |
| random-same-cluster | 0.073 | 0.813 | 0.25 | 100.0% |
| Raw SIF | Action AUC | Variation ( ) | Realistic | |
|---|---|---|---|---|
| Three clips per group | ||||
| True retargeted counterpart | +0.87 (+0.77–+0.99) | 0.61 (0.60–0.63) | 1.00 (1.00–1.00) | 99 (99–99)% |
| Random same-group clip | +0.08 (+0.02–+0.18) | 0.62 (0.61–0.63) | 0.73 (0.69–0.77) | 99 (99–100)% |
| Unpaired objective | +0.24 (+0.11–+0.36) | 0.50 (0.49–0.52) | 0.15 (0.02–0.41) | 98 (89–100)% |
| Averaging objective | +0.11 (+0.07–+0.20) | 0.60 (0.58–0.63) | 0.29 (0.26–0.35) | 46 (33–83)% |
| Model trained on true pairs | +0.89 (+0.85–+0.99) | 0.61 (0.60–0.63) | 1.00 (0.99–1.00) | 98 (97–99)% |