Organizations: The University of Hong Kong, Hong Kong · Shandong University, China · Hong Kong University of Science and Technology, Hong Kong · Adobe Research, UK · TransGP, Hong Kong
We present MotionPersona, a generative framework for character-aware locomotion control, in which the motion for a command depends on the captured persona, the body shape, and the character's style. Unlike style, which one performer can vary at will, persona and body shape are coupled in capture: each performer is observed in only one body. The captured data therefore cannot uniquely determine which motion characteristics should follow the persona and which should change with the body, leaving unseen persona-body combinations unconstrained. We capture 48 performers, aged 5 to 68, under the same nine styles and seven commands, 44 of them with persona annotation. From this repeated-measures design, we identify two robust associations between body shape and gait. These measurements guide a cross-body specification of which characteristics should change and which should be preserved. We implement this specification through a physically informed retargeting pipeline, producing cross-body training data while penalizing penetration and foot skating. On this data we train a single generative controller. A shape-aware VAE compresses each motion block into a few latent tokens and renders them on a conditioned target body under explicit geometric supervision; over these tokens, a latent flow-matching prior generates persona- and style-conditioned motion in two sampling steps. The controller covers all captured personas, a wide family of SMPL-X target bodies, and nine styles in one model, and runs at 27 ms per block on two threads of a laptop CPU. We verify the framework at every stage, following the same gait descriptors from captured to retargeted to generated motion and sweeping each axis in isolation. To our knowledge, this is the first real-time locomotion controller that carries part of a captured persona's performer-specific variation across independently selected body shapes and styles.
Figures & tables
Figure 1 . We present a generative character-aware locomotion controller whose three controls specify a captured motion persona (who is moving), a target body (which body carries the motion), and a performed style (how the character is moving) under varying locomotion commands. Within the captured catalog and target-body family, one shared model recombines the three controls in real time. Code and data will be fully released.
Figure 2 . Why real capture binds persona to body. (a) Each captured performer brings one body, so capture fills only the diagonal of the persona × body grid (schematic; 44 performers in our data); every other cell, such as one persona on another body, is never observed, and the cross design of Section 4 constructs it by retargeting. (b) Within one performer, an 8 ∘ forward lean splits into body, persona, and remainder as B+P+u in many ways; moving any g from P to B explains the clips equally well.
Figure 3 . Wider hips go with more forward and side trunk lean, the two associations assigned to shape. Grey points are the adult performers’ effects on trunk lean, the performer term of Equation 2 , against hip width; blue points are the means of the four hip-width quartiles with standard errors. The dashed line and the band are the fit over all 39 adults and its bootstrap 95% interval; the pair slope comes from the twelve metadata-matched adult pairs that share all three attributes but differ in body.
Figure 4 . Cross-body retargeting. One frame of an angry forward walk on the source body (left) and seven target bodies of increasing height and girth (right: a captured child and six grid bodies), shown with a fixed camera rig and SMPL-X body masses. Stance widens, arms clear the torso and thighs, and trunk lean follows the assigned hip-width associations. On the widest body the plain copy penetrates by up to 14 cm; the optimizer brings this to 1.6 cm, and the mean forward lean runs from −6.7∘ on the child body to +3.5∘ on the widest.
Capture
Assigned
MotionBuilder
Ours
Sign
Forward lean, ∘
2.0
1.4
0.0
1.9
0.92
Side lean, ∘
0.4
0.6
0.0
0.6
0.79
Hand distance
0.23
–
0.00
0.03
0.54
Cadence (preservation target), Hz
0.23
–
0.00
0.02
0.54
Penetration, % of frames ↓
8.9
0
8.2
0.0
–
Lowest toe, cm
−4.1
≥0
−4.0
−0.6
–
Table 1 . Ours moves the lean toward the real partner and removes floor penetration; MotionBuilder does neither. Upper rows: medians over 24 transfers within twelve pairs; Capture is the real difference between the two performers, Assigned the part the associations predict from their hip widths, and Sign the fraction of transfers in which ours moves toward the partner. Lower rows: 100 clips both retargeters produced on the same body.
Figure 5 . The controller: a shape-aware VAE realizes the motion on the target body, and a latent prior selects it. Dashed outlines mark the modules each stage trains. The shape-aware VAE decoder (a) realizes every block on the target body, continues the previous block, and is trained with explicit forward-kinematic and contact losses. The prior (b) generates the next latent tokens from the command, target shape, persona ID and typed attributes, and style, in two steps.
Condition
Form
Enters
Drop
History
last 5 frames
decoder
–
1 token, codec encoder
prior
15%
Trajectory
45 positions, 45 facings
prior
–
Persona ID s
44-way learned embedding
prior
–
Attributes as
three typed embeddings
prior
–
Style y
label embedding
prior
–
Table 2 . Condition design and training dropout. Only history is dropped, enabling guidance in Equation 13 .
Step
ms
Featurization of the history and the trajectory
0.3
History encoder, one token
1.1
Prior, two Euler steps
13.7
Decoder, per-frame queries and temporal head
5.6
Block, M3 Max CPU, two threads
26.5
Block, RTX 5080 GPU
9.9
Table 3 . Block runtime. Steps are timed separately on two threads of an Apple M3 Max CPU and include individual call overheads; totals are measured end-to-end, not summed.
Quality
Control
Spread kept
Test set
FPD pose ↓
FPD contact ↓
Jerk ↓
Skate ↓
IoU ↑
Pen. % ↓
Seam ↓
Rec. ↑
Style ↑
Dyn. ↑
Amp. ↑
MotionPersona-X (data)
own body
–
0.83
0.51
0.34
0.57
0.5
–
0.889
0.80
1.00
1.00
CAMDM ( Chen et al., 2024 )
own body
0.30
1.90
0.65
0.60
0.49
49.0
3.52
0.323
0.32
0.43
0.55
cross-body, seen
0.28
3.28
0.51
0.53
0.50
21.8
3.10
0.257
0.51
⋅
⋅
cross-body, withheld
0.27
3.25
0.51
0.53
0.51
21.8
3.10
0.255
0.42
⋅
⋅
held-out bodies
0.28
3.42
0.51
0.53
0.50
23.9
3.11
0.256
0.40
⋅
⋅
Table 4 . Ours leads the baseline in control and in most quality columns on all four test sets, and its quality holds off the performer’s body. Own body: every persona in every style on its own body; cross-body: every persona in every style on training bodies it was paired with, on training bodies whose pairing was withheld, and on bodies training never saw. Jerk is the median third difference of the toe positions in cm per frame 3 (the data value is measured on the reference takes), skate and seam in cm per frame, penetration in percent of frames, recovery and style accuracy as retrieval rates, with the held-out takes retrieved against each other as the data reference; spread kept exists on the own-body set only. Own body + retargeter is the modular pipeline: our controller generates on the performer’s own body, and our retargeter carries the result to the target body. Bold marks the best method per test set.
Figure 6 . The baseline loses the floor at both ends of the height range; ours keeps it. One persona walks straight on three sweep bodies (thumbnails at a common scale); blue marks hovering feet, red the sole below the floor. Numbers are the median height of the lowest foot vertex above the floor over 19 s; within ±1 cm counts as on the floor.
Figure 7 . One control at a time, with the command fixed. (a) One persona on five bodies, ordered by hip width: the trunk (blue, against the grey vertical) leans back less as the hips widen. (b) Five personas on one body: the foot-contact strips, four seconds of the left and right foot, show five rhythms. (c) Five styles on one persona and body change the whole manner of moving.
Figure 8 . The single-axis test, scored against the data. Each column changes one control and holds the other two and the command. (a) Body: forward lean and step length against hip width for all 44 personas on the sixteen sweep bodies, in the held-out takes (black), their codec reconstruction (grey, dashed), and the controller’s rollouts (blue); lines are means, bands the 10th to 90th percentile, numbers the mean and spread of the per-persona slopes per metre. (b, c) Persona and style: per descriptor, the spread the controller produces across the 44 personas or the nine styles, as a fraction of the data’s spread on the same body with each command’s mean removed; (c) averages over the personas; median over bodies with a bootstrap interval, 1 reproduces the data. Dots name the factor that owns the descriptor in the grid; open rows are fixed by the command and should stay near 0; ticks in (b) are a prior trained without a persona input.
Lean slope
Preserved slope
Other bodies
fwd
side
cadence
step
Pen. % ↓
Pres. ↓
Rec. ↑
target
150
65
0
0
0
0
MotionPersona-X (data)
150
65
−0.3
+0.1
0.3
–
–
Codec reconstruction
155
55
0.0
0.0
0.5
–
–
Captured grid
22
−8
−0.2
+0.1
60.9
0.25
0.235
Geometric copies
−77
6
+0.7
−0.2
0.0
0.97
0.341
Table 5 . What each step of the cross design contributes. The same controller trained on the captured grid, on the geometric copies, on the optimizer’s output without the lean targets, and on MotionPersona-X, with the data and the codec’s reconstruction above. Slopes per metre of hip width on the sweep bodies, targets under the headers (over all 44 personas, ours gives 115 and 28); penetration and preservation on the same bodies, recovery on the held-out bodies. Bold marks the trained row nearest its target; the copies row does not touch the floor, so its penetration is not bolded. Pres. is the relative change of cadence and step length from the own-body reference.
Figure 9 . The baseline averages two performers’ pelvis bob; ours keeps part of the difference. Left: side-view strobes over 1.4 s, each performer on her own body at a common scale, with the pelvis trace exaggerated five times. Right: pelvis height minus its trend at true scale; labels give the peak-to-peak amplitude over 15 s.
FPD ↓
ms
Backbone
pose
contact
Pen. % ↓
Skate ↓
IoU ↑
Style ↑
Spread ↑
GPU
CPU
Data
–
0.83
0.5
0.34
0.57
0.80
1.00
–
–
FM pose, 2 steps
1.74
288
2.2
–
0.00
0.52
–
6.2
44
DDPM pose, 4 steps
0.28
1.90
18.9
0.57
0.49
0.37
0.61
9.9
87
FM token, 2 steps
0.32
1.70
1.0
0.49
0.54
0.74
0.70
9.9
27
Table 6 . One transformer trained three ways on MotionPersona-X, own-body set: flow matching in pose space, four-step diffusion in pose space with CAMDM’s sampler, and two-step flow matching in the codec’s token space, the controller. Spread is the amplitude group of Table 4 ; ms per block on an RTX 5080 and on two CPU threads. Bold marks the best of the two rows that walk; the first does not walk under either commit schedule and is reported under the three-frame one.
Criterion
Ours
CAMDM
Realism of the motion ↑
3.97
3.22
Match to the persona ↑
4.17
2.59
Match to the style ↑
3.91
3.82
Body-shape artifacts ↓
1.26
2.17
Responsiveness ↑
3.71
3.31
Follows the command ↑
4.12
3.79
Table 7 . Practitioners rate ours higher on five criteria and level on style: mean of all ratings collected from 28 professionals on a five-point scale. Lower is better for artifacts; bold marks the better controller.
Reconstruction
Rollout
Codec
IoU ↑
Lean → 1
Skate ↓
IoU ↑
Seam ↓
Pen. % ↓
Lean → 1
Reference
1.00
1.03
0.49
0.54
3.45
1.0
0.70
No β in the decoder
1.00
0.99
0.55
0.35
3.46
42.0
0.61
No boundary frames
1.00
1.02
0.62
0.53
3.81
0.9
1.17
Table 8 . Ablations on the codec, on MotionPersona-X: one decision changed and the prior retrained. Reconstruction columns are frame-aligned with the source, with the reference-label IoU; rollout columns follow the protocol on the own-body set. Lean is the forward-lean slope as a fraction of the data’s, after reconstruction and after rollout, with a target of 1; bold marks the best measured row, and for Lean the row nearest 1.
FPD ↓
Prior
pose
contact
Skate ↓
IoU ↑
Track. ↓
Rec. ↑
Style ↑
Spread ↑
Lean → 1
Reference
0.32
1.70
0.49
0.54
5.5
0.632
0.74
0.70
0.70
Dropout 0
0.29
2.23
0.53
0.48
2.9
0.494
0.73
0.66
1.28
No shape token
0.33
1.90
0.65
0.48
7.8
0.629
0.73
0.76
0.61
No persona
0.48
2.24
0.52
0.51
5.9
0.318
0.53
0.19
1.00
ID only
0.32
1.70
0.49
0.54
5.4
0.644
0.66
0.69
0.82
Table 9 . Ablations on the prior, on MotionPersona-X, own-body set: one decision changed per row on the final codec. Track. is the root error in cm, Rec. persona recovery, Style the style accuracy, Spread the amplitude group, Lean the forward-lean slope as a fraction of the data’s on the sweep bodies, with a target of 1; bold marks the best measured row, and for Lean the row nearest 1. Rows within 0.03 of skate or 0.04 of recovery of the reference are within one training seed’s spread and are not ranked.
Figure 10 . In-betweening on a 105 cm body: MotionBricks ( Wang et al., 2026 ) , which has no body input, sinks the feet into the floor for the whole gap, whereas ours keeps them where the data has them in this example. Both receive the same keyframes (gold) and root path of a backward walk and generate the 40 frames between them. (a) Poses spread horizontally for legibility, with the feet shown at frame 5. (b) Lowest toe joint height: minimum +1.8 cm in the data, +1.9 cm for ours, and −2.1 cm for MotionBricks.
Figure 11 . Speech-driven gesture: one speech clip of a BEAT2 speaker drives five other speakers on their own bodies, short to tall, each in its own manner. Left, the speaker’s motion capture; feet locked after generation.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Performer profile and source body
Style
Command
Take name
Age
Gender
Height (cm)
Weight (kg)
Beta
Role
Affiliation
Dominance
Label
Group
Base movement
s001_angry_fw
5
M
112
17
[-5.68, 1.26, …]
Child
Reserved
Compliant
Angry
Affective
forward walking
s015_neutral_bw
22
F
165
47
[0.96, 0.27, …]
Student
Sociable
Assertive
Neutral
Affective
backward walking
s032_twofootjump_fr
23
M
189
80
[1.44, -1.35, …]
Athlete
Gregarious
Dominant
Two-foot jump
Locomotion variant
forward running
Appendix
Table 10 . Example annotations. Each clip is linked to a closed-set performer ID, three typed persona attributes, the performer’s source-body metadata, a style label, and a locomotion command. Beta is the 10-dimensional SMPL-X shape vector; age and gender describe the capture cohort but are not persona conditions.
Locomotion coverage
Performers
Scale and annotation
Dataset
Accessible
Forward
Backward
Sideways
Fingers
#Perf.
Age
Height
Weight
Text
#Seq
Hours
SMPL
Edinburgh ( 2017 )
✓
✓
✓
✓
✗
–
✗
✗
✗
✗
80
1
✗
LAFAN1 ( 2020 )
✓
✓
✓
✓
✗
5
✗
✗
✗
✗
77
4.6
✗
BFA ( 2020b )
✓
✓
✗
✗
✗
1
✗
✗
✗
✗
33
1.5
✗
100STYLE ( 2022 )
✓
✓
✓
✓
✗
1
✗
✗
✗
✗
810
18.75
✗
Multi-sub ( 2024 )
✗
✓
✓
✗
✗
12
✗
154–195
✗
✗
–
4
✗
Appendix
Table 11 . Comparison with locomotion datasets. MotionPersona combines performer diversity, repeated styles, repeated commands, and SMPL-X bodies. “Text” denotes performer- or style-level annotation rather than motion-content captions.
Figure 12 . Top: eight of the SMPL-X bodies fitted to the performers. Bottom: the distribution of height, weight, and age. Blue points are male performers, orange points female.
Performer
Ang
Big
Dep
Dru
Fea
Hap
Neu
Swi
Jmp
Sum
Performer
Ang
Big
Dep
Dru
Fea
Hap
Neu
Swi
Jmp
Sum
P01
6
6
6
7
6
6
6
6
6
55
P23
6
6
6
6
6
6
6
6
0
48
P02
7
7
7
7
7
7
7
7
7
63
P24
7
7
7
7
7
7
7
7
7
63
P03
6
4
5
3
3
1
7
1
5
35
P25
7
7
7
7
7
7
7
7
7
63
P04
7
7
7
7
7
7
6
7
7
62
P26
7
7
7
7
7
7
6
7
7
62
P05
7
7
7
7
7
7
7
7
7
63
P27
7
7
7
7
7
7
7
7
7
63
P06
7
7
7
7
7
7
7
7
7
63
P28
7
7
7
7
7
7
7
7
7
63
Appendix
Table 12 . Clips per performer and style in the analysed grid; each style holds at most seven clips, one per command. Sixteen performers fill all 63 cells, and P33, P43, P37, P03, and P29 miss many. Columns are angry, big step, depressed, drunk, fear, happy, neutral, swimming, and two-foot jump.
Group
Descriptor
Definition and estimator
Unit
Timing
Cadence
Steps per second, 2/T , where the stride period T is the first autocorrelation peak of each ankle’s height; the median over 10-second sliding windows
steps/s
Cadence (events) †
Number of contact onsets of both feet divided by the duration; kept only as a check, since it undercounts steps
steps/s
Contact ratio
Fraction of frames in which a foot is in contact; contact starts when the ankle or toe is lower than max(3%leg,2cm) and ends when both are higher than max(6%leg,4cm)
fraction
Contact ratio (labels) †
The same from the stored contact labels (foot speed below a threshold and height at most 3 cm), which miss most stance frames in running
fraction
Double support
Fraction of frames with both feet in contact
fraction
Step regularity
Coefficient of variation of the step intervals (higher is less regular)
–
Appendix
Table 13 . The 27 gait descriptors, computed in the pelvis-facing frame from forward-kinematics joint positions. Lengths are divided by the performer’s leg length, hip width, or half shoulder width, so similarity scaling leaves them unchanged. Daggers mark the four reference quantities, which enter only the variance decomposition and the first body fits, as checks.
Descriptor
Style
Command
Performer
Interaction
Residual
Noise
Reliability
Speed
0.02
0.21
0.52
0.10
0.16
0.09
0.86–0.93
Jerk
0.03
0.13
0.51
0.20
0.14
0.02
0.76–0.96
Sway
0.02
0.21
0.47
0.11
0.20
0.09
0.86–0.87
Contact ratio
0.10
0.28
0.33
0.09
0.20
0.09
0.87–0.92
Forward lean
0.25
0.02
0.25
0.36
0.13
0.04
0.73–0.94
Cadence
0.04
0.16
0.21
0.18
0.42
0.02
0.73–0.74
Appendix
Table 14 . Variance decomposition on the grid, as fractions of total variance under the additive model in the main paper. Residual is what the four terms leave; for cadence most of it is estimator variance. Noise is the robust within-clip floor, from the median absolute deviation over 10-second segments. Reliability is the split-half correlation of the performer effect, across style groups and across walking versus running.
Order A: style first
Order B: performer first
Noise floor
Split-half reliability
Descriptor
Style
Command
Performer
Interaction
Residual
Performer
Command
Style
Raw
Robust
Styles
Walk/run
Cadence
0.04
0.16
0.21
0.18
0.42
0.21
0.16
0.04
0.65
0.02
0.73
0.74
Cadence (events) †
0.04
0.27
0.30
0.12
0.27
0.30
0.27
0.04
0.22
0.11
0.81
0.88
Speed
0.02
0.21
0.52
0.10
0.16
0.52
0.20
0.03
0.24
0.09
0.86
0.93
Step length
0.16
0.15
0.27
0.12
0.30
0.27
0.14
0.17
0.54
0.12
0.90
0.87
Contact ratio
0.10
0.28
0.33
0.09
0.20
0.33
0.28
0.09
0.24
0.09
0.92
0.87
Appendix
Table 15 . Variance decomposition of all 27 descriptors, as fractions of the total sum of squares. The performer share exceeds the robust noise floor for every descriptor except step regularity and head sway, and changes by at most 0.01 when the performer term is fitted first. Interaction and residual are identical in both orders and listed once; the robust noise floor uses a MAD-based variance, because about 3% of the 10-second cadence and step-length estimates are octave errors.
Body, all 44
Body, 39 adults
Body, nested
Attributes
Descriptor
Full
Compact
Full
Compact
All
Adults
Strongest adult correlate ( r )
All
Adults
Body residual
Cadence
− 0.03
− 0.17
− 0.00
− 0.08
− 0.12
− 0.03
arm/leg ( − 0.20)
− 0.01
− 0.02
− 0.12
Cadence (events) †
− 0.08
− 0.11
− 0.04
− 0.02
–
–
BMI ( − 0.32)
0.08
0.09
0.07
Speed
− 0.30
− 0.23
0.06
− 0.04
− 0.13
− 0.07
arm/leg ( − 0.25)
0.02
0.07
− 0.06
Step length
− 0.29
− 0.27
0.24
0.20
− 0.32
0.22
BMI ( − 0.33)
0.00
0.09
− 0.11
Contact ratio
− 0.07
− 0.07
− 0.12
− 0.08
− 0.16
0.07
arm/leg ( − 0.33)
− 0.22
− 0.23
− 0.21
Appendix
Table 16 . Body measures and attributes predict little of the performer effect. Values are leave-one-performer-out R2 , negative when worse than the mean; full uses the nine body measures, compact leg length, BMI, child, and female, and nested selects one or two measures inside each training fold (reference quantities not fitted). Attributes are the role one-hot with ordinal affiliation and dominance; the last column fits them to the residual of the better body model.
Leave-one-out R2
Correlation of the pair
Descriptor
Body pair
Pair
Age
Pair + age
Pair, under 53
Raw r
r after age
Speed
height, BMI
0.07
0.30
0.27
− 0.10
+ 0.24, − 0.21
+ 0.05, + 0.02
Step length
height, BMI
0.26
0.37
0.38
− 0.07
+ 0.32, − 0.33
+ 0.14, − 0.12
Foot clearance
height, weight
0.16
0.20
0.20
− 0.03
+ 0.21, − 0.14
+ 0.05, − 0.12
Pelvis bob
height, arm/leg
0.13
0.23
0.29
0.02
+ 0.29, − 0.50
+ 0.12, − 0.43
Pelvis sway
height, BMI
0.06
0.19
0.15
− 0.09
+ 0.24, − 0.19
+ 0.08, + 0.00
Appendix
Table 17 . Age control and partial correlation for the 15 descriptors whose best adult pair of body measures has a positive leave-one-out R2 . Refitting the pair on the 31 adults younger than 53 leaves an R2 of at most 0.02 for speed, step length, foot clearance, pelvis bob and sway, knee range, and joint speed, while forward and side lean keep their hip-width correlation after partialling out age. The pair is chosen on the same 39 adults, so its R2 is optimistic.
Role / affiliation / dominance
Performers
Pairs
Athlete/sociable/assertive
P12 (189/0.88), P22 (180/0.85), P38 (168/0.74)
3
Professional/moderate/dominant
P34 (160/0.79), P39 (160/0.72)
1
Retiree/moderate/compliant
P01 (159/0.74), P33 (173/0.73)
1
Student/gregarious/assertive
P15 (182/0.85), P32 (175/0.75)
1
Student/reserved/moderate
P28 (160/0.69), P29 (163/0.73), P36 (171/0.80)
3
Student/withdrawn/submissive
P03 (165/0.73), P16 (179/0.80), P31 (164/0.72)
3
Appendix
Table 18 . The twelve metadata-matched adult pairs come from six attribute cells shared by two or three adults of different body. Each performer is listed with height (cm) and leg length (m); the two child cells, with two more pairs, are left out.
Descriptor
Body measure
Pairs
Adult slope
Paired slope
95% interval
Agree
ρ
p
Forward lean
hip w.
12
+ 226
+ 131
[ + 41.3, + 249]
0.67
0.71
0.01
Forward lean
BMI
12
+ 0.48
+ 0.177
[ − 0.0592, + 0.612]
0.58
0.59
0.04
Forward lean
weight
12
+ 0.0966
+ 0.0721
[ + 0.00871, + 0.193]
0.67
0.72
0.01
Step length
BMI
12
− 0.00855
+ 0.00161
[ − 0.00598, + 0.00883]
0.50
− 0.06
0.85
Step length
weight
12
− 0.000453
+ 0.000497
[ − 0.00114, + 0.002]
0.58
0.13
0.68
Knee range
BMI
12
− 0.679
− 0.117
[ − 1.52, + 0.686]
0.50
− 0.29
0.35
Appendix
Table 19 . Within the metadata-matched adult pairs, only forward lean (on hip width, BMI, and weight) and side lean (on hip width) keep a rank correlation with p<0.05 . The adult slope is fitted across the 39 adults, the paired slope through the origin on the within-pair differences with a bootstrap 95% interval, in descriptor units per unit of the body measure (hip width and leg length in metres, height in cm, weight in kg). Agree is the fraction of pairs whose difference has the sign of the adult slope; ρ is the Spearman correlation of the pair differences.
Figure 13 . The 44 annotated performers grouped into their 33 role, affiliation, and dominance cells and plotted at their body heights. Twenty-five cells hold one performer; the other eight hold two or three, 14 pairs in all, which provide the metadata-matched identities of the real-pair check, not one persona in several bodies.
Age, 39 adults
53 and over
R2
Descriptor
r
R2
p
+20 years
d
p
Flag
Under 53
Cadence
− 0.20
− 0.07
0.34
− 0.102
− 0.41
0.27
− 0.07
− 0.08
Speed
− 0.61
0.30
0.00
− 0.192
− 1.54
0.00
0.23
− 0.00
Step length
− 0.65
0.37
0.00
− 0.0687
− 1.78
0.00
0.30
− 0.03
Contact ratio
0.22
− 0.10
0.62
+ 0.0235
0.57
0.46
–
–
Double support
0.35
0.00
0.05
+ 0.0576
0.98
0.13
–
–
Appendix
Table 20 . Age lowers speed, step length, foot clearance, pelvis bob and sway, knee and hip range, head bob, and joint speed among the adults, and the effect is carried by the eight performers aged 53 and over. Effects per +20 years are in the units of Table 13 ; d and the Welch p compare the eight with the 31 younger adults. The last two columns are the leave-one-out R2 of the age-53 flag and of age within the younger adults alone.
P43
P37
P13
P04
P26
Difference
% of mean
d
105 cm, 6 y
110 cm, 5 y
112 cm, 5 y
118 cm, 8 y
Small
142 cm, 9 y
P26 / small
Cadence
+ 0.548
+ 22%
1.15
+ 0.3
+ 0.5
− 0.3
+ 1.1
0.42
+ 5.1
12.24
Speed
+ 0.248
+ 25%
0.90
+ 2.2
+ 2.5
+ 0.7
+ 1.2
1.48
− 1.8
− 1.20
Step length
+ 0.0743
+ 17%
0.77
+ 2.3
+ 3.0
+ 0.9
+ 0.6
1.48
− 2.3
− 1.59
Contact ratio
− 0.0308
− 4%
− 0.36
+ 0.4
− 1.8
− 0.8
+ 0.1
–
+ 0.3
–
Double support
− 0.0303
− 5%
− 0.24
+ 0.4
− 1.2
− 0.8
+ 0.1
–
+ 0.4
–
Appendix
Table 21 . The children differ from the adults most in foot clearance and step width, but the tallest child is not an interpolation between the small children and the adults. Difference is children minus adults on the performer effect in the units of Table 13 , also as a percentage of the grand mean; per-child values are z -scores against the adults, with P43’s from two clips. Small is the mean z of P37, P13, and P04, and the last column is P26’s z as a fraction of it.
Forward lean
Side lean
Subset
Slope [95% interval]
R2
Slope [95% interval]
R2
Angry
+ 3.06 [ + 0.18, + 6.21]
+ 0.00
+ 0.43 [ − 0.05, + 1.12]
− 0.06
Big step
+ 5.11 [ + 1.80, + 9.34]
+ 0.10
+ 0.77 [ + 0.29, + 1.27]
+ 0.07
Depressed
+ 9.18 [ + 3.69, + 15.29]
+ 0.10
+ 1.21 [ + 0.43, + 1.99]
+ 0.08
Drunk
+ 6.51 [ + 1.05, + 13.45]
+ 0.04
+ 1.57 [ + 0.54, + 2.59]
+ 0.14
Fear
+ 10.75 [ + 4.90, + 18.97]
+ 0.12
+ 1.01 [ + 0.18, + 2.18]
− 0.01
Appendix
Table 22 . The slopes of forward and side lean on hip width are positive in every style and in walking, running, and transitions, although several single-style intervals include zero. Slopes are degrees per +2 cm of hip width, fitted on the adults’ per-performer means, with bootstrap 95% intervals and leave-one-out R2 .
n
p
True R2
Median
5%
95%
P(>0.15)
P(>0)
39
1
0.00
− 0.09
− 0.14
0.02
0.00
0.07
39
1
0.05
− 0.04
− 0.12
0.14
0.05
0.32
39
1
0.10
0.01
− 0.11
0.20
0.11
0.54
39
1
0.15
0.06
− 0.10
0.28
0.24
0.68
39
1
0.20
0.10
− 0.06
0.31
0.40
0.84
39
1
0.30
0.22
0.02
0.44
0.74
0.96
Appendix
Table 23 . At 39 or 44 performers, leave-one-out R2 falls well below the true R2 , and a true effect of 0.15 exceeds 0.15 in at most a quarter of the simulations. Features are Gaussian with one true linear effect; p=1 is univariate least squares, p=4 ridge on four features of which one is real. The columns give the median and the 5th and 95th percentiles of the leave-one-out R2 , and the probability that it exceeds 0.15 and 0.
39 adults
All 44
Descriptor
Share
Reliability
Shape
+ age
Shape
+ age
Child
Cadence
0.21
0.84
0.00
0.00
0.00
0.00
0.00
Speed
0.52
0.92
0.00
0.30
0.01
0.29
0.00
Step length
0.27
0.93
0.22
0.37
0.03
0.28
0.00
Contact ratio
0.33
0.93
0.07
0.07
0.04
0.04
0.00
Double support
0.28
0.93
0.00
0.00
0.03
0.03
0.00
Appendix
Table 24 . A shape predictor fitted on the adults explains 7.5% of the performer differences and 2.0% of the total variance; allowing age nearly doubles both. Share is the performer share sd , and reliability the Spearman–Brown split-half reliability that bounds any predictor. The adult columns use shape alone or shape and age; the columns for all 44 use shape with the child flag, shape and age, or the child flag alone.
Term
What it asks
Depends on
w
Fidelity (rotation, root)
stay close to the kinematic copy
–
1, 0.01
Penetration
penalize block overlap beyond an allowance
block distance fields
1
Ground
penalize sole corners below the floor
foot geometry
1
Stance
penalize contact height away from the floor
foot geometry
1
Slide
penalize corner velocity during contact
foot geometry
5
Anchor
penalize drift from the copy’s contact-run anchors
foot geometry
10
Appendix
Table 25 . Terms of the retargeting objective, what in the body each depends on, and the weights of the production run.
Figure 14 . The evaluation trajectory, colored by commanded speed; every test case follows this one-minute control sequence, scaled with leg length.
Shape-aware VAE
Latent FM prior
Raw-space baseline
Input
45 frames
9 tokens × 48
55 frames
Token span
5 frames
–
–
Decoder read-out
5 frame queries per token; temporal head 3× depthwise conv, kernel 5, residual
–
–
Layers × width
6 + 6 × 384
8 × 384 (DiT)
8 × 256
Feed-forward, heads
1024, 4
1024, 4
1024, 4
Dropout
0.1
0.2
0.2
Appendix
Table 26 . Architecture and training of the two stages and of the raw-space baseline. Every epoch is about 6.1 M windows: 6,000 batches of 1,024 split over the GPUs for the two stages, and 2,667 steps of 2,304 windows for the baseline, whose per-GPU batch of 384 with two accumulation steps fits the memory of one card. The baseline pins its history prefix at every denoising step and therefore trains without history dropout.
Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches assume identity-neutral motion and generate movements using a canonical body representation, ignoring the strong influence of body morphology on motion dynamics. In practice, attributes such as body proportions, mass distribution, and age significantly affect how actions are performed, and neglecting this coupling often leads to physically inconsistent motions. We propose an identity-aware motion generation framework that explicitly models the relationship between body morphology and motion dynamics. Instead of relying on explicit geometric measurements, identity is represented using multimodal signals, including natural language descriptions and visual cues. We further introduce a joint motion-shape generation paradigm that simultaneously synthesizes motion sequences and body shape parameters, allowing identity cues to directly modulate motion dynamics. Extensive experiments on motion capture datasets and large-scale in-the-wild videos demonstrate improved motion realism and motion-identity consistency while maintaining high motion quality. Project page: https://vjwq.github.io/IAM
Existing humanoid locomotion systems primarily focus on stability and task execution, while integrating expressiveness with explicit locomotion control remains challenging. We propose EMoG, an emotion-modulated gait generation framework for expressive humanoid locomotion. EMoG introduces an emotional-style code with continuously adjustable intensity. Conditioned on this code and physical commands, a lightweight MLP generates expressive, command-consistent periodic gait trajectories in real time, which are tracked by a unified reinforcement learning policy for physical execution. To support training, we collect a large-scale emotion-annotated gait dataset from professional performers and develop an automated pipeline to extract physically consistent periodic gait cycles. EMoG also integrates an LLM-based parser that converts free-form language into emotional style and motion parameters for interactive control. Experiments demonstrate that our system achieves continuous gait-style modulation with perceptible expressive cues while maintaining command tracking. EMoG provides a practical approach to parameterized emotional-style walking for human-robot interaction.
Yi Lu, Tianhao Jiang, Honglong Tian +6
School of Electronic Science and Engineering, Nanjing University, Nanjing 210023, China. · School of Intelligence Science and Technology, Nanjing University, Suzhou 215163, China. · School of Artificial Intelligence, Nanjing University, Nanjing 210023, China. +1
Developing controllers capable of completing a wide range of tasks in a natural and life-like manner is a key challenge in enabling practical applications of physics-based character animation. In this work, we introduce Generative Pretrained Controllers (GPC), which leverage tokenization and next-token modeling to create general-purpose, reusable generative controllers from large-scale motion datasets. Our framework utilizes end-to-end reinforcement learning to jointly optimize a "motion vocabulary", modeled via Finite Scalar Quantization (FSQ), along with a corresponding control policy that can map the discrete codes to physics-based controls. After the "codebook" has been learned, the underlying structure of this large vocabulary is modeled by training a GPT-style autoregressive transformer, leading to a powerful generative controller that generates controls for a physically simulated character by performing next-token prediction. Once the generative controller has been trained, we propose a suite of adaptation techniques for finetuning the controller for new downstream tasks. Our proposed framework greatly simplifies the training process compared to previous tokenized methods, and achieves a 99.98% success rate in reproducing a vast corpus of motion clips. The generative controller exhibits a variety of natural emergent behaviors, such as responsive behaviors to perturbations and recovery behaviors after falling. This results in highly robust general purpose controllers for a variety of downstream applications.
Yi Shi, Yifeng Jiang, Chen Tessler +1
Simon Fraser University, Canada and NVIDIA, USA · NVIDIA, USA · Simon Fraser University, Canada and NVIDIA, Canada