Organizations: School of Mechanical Engineering, Yonsei University, Seoul 03722, Korea · Power Generation Lab, Korea Power Research Institute, 105, Munji-Ro, Yuseong-Gu, Daejeon, 34056, South Korea
Human motion prediction is a crucial capability for advanced robotic systems that interact with humans. In facilities with dynamic human-robot collaboration settings, robots must anticipate human movements to ensure safety, prevent collisions, and optimize cooperative tasks. Traditionally, motion forecasting is treated as a sequential modeling problem using historical pose data, but achieving long-term accuracy and physical realism remains challenging. We present Adversarial Motion Transformer (AdvMT), a novel approach that integrates a Transformer-based motion encoder with a temporal continuity discriminator to address these challenges. The Transformer captures rich spatio-temporal dependencies across human joints, while adversarial training with a continuity discriminator enforces smooth, natural motion trajectories that adhere to biomechanical constraints. Our training scheme includes a bone-length consistency term and adversarial loss to reduce common artifacts like pose freezing or unnatural transitions. In experiments on the Human3.6M motion dataset, AdvMT achieves state-of-the-art long-horizon prediction accuracy while also delivering robust short-term predictions. These improvements strengthen the prediction foundation for physical AI in manufacturing and human-robot collaboration, where anticipating human motion is a prerequisite for safe and efficient robot coordination.
Figures & tables
Fig. 1: Illustration of human-humanoid robot interaction enabled by human motion prediction. The humanoid predicts the future motion of a human approaching on a collision path and adjusts its trajectory accordingly, enabling safe and efficient collaboration.
Fig. 2: Left : Overview of our proposed AdvMT network to predict future human motion by observing historical motion. Right : The human body joints link structure consisting of human body parts: the torso and head, left leg, right leg, left arm, and right arm.
Fig. 3: Overall AdvMT architecture with dual branches. The motion encoder branch is a Transformer-based network that encodes historical human motion and generates future pose predictions. The temporal continuity discriminator observes sequences of predicted poses and outputs a realism score. During training, the discriminator’s feedback guides the encoder to produce smooth, human-like motion transitions. This two-branch adversarial design ensures the predicted motion is both accurate and physically natural.
Walking
Eating
Time (ms)
160
400
560
720
880
1000
160
400
560
720
880
1000
Res. Sup. ∗ [ 13 ]
40.9
66.1
71.6
72.5
76.0
79.1
31.5
61.7
74.9
85.9
93.8
98.0
convSeq2Seq ∗ [ 16 ]
33.5
63.6
72.2
77.2
80.9
82.3
22.4
48.4
61.3
72.8
81.8
87.1
HisRepeat ∗ [ 19 ]
19.5
39.8
47.4
52.1
55.5
58.1
14.0
36.2
50.0
61.4
70.6
75.7
BiTGAN † [ 26 ]
–
–
49.8
55.0
58.5
60.5
–
–
48.5
59.2
68.2
73.0
siMLPe ‡ [ 38 ]
–
39.6
46.8
–
–
55.7
–
36.1
49.6
–
–
74.5
TABLE I: Comparison for short-term ( ≤ 400 ms) and long-term ( > 400 ms) prediction on Human3.6M [ 6 ] across four action categories. The symbols ∗ , † , and ‡ indicate that the corresponding results are taken from [ 19 ] , [ 26 ] , and [ 38 ] , respectively.
Directions
Greeting
Phoning
Time (ms)
560
720
880
1000
560
720
880
1000
560
720
880
1000
Res. Sup. ∗ [ 13 ]
101.1
114.5
124.5
129.1
126.1
138.8
150.3
153.9
94.0
107.7
119.1
126.4
HisRepeat ∗ [ 19 ]
73.8
88.1
100.1
106.4
101.9
118.4
132.7
138.8
67.4
82.9
96.5
105.0
BiTGAN † [ 26 ]
73.3
87.9
99.7
106.3
101.1
117.8
131.4
136.4
67.3
82.3
94.9
103.2
siMLPe ‡ [ 38 ]
73.1
–
–
106.7
99.8
–
–
137.5
66.3
–
–
103.3
AdvMT (ours)
79.2
90.1
99.4
103.5
95.1
104.5
114.1
118.5
68.4
79.6
88.5
93.7
TABLE II: Mean Per Joint Position Error (MPJPE) for long-term prediction ( > 400 ms) on Human3.6M [ 6 ] across remaining action categories. The symbols ∗ , † , and ‡ indicate that the corresponding results are taken from [ 19 ] , [ 26 ] , and [ 38 ] , respectively.
Fig. 5: Qualitative results up to 2 seconds future motion prediction for walking , eating , phoning , and walking together actions from Human3.6M dataset [ 6 ] . For visualization purposes, the predictions are downsampled to 5 frames per second. Ground truth poses are drawn in purple and green, whereas the future predictions are marked in blue and red colors. Best visualized in zoomed view.
Walking
Eating
Smoking
Discussion
Time (ms)
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
HisRepeat [ 19 ]
59.1
61.6
66.5
72.2
73.4
82.6
87.9
90.7
93.1
96.1
77.2
84.0
89.4
94.8
101.8
131.3
138.6
144.3
149.0
151.1
AdvMT (ours)
56.3
59.7
65.4
71.0
73.2
63.6
66.6
68.7
70.5
71.4
84.1
89.4
92.2
95.1
99.1
106.3
109.4
114.1
117.7
121.1
Directions
Greeting
Phoning
Posing
Time (ms)
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
HisRepeat [ 19 ]
116.3
121.4
126.7
130.0
132.9
148.2
152.4
153.1
150.2
150.6
118.9
132.4
144.2
153.4
162.3
202.5
220.4
233.3
246.4
254.6
TABLE III: Extended motion prediction until 2 seconds on Human3.6M [ 6 ] across all 15 action categories. HisRepeat [ 19 ] predictions were generated using the trained models provided by the authors. Bold indicates lower error.
Methods
Time (ms)
160
400
560
880
1000
Baseline [ 23 ]
44.2
79.7
92.2
118.9
126.6
AdvMT ( LMPJPE )
45.8
77.2
88.9
112.9
119.7
AdvMT ( LMPJPE + Lbone + LDK )
33.2
65.3
80.3
100.8
106.6
TABLE IV: Ablation study results on the Human3.6M dataset [ 6 ] .
Human motion describes the three-dimensional full-body movement of a person. Anticipating such motion holds significant relevance across a wide range of application domains such as human-robot interaction, autonomous driving, animation, and healthcare. In recent research, spatial and temporal dependencies are modeled by bidirectional attention mechanisms. These typically anticipate human motion in an autoregressive manner which could cause an accumulation of errors over time. As a consequence, they solely focus on local pose forecasting. To address these limitations, we propose a non-autoregressive transformer based on spatio-temporal attention, and train it not only for local pose anticipation, but also for global motion prediction in space. Furthermore, to enhance its applicability in real-world scenarios, our model is also trained to recover missing joints due to occlusions, and is capable of processing varying lengths of history observations. Our code is publicly available at https://github.com/Q-Y-Yang/Prediction-of-Local-and-Global-Human-Motion.
Qiaoyue Yang, Sven Heutger, Christopher Niemann +3
Existing Stochastic 3D Human Motion Prediction models are fundamentally constrained by hard-coding the skeleton kinematics, severely limiting generalization, preventing cross-dataset training, and requiring complex data retargeting. We introduce EquiFusion, the first kinematics-agnostic model to solve this bottleneck, implementing a latent diffusion model with a permutation equivariant architecture. EquiFusion treats the kinematics' connectivity as an explicit input parameter, ensuring its internal computations are inherently agnostic to joint ordering and graph structure. This novel design enables truly cross-dataset generalization to unseen kinematics and unlocks novel zero-shot directions, such as motion prediction from partial or occluded observations and targeted limb generation. EquiFusion achieves state-of-the-art results on major benchmarks, being up to 75% more compact than previous kinematics-specific methods, while achieving faster training and inference. EquiFusion thus establishes a new, flexible standard for robust human motion prediction. Model and training code are available at https://ceveloper.github.io/publications/equifusion/.
Retargeting human motion to humanoid robots is critical for teleoperation, imitation learning and human-robot interaction. However, it remains challenging because of substantial morphological discrepancies between humans and robots, including differences in skeletal topology, limb proportions and degrees of freedom, as well as the scarcity of paired motion data. This paper presents Human2Humanoid, an unsupervised motion retargeting framework that transfers human motions to humanoid robot behaviors with high fidelity. To bridge the domain gap under unpaired data, we adopt a CycleGAN-based architecture equipped with a skeleton-aware graph convolutional network to capture topology-dependent motion features. To address cross-domain scale mismatches, we introduce a morphology-invariant end-effector consistency loss that aligns normalized end-effector trajectories to preserve motion semantics across embodiments. To improve physical plausibility and reduce contact artifacts, we impose explicit physics-aware feasibility constraints to encourage reproduction of the contact patterns in the source motion. Experimental results show that the proposed method successfully retargets human motion to the Unitree G1 humanoid robot without paired data, and outperforms existing methods in both downstream controllability and physical feasibility.
Tianchen Huang, Feiyang Yuan, Junchi Gu +5
Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China