Organizations: School of Mechanical Engineering, Yonsei University, Seoul 03722, Korea · Power Generation Lab, Korea Power Research Institute, 105, Munji-Ro, Yuseong-Gu, Daejeon, 34056, South Korea
Human motion prediction is a crucial capability for advanced robotic systems that interact with humans. In facilities with dynamic human-robot collaboration settings, robots must anticipate human movements to ensure safety, prevent collisions, and optimize cooperative tasks. Traditionally, motion forecasting is treated as a sequential modeling problem using historical pose data, but achieving long-term accuracy and physical realism remains challenging. We present Adversarial Motion Transformer (AdvMT), a novel approach that integrates a Transformer-based motion encoder with a temporal continuity discriminator to address these challenges. The Transformer captures rich spatio-temporal dependencies across human joints, while adversarial training with a continuity discriminator enforces smooth, natural motion trajectories that adhere to biomechanical constraints. Our training scheme includes a bone-length consistency term and adversarial loss to reduce common artifacts like pose freezing or unnatural transitions. In experiments on the Human3.6M motion dataset, AdvMT achieves state-of-the-art long-horizon prediction accuracy while also delivering robust short-term predictions. These improvements strengthen the prediction foundation for physical AI in manufacturing and human-robot collaboration, where anticipating human motion is a prerequisite for safe and efficient robot coordination.
Figures & tables
Fig. 1: Illustration of human-humanoid robot interaction enabled by human motion prediction. The humanoid predicts the future motion of a human approaching on a collision path and adjusts its trajectory accordingly, enabling safe and efficient collaboration.
Fig. 2: Left : Overview of our proposed AdvMT network to predict future human motion by observing historical motion. Right : The human body joints link structure consisting of human body parts: the torso and head, left leg, right leg, left arm, and right arm.
Fig. 3: Overall AdvMT architecture with dual branches. The motion encoder branch is a Transformer-based network that encodes historical human motion and generates future pose predictions. The temporal continuity discriminator observes sequences of predicted poses and outputs a realism score. During training, the discriminator’s feedback guides the encoder to produce smooth, human-like motion transitions. This two-branch adversarial design ensures the predicted motion is both accurate and physically natural.
Walking
Eating
Time (ms)
160
400
560
720
880
1000
160
400
560
720
880
1000
Res. Sup. ∗ [ 13 ]
40.9
66.1
71.6
72.5
76.0
79.1
31.5
61.7
74.9
85.9
93.8
98.0
convSeq2Seq ∗ [ 16 ]
33.5
63.6
72.2
77.2
80.9
82.3
22.4
48.4
61.3
72.8
81.8
87.1
HisRepeat ∗ [ 19 ]
19.5
39.8
47.4
52.1
55.5
58.1
14.0
36.2
50.0
61.4
70.6
75.7
BiTGAN † [ 26 ]
–
–
49.8
55.0
58.5
60.5
–
–
48.5
59.2
68.2
73.0
siMLPe ‡ [ 38 ]
–
39.6
46.8
–
–
55.7
–
36.1
49.6
–
–
74.5
TABLE I: Comparison for short-term ( ≤ 400 ms) and long-term ( > 400 ms) prediction on Human3.6M [ 6 ] across four action categories. The symbols ∗ , † , and ‡ indicate that the corresponding results are taken from [ 19 ] , [ 26 ] , and [ 38 ] , respectively.
Directions
Greeting
Phoning
Time (ms)
560
720
880
1000
560
720
880
1000
560
720
880
1000
Res. Sup. ∗ [ 13 ]
101.1
114.5
124.5
129.1
126.1
138.8
150.3
153.9
94.0
107.7
119.1
126.4
HisRepeat ∗ [ 19 ]
73.8
88.1
100.1
106.4
101.9
118.4
132.7
138.8
67.4
82.9
96.5
105.0
BiTGAN † [ 26 ]
73.3
87.9
99.7
106.3
101.1
117.8
131.4
136.4
67.3
82.3
94.9
103.2
siMLPe ‡ [ 38 ]
73.1
–
–
106.7
99.8
–
–
137.5
66.3
–
–
103.3
AdvMT (ours)
79.2
90.1
99.4
103.5
95.1
104.5
114.1
118.5
68.4
79.6
88.5
93.7
TABLE II: Mean Per Joint Position Error (MPJPE) for long-term prediction ( > 400 ms) on Human3.6M [ 6 ] across remaining action categories. The symbols ∗ , † , and ‡ indicate that the corresponding results are taken from [ 19 ] , [ 26 ] , and [ 38 ] , respectively.
Fig. 5: Qualitative results up to 2 seconds future motion prediction for walking , eating , phoning , and walking together actions from Human3.6M dataset [ 6 ] . For visualization purposes, the predictions are downsampled to 5 frames per second. Ground truth poses are drawn in purple and green, whereas the future predictions are marked in blue and red colors. Best visualized in zoomed view.
Walking
Eating
Smoking
Discussion
Time (ms)
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
HisRepeat [ 19 ]
59.1
61.6
66.5
72.2
73.4
82.6
87.9
90.7
93.1
96.1
77.2
84.0
89.4
94.8
101.8
131.3
138.6
144.3
149.0
151.1
AdvMT (ours)
56.3
59.7
65.4
71.0
73.2
63.6
66.6
68.7
70.5
71.4
84.1
89.4
92.2
95.1
99.1
106.3
109.4
114.1
117.7
121.1
Directions
Greeting
Phoning
Posing
Time (ms)
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
1200
1400
1600
1800
2000
HisRepeat [ 19 ]
116.3
121.4
126.7
130.0
132.9
148.2
152.4
153.1
150.2
150.6
118.9
132.4
144.2
153.4
162.3
202.5
220.4
233.3
246.4
254.6
TABLE III: Extended motion prediction until 2 seconds on Human3.6M [ 6 ] across all 15 action categories. HisRepeat [ 19 ] predictions were generated using the trained models provided by the authors. Bold indicates lower error.
Methods
Time (ms)
160
400
560
880
1000
Baseline [ 23 ]
44.2
79.7
92.2
118.9
126.6
AdvMT ( LMPJPE )
45.8
77.2
88.9
112.9
119.7
AdvMT ( LMPJPE + Lbone + LDK )
33.2
65.3
80.3
100.8
106.6
TABLE IV: Ablation study results on the Human3.6M dataset [ 6 ] .
Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China