Human pose forecasting predicts future human poses from past observations for applications including action recognition, sports analysis, and human-robot interaction. However, progress in the field is difficult to assess because reported results often rely on heterogeneous preprocessing, inconsistent metric implementations, and partially released artifacts. This paper re-examines human pose forecasting as an evaluation problem by auditing a wide range of recent forecasting methods, identifying several reproducibility issues, and introducing a unified training and evaluation pipeline for absolute pose forecasting. Under this unified protocol, mature cross-domain sequence models adapted from speech recognition provide strong baselines and obtain the lowest errors among the evaluated models. To study deployment-relevant noise, a paired clean/detector-noisy dataset variant is introduced using poses generated by a pose-estimation model. The experiments show that detector-generated poses cause substantial performance degradation compared with clean motion-capture inputs, but that part of this degradation can be recovered by adapting the forecaster on detector-generated pose sequences.
Figures & tables
Figure 1 : Example of an absolute pose forecast of a walking person. The green skeletons visualize the input sequence of the prediction model, the red ones the predicted future poses, and the blue ones the ground-truth labels.
Method
Reproduction outcome
Main issue
MotionMixer [ 5 ]
corrected result lower
averaging/counting bug
STSGCN [ 31 ]
corrected result lower
MPJPE implementation bug
GMFnet [ 30 ]
checkpoint reproduced
training/checkpoint mismatch
SPGSN [ 20 ]
retrained result lower
unresolved training mismatch
HisRepItself [ 38 ]
reproduced by code
–
AuxFormer [ 40 ]
checkpoint reproduced
training/checkpoint mismatch
Table 1 : Summary of reproducibility audit. The table reports the dominant issue observed for each method. “Checkpoint reproduced” means that the released checkpoint matched the reported result, but training from the released code did not reproduce the same performance. “Reproduced by code” means that the model was successfully trained and evaluated using the publicly released code. “Corrected result lower” means that an evaluation issue was identified and the corrected score was worse than the originally reported score. Detailed numerical comparisons and forensic notes are provided in Appendix B .
Method
Type
Size
MPJPE
FPS
FCE
FADE
Repeat last frame
-
-
301
-
-
301
Last delta average
-
-
264
-
-
264
Ridge Regression
-
-
185
-
-
185
TBiFormer ’2023
abs
5.64M
301
293
7
302
PGBIG ’2022
rel
2.65M
279
35
57
287
GMFnet ’2024
rel
11.3M
279
51
39
284
Table 2 : Unified comparison of absolute and relative methods in the task of absolute pose forecasting on Human3.6M . MPJPE and FADE are measured at timestep 1000ms . The error metrics are in millimeters. FPS is tested on a computer with an AMD-9900X and a single Nvidia-RTX4080. A batch-size of 1 was used for testing. Note that the size and speed of some models depends on the number of input and output timesteps (here: 50 in, 25 out).
Method
Size
MPJPE
FPS
FCE
FADE
DeepSpeech ’2014
2.35M
198
653
3
198
QuartzNet ’2020
3.17M
168
1300
2
168
Squeezeformer ’2022
3.31M
149
960
2
149
Conformer ’2020
3.22M
147
1049
2
147
MotionConformer
9.23M
143
929
2
143
Table 3 : Results of converted speech models on Human3.6M .
Method
Size
MPJPE 400 / 1000
FPS
FADE 400 / 1000
Repeat last frame
-
201 / 408
-
201 / 408
Last delta average
-
167 / 398
-
167 / 398
Ridge Regression
-
130 / 296
-
130 / 296
JRTransformer
3.79M
99.0 / 245
610
99.4 / 245
DeformMLP
1.29M
98.3 / 242
462
98.8 / 243
EqMotion
3.37M
90.6 / 217
41
96.1 / 222
Table 4 : Comparison of single person forecasting on CMU-MoCap dataset. The models were trained to predict a maximum time-range of 1000ms , the forecast error was measured at the timesteps 400ms and 1000ms .
Method
Size
MPJPE 1000/2000/3000
FPS
FADE 1000/2000/3000
Repeat last frame
-
356 / 562 / 706
-
356 / 562 / 706
Last delta average
-
373 / 773 / 1198
-
373 / 773 / 1198
Ridge Regression
-
276 / 507 / 692
-
276 / 507 / 692
DeformMLP
1.29M
241 / 454 / 656
226
242 / 455 / 657
JRTransformer
4.38M
255 / 446 / 602
562
255 / 446 / 602
EMPMP
3.91M
233 / 385 / 489
201
234 / 386 / 490
Table 5 : Comparison of single person forecasting on CMU-MoCap dataset. The models were trained to predict a maximum time-range of 3000ms , the forecast error was measured at 1000ms , 2000ms and 3000ms . Again, twice the output duration was used as input duration. The reason for the difference of repeat last frame method at timestep 1000ms to the table before is the windowing of the dataset. Here, the window is much larger, resulting in fewer samples, because the windows start later and end earlier.
Method
Size
MPJPE
FPS
FCE
FADE
Repeat last frame
-
272
-
-
272
Last delta average
-
318
-
-
318
Ridge Regression
-
277
-
-
277
EqMotion
3.37M
234
16
125
249
EMPMP
314K
227
209
10
228
MotionConformer
10.2M
207
853
2
207
Table 6 : Forecasting close two-person interactions on CHi3D , errors at timestep 1000ms .
Method
MPJPE
40ms
400ms
1000ms
EMPMP
4.3
59.7
165
EqMotion
4.0
52.0
156
MotionConformer
4.0
51.6
149
Table 7 : Zero-shot transfer capabilities from a large mixture of datasets to Human3.6M , at the given timesteps.
Method
detector target
ground-truth target
400ms
1000ms
400ms
1000ms
EqMotion
513
842
532
856
EMPMP
100
215
123
229
MotionConformer
94.2
214
118
228
Table 8 : Influence of skeleton joints generated by a 3D-pose-estimation network on model performance, measured as MPJPE in millimeters at the given timesteps. The left section uses the network predictions as inputs and targets (which would be the measurable error in live systems), and the right one the predictions for input and ground-truth labels as targets (which would be the real error).
Method
detector target
ground-truth target
400ms
1000ms
400ms
1000ms
MotionConformer
66.9
159
94.8
177
EMPMP
76.4
179
104
195
EqMotion
68.5
168
96.1
185
MotionConformer
69.1
165
98.3
183
Table 9 : Improvements through pseudo-label training. Finetuning the pretrained model on noisy predictions in the upper part, training all models from scratch in the lower part. Columns are separated the same as in Table 8
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2 : Example of a relative pose forecast of a walking person. The green skeletons visualize the input sequence of the prediction model, and the red ones the predicted future poses. Only the predicted joints are visualized, therefore this person does not have a hip (it is fixed to the same spot).
Method
MPJPE (400/1000ms)
Replicated
Repeat last frame
88.2 / 136.6
-
Ridge Regression
76.5 / 128.2
-
Res. Sup. [ 24 , 38 ] ’2017
88.3 / 136.6
-
IAFormer [ 39 ] ’2024
50.8 / 130.4
-
MotionMixer [ 5 ] † ’2022
59.3 / 111.0
65.4 / 117.9
STSGCN [ 31 ] † ’2021
38.3 / 75.6
67.5 / 117.0
Appendix
Table 10 : Comparing recent approaches for relative pose forecasting on Human3.6M dataset, averaged over all actions. The mean per joint position error (MPJPE) is measured in millimeters at two frames 400 ms and 1000 ms in the future (using the same model for both timesteps). Approaches marked with a † had errors in their evaluation code which were fixed in the replicated column.
Method
Reporting Paper
MPJPE (1000ms / 3000ms)
HisRepItself
MRT [ 36 ]
1.05 / 1.58
TBiFormer [ 27 ]
134 / 349
SoMoFormer [ 34 ]
0.50 / 1.42
JRTransformer [ 42 ]
14.9 / 30.7
MRT
MRT [ 36 ]
0.79 / 1.22
TBiFormer [ 27 ]
148 / 352
Appendix
Table 11 : Comparing reported results of recent approaches for absolute pose forecasting on the CMU-MoCap dataset proposed by MRT [ 36 ] . Even though the descriptions of the experiments are similar, the reported results differ substantially between papers, showing general replicability problems.
Figure 3 : Measured MPJPE over time for MotionConformer , together with local one-step extrapolation segments used to validate the FADE approximation. Each segment starts at a measured timestep and follows the linear increase assumed by FADE . The small vertical offset between the segment endpoint and the next measured MPJPE value indicates the local approximation error.
Method
Size
MPJPE
FPS
FADE
Conformer
3.22M
147
1049
147
increased model size
8.41M
144
932
144
moved time reduction
3.59M
146
965
146
MotionConformer
9.23M
143
929
143
no spec augmentation
144
no convolutional modules
7.89M
145
1024
145
Appendix
Table 12 : Implemented improvements and ablations of MotionConformer , tested on Human3.6M .
Method
detector target
ground-truth target
400ms
1000ms
400ms
1000ms
EMPMP
82.5
188
108
204
MotionConformer
82.4
185
106
199
Appendix
Table 13 : Improvements through artificial noise pre-training. The left/right columns are separated into prediction/groundtruth targets, as in Table 8 , showing the measurable/real error.
Action-Class
MPJPE
MPJPE
above
MPJPE
Difference
Joint-Type
(clean)
(input)
10cm %
(noisy)
(noisy-clean)
Directions
140
54
5
156
16
Discussion
174
53
5
196
22
Eating
90
52
6
103
13
Greeting
235
96
37
262
27
Phoning
101
57
8
124
23
Appendix
Table 14 : Errors of MotionConformer per action and joint type on the Human3.6M dataset using ground-truth or detector-noisy joints.
Figure 4 : Example of improvements through finetuning on Human3.6M dataset. In blue the ground-truth, in orange the prediction. The walking movement was already well continued in (a), but especially the stride distances improved after finetuning. For better visualization some intermediate timesteps are not displayed.
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
Dairu Liu, Zekun Qi, Jiayu Zeng +11
1Nankai University · 3Galbot · 2Tsinghua University +3
This paper revisits camera pose estimation through the lens of self-supervised pretraining, focusing on inverse-dynamics pretraining as a scalable alternative to the current trend of fully supervised training with 3D annotations. Concretely, we employ inverse- and forward-dynamics models to learn latent action representations, similar to Genie from large-scale driving videos. Our idea is simple yet effective. Existing methods use latent actions in their original capacity, that is, as action conditioning of world-models or as proxies of robot action parameters in policy networks. Our method, dubbed LA-Pose, repurposes the latent action features as inputs to a camera pose estimator, finetuned on a limited set of high-quality 3D annotations. This formulation enables accurate and generalizable pose prediction while maintaining feed-forward efficiency. Extensive experiments on driving benchmarks show that LA-Pose achieves competitive and even superior performance to state-of-the-art methods while using orders of magnitude less labeled data. Concretely, on the Waymo and PandaSet benchmarks, LA-Pose achieves over 10% higher pose accuracy than recent feed-forward methods. To our knowledge, this work is the first to demonstrate the power of inverse-dynamics self-supervised learning for pose estimation.
Human motion describes the three-dimensional full-body movement of a person. Anticipating such motion holds significant relevance across a wide range of application domains such as human-robot interaction, autonomous driving, animation, and healthcare. In recent research, spatial and temporal dependencies are modeled by bidirectional attention mechanisms. These typically anticipate human motion in an autoregressive manner which could cause an accumulation of errors over time. As a consequence, they solely focus on local pose forecasting. To address these limitations, we propose a non-autoregressive transformer based on spatio-temporal attention, and train it not only for local pose anticipation, but also for global motion prediction in space. Furthermore, to enhance its applicability in real-world scenarios, our model is also trained to recover missing joints due to occlusions, and is capable of processing varying lengths of history observations. Our code is publicly available at https://github.com/Q-Y-Yang/Prediction-of-Local-and-Global-Human-Motion.
Qiaoyue Yang, Sven Heutger, Christopher Niemann +3