Video games offer scalable environments for studying perception and control in embodied agents. Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on ∼1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affect an IDM's outcome and we analyse our models on per-key and balanced metrics such as F1macro. Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics. The application of the same architecture and training recipe to Cyberpunk 2077 reveals uneven performance across game mechanics. Our per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.
Figures & tables
Figure 1 : Sample sequences and key-press distributions for Trackmania (a) and Cyberpunk 2077 (b). Distributions are on a log scale: key usage is highly imbalanced.
Game
Training sequences
Validation sequences
Testing sequences
Trackmania
30,000 ( ∼25 mins)
20,294 ( ∼17 mins)
27,984 ( ∼23 mins)
Cyberpunk 2077
72,192 ( ∼60 mins)
25,600 ( ∼21 mins)
24,576 ( ∼20 mins)
Table 1 : Number of overlapping 5-frame sequences per dataset. Each split’s sequences belong to different recordings that are disjoint.
Model
Configuration
Acc.
F1
Tout
Arch.
Res.
Flow
Initialization
Loss
Epochs
micro
macro
W
A
S
D
Sp
CNN-512
5
CNN
5122
N
none
soft-F1
100
.489
.499
.367
.902
.434
.075
.415
.008
ViT-E2E-512
5
ViT
5122
N
none
soft-F1
100
.887
.817
.637
.907
.698
.567
.755
.256
HT-RGB-E2E
1
HT
5122
N
none
soft-F1
50
.929
.880
.746
.921
.837
.648
.823
.502
HT-RGBF-E2E
1
HT
5122
Y
none
soft-F1
50
.929
.878
.791
.920
.829
.655
.822
.728
HT-RGBF-PT-F1
1
HT
5122
Y
contrastive
soft-F1
20+25
.938
.895
.793
.920
.869
.688
.870
.618
Table 2 : Accuracy Acc. and F1 metrics on the test datasets; 20+25 denotes 20 contrastive and 25 supervised epochs. All rows refer to Trackmania except the last one marked ∗ , which uses Cyberpunk 2077. Bold and underlined mark the highest and second-highest distinct Trackmania values in each metric column, respectively. For Cyberpunk 2077 we report for comparison only navigation keys that are in common with Trackmania, with Accuracy, F1macro , and F1micro over that set.
Figure 2 : Correctly predicted Trackmania sequences for HT-RGBF-PT-F1 (attention in the lower row). Pred matches GT for the supervised frame t . Better seen at 6x zoom.
Figure 3 : Incorrectly predicted Trackmania sequences for HT-RGBF-PT-F1 (attention in the lower row). Reverse-right ( S,D ) is predicted as forward-left ( W,A ), indicating lack of 3D understanding, while airborne braking and steering ( S,D ) is predicted as ( W,A ) as it is visually ambiguous. Better seen at 6x zoom.
Figure 4 : Correctly classified Cyberpunk 2077 sequences for CP-HT-RGBF-PT-F1 (attention in the lower row). Exact backward-motion match and partial right-movement match after being shot. Better seen at 6x zoom.
Figure 5 : Incorrectly classified Cyberpunk 2077 sequences for CP-HT-RGBF-PT-F1 (attention in the lower row). Both examples invert backward motion; in the second a false left movement is also predicted. Better seen at 6x zoom.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Experiment
Configuration
Purpose
CNN-128
1282 , 15-channel RGB, soft-F1
resolution baseline
CNN-512
5122 , otherwise CNN-128
resolution ablation
ViT-E2E-512
5122 , RGB, d/L/H=384/2/4 , soft-F1
ViT baseline
ViT-DINO-512
ViT-E2E with DINOv3 features
spatial-pretraining ablation
HT-RGB-E2E
5122 , 15-channel RGB, d/L/H=256/4/8 , soft-F1
HT baseline
HT-RGB-BCE
HT-RGB-E2E with BCE
loss ablation
Appendix
Table 3 : Experiment configurations for Tables 2 and 4 . E2E is end-to-end; d/L/H is transformer embedding dimension, number of blocks, and number of attention heads.
Model
Configuration
Acc.
F1
Tout
Arch.
Res.
Flow
Initialization
Loss
Epochs
micro
macro
W
A
S
D
Sp
CNN-128
5
CNN
1282
N
none
soft-F1
100
.469
.492
.361
.902
.430
.074
.392
.009
CNN-512
5
CNN
5122
N
none
soft-F1
100
.489
.499
.367
.902
.434
.075
.415
.008
ViT-E2E-512
5
ViT
5122
N
none
soft-F1
100
.887
.817
.637
.907
.698
.567
.755
.256
ViT-DINO-512
5
ViT
5122
N
DINOv3
soft-F1
100
.870
.790
.624
.898
.696
.452
.664
.411
HT-RGB-E2E
1
HT
5122
N
none
soft-F1
50
.929
.880
.746
.921
.837
.648
.823
.502
Appendix
Table 4 : Accuracy Acc. and F1 metrics on the Trackmania test split (for the secondary ablations reported in Table 3 ); 20+25 denotes 20 contrastive and 25 supervised epochs. Bold and underlined mark the highest and second-highest distinct values in each metric column, respectively.