cs.AISep 29, 2026

Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics

Authors: Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio

Organizations: NVIDIA

Abstract

Video games offer scalable environments for studying perception and control in embodied agents. Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on ∼\sim1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affect an IDM's outcome and we analyse our models on per-key and balanced metrics such as F1macroF_1^{macro}. Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics. The application of the same architecture and training recipe to Cyberpunk 2077 reveals uneven performance across game mechanics. Our per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement

    Apr 20, 2026Kerui Li, Zhe Jing, Xiaofeng Wang +6Inverse DynamicsEmbodied Artificial Intelligence

  2. KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding

    Jul 22, 2026Zeyu Liu, Zhangzhe Zhu, Yang Zhang +3Articulated KinematicsWorld Models

  3. World Models as Group Actions

    May 23, 2026Zijie Wang, Wei Zhang, Weiming Zhang +4Efficient World-Action ModelVideo World Models