World-model agents are usually evaluated in simulators that can wait for the policy; live games impose the opposite constraint, requiring capture, prediction, and action before the next frame. We present DashVMC, which learns a compact, action-conditioned world model from approximately two hours of recorded Geometry Dash gameplay. To test whether the learned dynamics are actionable, a controller is initialized by behavioural cloning (BC) and refined with Proximal Policy Optimization (PPO) entirely in frozen-model rollouts, without further interaction with the live game. Across three controller seeds, the refined policies survive longer than their BC initializations on all three official levels and a held-out community layout. At deployment, the baseline skips visual generation and sustains a 60-Hz capture-to-action loop on a consumer GPU. Action-conditioned continuations and rollout diagnostics show that the model remains useful for control despite imperfect long-horizon fidelity.
Figures & tables
Figure 1 : Learning in imagination and acting in the live game. (a) A recorded context seeds the frozen world model; the action model chooses what to do, the world model predicts the consequence, and PPO updates only the action model. (b) In the baseline live path, the encoder and temporal context provide the state used for action selection, while visual generation is skipped.
AdamW; 200 × 500 steps; batch 512; LR 2×10−3→5×10−5 ; weight decay 0.01; dropout 0.1; death oversampling 5× .
Max. death F1 (epoch 139).
BC
Three seeds; AdamW; 50 epochs/seed; batch 512; LR 10−3 ; weight decay 10−4 ; positive jump weight 1.5.
Min. loss (10; 10; 11).
PPO
Three seeds; Adam; 15,000 iterations/seed; 512 rollouts/iteration; four update epochs; minibatch 512; LR 10−4 ; entropy/critic weights 0.01/0.5; ratio clip 0.2; max grad. norm 0.5.
Max. survival (12,420; 5,090; 12,280).
Table 1 : Training and checkpoint selection; tokenizer/dynamics learning rates use cosine decay.
Level
Policy
Seed 43
Seed 44
Seed 45
Stereo Madness
No-op
shared: 46.2±0.7
BC
120.2±91.8
130.6±76.0
156.0±100.2
PPO
280.1±23.3
280.4±33.5
308.8±68.6
Δ [95% CI]
159.9[122.7,195.4]
149.8[116.8,180.6]
152.8[106.7,199.8]
Back on Track
No-op
shared: 63.4±0.7
BC
123.5±81.5
111.1±60.3
120.2±69.6
Table 2 : Live survival on all five layouts. No-op entries are mean frames ± standard deviation over 10 attempts and are shared across seeds. BC/PPO entries use 25 attempts; Δ is PPO minus its exact parent BC [95% attempt-level bootstrap interval]. The first three are official levels, Stereo Madness Copy begins in ship form, and Stereo INSANE Nerfed is held out.
Controller input
Dream
Stereo
Back
Polargeist
INSANE Nerfed
zt+ht
30.13
280.1±23.3
260.0±55.5
65.2±33.4
290.4±48.6
zt only
29.02
277.9±19.2
294.9±64.4
128.0±65.1
248.1±77.3
Table 3 : Seed-43 controller-input ablation (mean frames ± sample SD). Dream is best survival on the same 512 fixed contexts; live results use 25 attempts for the temporal-state reference and 10 for spatial-only.
Model
Run 1
Run 2
Mean
DashVMC
11.04
11.06
11.05
DIAMOND
50.13
48.46
49.30
IRIS
176.94
178.73
177.83
Table 4 : Batch-1 native imagined-transition latency on an RTX 2060 (synchronized run means in milliseconds). Interfaces differ, so this compares operational cost rather than quality.
Cohort
Horizon
Token acc. ↑
PSNR ↑
Token JS ↓
Standard
1
30.14%
32.83
0.0037
5
19.29%
28.25
0.0045
10
14.06%
25.23
0.0052
20
8.26%
22.45
0.0079
45
2.91%
19.99
0.0208
Extended
45
3.02%
19.78
0.0351
Table 5 : Recorded-action rollout diagnostics. Accuracy and decoder-space PSNR compare predicted with recorded grids; JS compares pooled token marginals. Standard: 1,024 starts from 153 trajectories; extended: 256 starts from four long trajectories.
Figure 2 : Action-conditioning diagnostic from one shared four-frame context. The rows are aligned at t+1 and t+5 : changing only the first future action makes the idle branch collide and terminate, while the jump branch clears the spike and lands by t+13 .
Variant
Death F1 ↑
Action adv. ↑
Mean acc. (%) ↑
Mean PSNR ↑
Mean JS ↓
Ours
.794
.165
14.93
25.75
.0084
No SLS
.808
.196
14.33
25.65
.0115
No CPC
.787
.148
13.89
25.56
.0104
Uniform LS
.826
.156
12.90
25.41
.0141
Table 6 : Matched dynamics-loss ablations on fixed development data. Ours uses structured SLS and CPC; other row names denote the changed component. Action advantage is the episode-mean factual-versus-flipped NLL difference. Rollout columns average horizons 1/5/10/20/45. Each row is one run; bold is best.
Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world models predict visual futures from supplied actions but do not serve as task policies. World-Action Models (WAMs) unify these objectives, but remain largely unexplored under the dynamics and open-ended interaction of video games. We introduce GameWAM, to our knowledge the first WAM for native closed-loop gameplay and GUI control. GameWAM jointly generates future visual observations and executable keyboard-mouse trajectories through parallel visual and action generative processes with block-causal conditioning and flow matching. To support joint world-action learning, we construct synchronized gameplay and GUI trajectories. To handle heterogeneous native control, GameWAM predicts a gameplay/GUI mode per action step and generates actions with mode-specific prediction distributions and continuous-action normalization. For long-horizon interaction, block-cycle control coordinates prediction, execution, and temporal context: it predicts beyond the committed horizon, executes short action blocks, replans from new observations, and hierarchically structures context from fine-grained within-cycle history to persistent cross-cycle history. Experiments demonstrate competitive task success with fewer executed native actions than the compared agents. We further uncover Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control. Project page is available at https://yunncheng.github.io/GameWAM/.
Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo +1
Fudan University · LIGHTSPEED · Independent Researcher +1
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.
Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decision, introducing latency that can limit real-time use. We introduce Code to Control, an approach that synthesizes Python controllers which execute directly as policies. Code to Control separates program structure from parameters. An LLM synthesizes the controller structure, while derivative-free search fits its parameters for continuous control using feedback from the environment. Once learned, the resulting controllers require neither LLM inference nor planning at decision time, enabling real-time gameplay and, under our timing protocol, faster action selection than a PPO policy. Across a suite of Atari games, Flappy Bird, and MuJoCo tasks, Code to Control outperforms planning-based program synthesis methods, remains competitive with deep reinforcement learning while using fewer environment interactions, transfers across substantial changes in environment dynamics, and scales to complex locomotion tasks.
Zergham Ahmed, Joshua B. Tenenbaum, Chris Bates +1
Harvard University · Massachusetts Institute of Technology · Florida Institute for Human and Machine Cognition