Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments. Its simulation backends maintain persistent world state and execute program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a common interface. To build the neural renderer, we adapt a pretrained video model to geometry conditions, distill it with our proposed Adversarial Forcing, and optimize inference for real-time interaction. Adversarial Forcing makes history prefilling differentiable through exact replay, so that losses on later predictions update how the renderer encodes prior observations, and adds real-data adversarial supervision to improve its visual quality. In AgentGarten, agents perceive the world through visual observations, interact with it in real time, and improve by distilling each round of experience into playbooks that subsequent agents inherit and refine. Our empirical study demonstrates a substantial gain in learning efficiency, with agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart. As new worlds can be written as code and rendered through the same interface, environments can scale in both number and difficulty alongside their agents, a step toward agents that keep evolving through interactive experience.
Figures & tables
Figure 1 : \captionleadfont Interaction and rendering in a code world (Eqs. 1 and 2 ). The agent submits actions and receives generated observations. The engine maintains scene state and exports structured conditions, here surface normals; the neural renderer combines them with an appearance reference, text, and cached visual history. New observations return to the agent and join the history.
Figure 2 : \captionleadfont A coding agent builds a code world from a single image. A living-room example. The agent calls perception and generation models as tools, writes a scene program from their outputs, and repairs it after inspecting rendered views and rollout checks. The expanded scene keeps the observed room and adds connected rooms beyond the input view as authored extensions. The input image supplies the appearance reference x0 .
Figure 3 : \captionleadfont Block-causal teacher-forcing mask, shown for two target blocks. Clean copies Pj encode history and see the reference, earlier clean blocks, and themselves; noisy copies Qj see the reference, clean blocks before j , and themselves, but never their own clean target. Text is visible to every row. The same layout is reused, with recorded latents, by the replay pass of Adversarial Forcing.
Figure 4 : \captionleadfont Exact replay, drawn as block-level attention masks. Row i holds block i ’s queries, and the columns are the key/value blocks it reads. The rollout runs one attention call per block and reads earlier blocks from a detached cache. Our replay runs the same calls but computes the history keys and values with gradients, preserving the rollout’s execution structure. An SGF-style replay computes the same mask in one full-sequence call, which changes numerical execution. Table 1 evaluates the difference. In the implementation each block issues two calls, one for its noisy queries and one for the clean publication that later blocks read.
Figure 5 : \captionleadfont Exact R1/R2 with a frozen backbone B and a trainable head hϕ . (1) Differentiate with respect to x only to obtain g and R . (2) Evaluate the backbone JVP to obtain features z and directions v . (3) Evaluate the explicit head JVP on detached (z,v) , returning both the logit and its directional derivative, which supplies s . An ordinary backward pass with respect to ϕ updates the head using the full discriminator objective.
Figure 6 : \captionleadfont The inference cache. It retains a sink prefix, which begins with the reference, and recent frames, with configurable capacities; older frames are evicted. The current block attends to every cached frame and to the text. After denoising, the clean block joins the recent window.
Figure 7 : \captionleadfont Qualitative comparison with Self Forcing. Each pair shows the same 30-second condition trajectory at four moments.
\mirrosth Measure
\mirrosth SGF-style (FlexAttention)
\mirrosth Ours (block-by-block SDPA)
\mirrosth Change
Replay error (relative L2 )
3.99%
0 (bitwise)
—
\mirrosrowalt Forward time
2.95 s
2.79 s
−5.4%
Peak allocated memory
50.83 GiB
50.92 GiB
+0.2%
Table 1: \captionleadfont Replay accuracy and cost. Relative L2 error between rollout and replay, measured on one H100 with 36 layers, 480×832 resolution, 61 latent frames, and BF16. Times are medians of three warmed-up runs of the replay forward pass and exclude backward and optimizer updates.
\mirrosth One served block of 16 frames (one H100, BF16)
\mirrosth Time
Condition encoding
10.6 ms
\mirrosrowalt Transformer: four denoising steps and cache publication
414.4 ms
Tiny decoder and host transfer
8.8 ms
\mirrosrowalt Total wall time
438.8 ms
Throughput
36.5 frames/s
Table 2: \captionleadfont Inference cost at steady state. Measured with five sink and 44 recent latent frames in the cache and four current latent frames, using the hand-written kernels, CUDA graph capture, and the tiny decoder. Stage times are median GPU times; the total is the mean wall time per block and also covers work outside the three stages.
Figure 8 : \captionleadfont Five code worlds, each at five moments of one rollout. In every column, the untextured scene geometry (top) conditions the generated observation (bottom). Racing frames contain first-person views from two agents sharing one world state. The scene programs specify coarse geometry; the renderer supplies appearance from the initial image, text, and visual history.
Figure 9 : \captionleadfont Hide-and-seek in two settings. Top: self-play reinforcement learning trains a policy network on object state and rewards. Our pretrained agents see first-person frames rendered by the neural renderer, act through short Python programs, and keep what they learn in a playbook of skill files. Bottom: three games from the run, seen from an overhead camera.
\mirrosth Strategy
\mirrosth Self-play RL [ 13 ]
\mirrosth Pretrained visual agents
training episodes
rounds played
Build shelters
≈25 million
4
\mirrosrowalt Use ramps to enter shelters
≈100 million
10
Table 3: \captionleadfont Milestones of physical tool use in hide-and-seek under two distinct paradigms. Self-play RL [ 13 ] trains tabula rasa policies over privileged object state across millions of episodes. In contrast, our study evaluates how pretrained agents, perceiving exclusively through the real-time neural renderer, ground general commonsense priors into closed-loop physical execution and adapt their strategies through written playbooks within a handful of games.
Figure 10 : \captionleadfont One round of practice. Agents read the task file, which states the goal, the actions, and the limits but no solution; play in parallel within a fixed budget, seeing only camera frames; and write a playbook. Playbooks are archived, and the next round’s agents start in fresh conversations from the task file and every earlier playbook.
Figure 11 : \captionleadfont The same loop in four further worlds. A frame from round 1 (top) and round 4 (bottom) of each run, with the world’s own measure. Table 4 lists every round.
\mirrosth World
\mirrosth Measure
\mirrosth Round 1
\mirrosth 2
\mirrosth 3
\mirrosth 4
Companion dog
engagement score
13
14
19
19
\mirrosrowalt One-lane bridge
seconds until both cars arrive
71
68
45
41
Herding
score out of 100
60
90.1
87.6
88.3
\mirrosrowalt
sheep penned, of four
3
4
4
4
Quarry loader
score out of 100
30
0
30
90.9
Table 4: \captionleadfont Outcome of every round in the four additional worlds. Each row is the world’s own terminal measure, one episode per round. For the bridge, lower is better.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
\mirrosth Stage
\mirrosth Updated parameters
\mirrosth Initialization
\mirrosth LR
\mirrosth Adam (β1,β2)
1. Geometry conditioning
self-attention projections and q/k norms; condition embedding
Cosmos 3-Nano
3×10−5
(0.9,0.99)
\mirrosrowalt 2. Causal adaptation
same as stage 1
stage 1
3×10−5
(0.9,0.99)
3. Distillation: student
complete video stream
stage 2
2×10−6
(0,0.999)
\mirrosrowalt fake score
complete video stream
stage 1
4×10−7
(0,0.999)
discriminator
head only; backbone is the frozen teacher
fresh
2×10−7
(0,0.999)
Appendix
Table 5: Training stages. Stages 1 and 2 train first on five-second and then on fifteen-second windows; each distillation run uses the teacher and fake score of the matching window length. The teacher is the stage-1 model and stays frozen. The text stream and the video autoencoder are frozen throughout.
Department of Mechanical & Aerospace Engineering, University of Wisconsin-Madison · School of Computer, Data, and Information Sciences, University of Wisconsin-Madison