AgentGarten: Code Worlds for Evolving Agents
Abstract
Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments. Its simulation backends maintain persistent world state and execute program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a common interface. To build the neural renderer, we adapt a pretrained video model to geometry conditions, distill it with our proposed Adversarial Forcing, and optimize inference for real-time interaction. Adversarial Forcing makes history prefilling differentiable through exact replay, so that losses on later predictions update how the renderer encodes prior observations, and adds real-data adversarial supervision to improve its visual quality. In AgentGarten, agents perceive the world through visual observations, interact with it in real time, and improve by distilling each round of experience into playbooks that subsequent agents inherit and refine. Our empirical study demonstrates a substantial gain in learning efficiency, with agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart. As new worlds can be written as code and rendered through the same interface, environments can scale in both number and difficulty alongside their agents, a step toward agents that keep evolving through interactive experience.
Figures & tables
| \mirrosth Measure | \mirrosth SGF-style (FlexAttention) | \mirrosth Ours (block-by-block SDPA) | \mirrosth Change |
|---|---|---|---|
| Replay error (relative ) | 3.99% | 0 (bitwise) | — |
| \mirrosrowalt Forward time | 2.95 s | 2.79 s | |
| Peak allocated memory | 50.83 GiB | 50.92 GiB |
| \mirrosth One served block of 16 frames (one H100, BF16) | \mirrosth Time |
|---|---|
| Condition encoding | 10.6 ms |
| \mirrosrowalt Transformer: four denoising steps and cache publication | 414.4 ms |
| Tiny decoder and host transfer | 8.8 ms |
| \mirrosrowalt Total wall time | 438.8 ms |
| Throughput | 36.5 frames/s |
| \mirrosth Strategy | \mirrosth Self-play RL [ 13 ] | \mirrosth Pretrained visual agents |
|---|---|---|
| training episodes | rounds played | |
| Build shelters | million | 4 |
| \mirrosrowalt Use ramps to enter shelters | million | 10 |
| \mirrosth World | \mirrosth Measure | \mirrosth Round 1 | \mirrosth 2 | \mirrosth 3 | \mirrosth 4 |
|---|---|---|---|---|---|
| Companion dog | engagement score | 13 | 14 | 19 | 19 |
| \mirrosrowalt One-lane bridge | seconds until both cars arrive | 71 | 68 | 45 | 41 |
| Herding | score out of 100 | 60 | 90.1 | 87.6 | 88.3 |
| \mirrosrowalt | sheep penned, of four | 3 | 4 | 4 | 4 |
| Quarry loader | score out of 100 | 30 | 0 | 30 | 90.9 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| \mirrosth Stage | \mirrosth Updated parameters | \mirrosth Initialization | \mirrosth LR | \mirrosth Adam |
|---|---|---|---|---|
| 1. Geometry conditioning | self-attention projections and q/k norms; condition embedding | Cosmos 3-Nano | ||
| \mirrosrowalt 2. Causal adaptation | same as stage 1 | stage 1 | ||
| 3. Distillation: student | complete video stream | stage 2 | ||
| \mirrosrowalt fake score | complete video stream | stage 1 | ||
| discriminator | head only; backbone is the frozen teacher | fresh |