Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments. Its simulation backends maintain persistent world state and execute program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a common interface. To build the neural renderer, we adapt a pretrained video model to geometry conditions, distill it with our proposed Adversarial Forcing, and optimize inference for real-time interaction. Adversarial Forcing makes history prefilling differentiable through exact replay, so that losses on later predictions update how the renderer encodes prior observations, and adds real-data adversarial supervision to improve its visual quality. In AgentGarten, agents perceive the world through visual observations, interact with it in real time, and improve by distilling each round of experience into playbooks that subsequent agents inherit and refine. Our empirical study demonstrates a substantial gain in learning efficiency, with agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart. As new worlds can be written as code and rendered through the same interface, environments can scale in both number and difficulty alongside their agents, a step toward agents that keep evolving through interactive experience.
Figures & tables
Figure 1 : \captionleadfont Interaction and rendering in a code world (Eqs. 1 and 2 ). The agent submits actions and receives generated observations. The engine maintains scene state and exports structured conditions, here surface normals; the neural renderer combines them with an appearance reference, text, and cached visual history. New observations return to the agent and join the history.
Figure 2 : \captionleadfont A coding agent builds a code world from a single image. A living-room example. The agent calls perception and generation models as tools, writes a scene program from their outputs, and repairs it after inspecting rendered views and rollout checks. The expanded scene keeps the observed room and adds connected rooms beyond the input view as authored extensions. The input image supplies the appearance reference x0 .
Figure 3 : \captionleadfont Block-causal teacher-forcing mask, shown for two target blocks. Clean copies Pj encode history and see the reference, earlier clean blocks, and themselves; noisy copies Qj see the reference, clean blocks before j , and themselves, but never their own clean target. Text is visible to every row. The same layout is reused, with recorded latents, by the replay pass of Adversarial Forcing.
Figure 4 : \captionleadfont Exact replay, drawn as block-level attention masks. Row i holds block i ’s queries, and the columns are the key/value blocks it reads. The rollout runs one attention call per block and reads earlier blocks from a detached cache. Our replay runs the same calls but computes the history keys and values with gradients, preserving the rollout’s execution structure. An SGF-style replay computes the same mask in one full-sequence call, which changes numerical execution. Table 1 evaluates the difference. In the implementation each block issues two calls, one for its noisy queries and one for the clean publication that later blocks read.
Figure 5 : \captionleadfont Exact R1/R2 with a frozen backbone B and a trainable head hϕ . (1) Differentiate with respect to x only to obtain g and R . (2) Evaluate the backbone JVP to obtain features z and directions v . (3) Evaluate the explicit head JVP on detached (z,v) , returning both the logit and its directional derivative, which supplies s . An ordinary backward pass with respect to ϕ updates the head using the full discriminator objective.
Figure 6 : \captionleadfont The inference cache. It retains a sink prefix, which begins with the reference, and recent frames, with configurable capacities; older frames are evicted. The current block attends to every cached frame and to the text. After denoising, the clean block joins the recent window.
Figure 7 : \captionleadfont Qualitative comparison with Self Forcing. Each pair shows the same 30-second condition trajectory at four moments.
\mirrosth Measure
\mirrosth SGF-style (FlexAttention)
\mirrosth Ours (block-by-block SDPA)
\mirrosth Change
Replay error (relative L2 )
3.99%
0 (bitwise)
—
\mirrosrowalt Forward time
2.95 s
2.79 s
−5.4%
Peak allocated memory
50.83 GiB
50.92 GiB
+0.2%
Table 1: \captionleadfont Replay accuracy and cost. Relative L2 error between rollout and replay, measured on one H100 with 36 layers, 480×832 resolution, 61 latent frames, and BF16. Times are medians of three warmed-up runs of the replay forward pass and exclude backward and optimizer updates.
\mirrosth One served block of 16 frames (one H100, BF16)
\mirrosth Time
Condition encoding
10.6 ms
\mirrosrowalt Transformer: four denoising steps and cache publication
414.4 ms
Tiny decoder and host transfer
8.8 ms
\mirrosrowalt Total wall time
438.8 ms
Throughput
36.5 frames/s
Table 2: \captionleadfont Inference cost at steady state. Measured with five sink and 44 recent latent frames in the cache and four current latent frames, using the hand-written kernels, CUDA graph capture, and the tiny decoder. Stage times are median GPU times; the total is the mean wall time per block and also covers work outside the three stages.
Figure 8 : \captionleadfont Five code worlds, each at five moments of one rollout. In every column, the untextured scene geometry (top) conditions the generated observation (bottom). Racing frames contain first-person views from two agents sharing one world state. The scene programs specify coarse geometry; the renderer supplies appearance from the initial image, text, and visual history.
Figure 9 : \captionleadfont Hide-and-seek in two settings. Top: self-play reinforcement learning trains a policy network on object state and rewards. Our pretrained agents see first-person frames rendered by the neural renderer, act through short Python programs, and keep what they learn in a playbook of skill files. Bottom: three games from the run, seen from an overhead camera.
\mirrosth Strategy
\mirrosth Self-play RL [ 13 ]
\mirrosth Pretrained visual agents
training episodes
rounds played
Build shelters
≈25 million
4
\mirrosrowalt Use ramps to enter shelters
≈100 million
10
Table 3: \captionleadfont Milestones of physical tool use in hide-and-seek under two distinct paradigms. Self-play RL [ 13 ] trains tabula rasa policies over privileged object state across millions of episodes. In contrast, our study evaluates how pretrained agents, perceiving exclusively through the real-time neural renderer, ground general commonsense priors into closed-loop physical execution and adapt their strategies through written playbooks within a handful of games.
Figure 10 : \captionleadfont One round of practice. Agents read the task file, which states the goal, the actions, and the limits but no solution; play in parallel within a fixed budget, seeing only camera frames; and write a playbook. Playbooks are archived, and the next round’s agents start in fresh conversations from the task file and every earlier playbook.
Figure 11 : \captionleadfont The same loop in four further worlds. A frame from round 1 (top) and round 4 (bottom) of each run, with the world’s own measure. Table 4 lists every round.
\mirrosth World
\mirrosth Measure
\mirrosth Round 1
\mirrosth 2
\mirrosth 3
\mirrosth 4
Companion dog
engagement score
13
14
19
19
\mirrosrowalt One-lane bridge
seconds until both cars arrive
71
68
45
41
Herding
score out of 100
60
90.1
87.6
88.3
\mirrosrowalt
sheep penned, of four
3
4
4
4
Quarry loader
score out of 100
30
0
30
90.9
Table 4: \captionleadfont Outcome of every round in the four additional worlds. Each row is the world’s own terminal measure, one episode per round. For the bridge, lower is better.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
\mirrosth Stage
\mirrosth Updated parameters
\mirrosth Initialization
\mirrosth LR
\mirrosth Adam (β1,β2)
1. Geometry conditioning
self-attention projections and q/k norms; condition embedding
Cosmos 3-Nano
3×10−5
(0.9,0.99)
\mirrosrowalt 2. Causal adaptation
same as stage 1
stage 1
3×10−5
(0.9,0.99)
3. Distillation: student
complete video stream
stage 2
2×10−6
(0,0.999)
\mirrosrowalt fake score
complete video stream
stage 1
4×10−7
(0,0.999)
discriminator
head only; backbone is the frozen teacher
fresh
2×10−7
(0,0.999)
Appendix
Table 5: Training stages. Stages 1 and 2 train first on five-second and then on fifteen-second windows; each distillation run uses the teacher and fake score of the matching window length. The teacher is the stage-1 model and stays frozen. The text stream and the video autoencoder are frozen throughout.
LLM/VLM-based digital agents have advanced rapidly thanks to scalable sandboxes for coding, web navigation, and computer use, which provide rich interactive training grounds. In contrast, embodied agents still lack abundant, diverse, and automatically generated 3D environments for interactive learning. Existing embodied simulators rely on manually crafted scenes or procedural templates, while recent LLM-based 3D generation systems mainly produce static scenes rather than deployable environments with verifiable tasks and standard learning interfaces. We introduce SimWorld Studio, an open-source platform built on Unreal Engine 5 for generating evolving embodied learning environments. At its core is SimCoder, a tool/skill-augmented coding agent that writes and executes engine-level code to construct physically grounded 3D worlds from language/image instructions. SimCoder self-evolves by using verifier feedback (e.g., compilation errors, physics checks, VLM critiques) to revise environments and autonomously add reusable tools and skills to its library. Generated worlds are exported as Gym-style environments for embodied agent learning. SimWorld Studio further enables co-evolution between environment generation and embodied learning: agent performance feedback guides SimCoder to generate adaptive curricula near the learner's capability frontier, so that environments become increasingly challenging as the embodied agent improves. Three case studies on embodied navigation show that self-evolution improves generation reliability, generated environments substantially improve embodied agent performance that generalizes to unseen benchmarks, and co-evolution yields an 18-point success-rate gain over fixed-environment learning and a 40-point gain over an untrained agent.
Large language models are increasingly expected to serve as general-purpose agents that interact with external, stateful tool environments. The Model Context Protocol (MCP) and broader agent skills offer a unified interface for connecting agents with scalable real-world services, but training robust agents remains limited by the lack of realistic environments and principled mechanisms for life-long learning. In this paper, we present \textbf{Agent-World}, a self-evolving training arena for advancing general agent intelligence through scalable environments. Agent-World has two main components: (1) Agentic Environment-Task Discovery, which autonomously explores topic-aligned databases and executable tool ecosystems from thousands of real-world environment themes and synthesizes verifiable tasks with controllable difficulty; and (2) Continuous Self-Evolving Agent Training, which combines multi-environment reinforcement learning with a self-evolving agent arena that automatically identifies capability gaps through dynamic task synthesis and drives targeted learning, enabling the co-evolution of agent policies and environments. Across 23 challenging agent benchmarks, Agent-World-8B and 14B consistently outperforms strong proprietary models and environment scaling baselines. Further analyses reveal scaling trends in relation to environment diversity and self-evolution rounds, offering insights for building general agent intelligence.
World models have emerged as a powerful paradigm for building interactive simulation environments, with recent video-based approaches demonstrating impressive progress in generating visually plausible dynamics. However, because these models typically infer dynamics from video and represent them in latent states, they do not explicitly enforce physical constraints. As a result, the generated video rollouts are not physically plausible, exhibiting unstable contacts, distorted shapes, or inconsistent motion. In this paper, we present an agentic framework constructing physics-based world models through executable simulation code. The framework coordinates planning, code generation, visual review, and physics analysis agents. The planning agent converts the natural language prompt into a structured scene plan, the code agent implements it as executable simulation code, and the visual review agent provide visual feedback while the physics analysis agent checks physical consistency. The code is iteratively revised based on the feedback until the simulation matches the prompt reqirements and physical constraints. Experimental results show that our framework outperforms advanced video-based models in physical accuracy, instruction fidelity and visual quality, which could be applied to various scenarios including driving simulation and embodied robot tasks.
Hongyu Wang, Jingquan Wang, Bocheng Zou +2
Department of Mechanical & Aerospace Engineering, University of Wisconsin-Madison · School of Computer, Data, and Information Sciences, University of Wisconsin-Madison