Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise
Organizations: Tsinghua University · Carnegie Mellon University · Genesis AI · Shanghai Qi Zhi Institute
Abstract
Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic pipeline that converts a text-only request into an executable, editable dynamic case. Built on Genesis, Text2Sim uses a hierarchical agentic structure that combines a Planner with specialized Writers, asset-generation tools, and an independent Critic. Compact skills (Debug Cards) distilled from graphics demonstrations provide role-specific physical guidance for execution-based repair. We evaluate physical quality, visual quality, and human preference on 42 held-out prompts spanning rigid, articulated, deformable, and cloth phenomena, with a paper-level split between experience construction and evaluation. We design automatic physical and visual scorers to evaluate the quality of the results, and Text2Sim achieves higher scores than all four state-of-the-art baselines on both metrics. In blinded user studies with these baselines, significantly more participants prefer Text2Sim than prefer the baselines, which is consistent with the results from our automatic scorers. The pipeline also supports a broad range of downstream applications; we select dataset construction and extension to multimodal input as two representative examples. We will release the code, the Debug Card library, and a dataset of generated cases, each pairing the text prompt and rendered video with the executable program, assets, physical parameters, controls, and recorded states.
Figures & tables
| Configuration | Scene | Body | Action | Render | Overall |
|---|---|---|---|---|---|
| Text2Sim | |||||
| GPT-5.6-Sol | |||||
| Qwen-3.8-Max | |||||
| Claude-Opus-5 | |||||
| Code2Worlds |
| Configuration | Content | Clarity | Polish | Appeal | Overall |
|---|---|---|---|---|---|
| Text2Sim | |||||
| GPT-5.6-Sol | |||||
| Qwen-3.8-Max | |||||
| Claude-Opus-5 | |||||
| Code2Worlds |
| Method | Task faithfulness | Physical plausibility | Visual quality | Overall preference |
|---|---|---|---|---|
| Text2Sim | ||||
| GPT-5.6-Sol | 50 (anchor) | 50 (anchor) | 50 (anchor) | 50 (anchor) |
| Qwen-3.8-Max | ||||
| Claude-Opus-5 | ||||
| Code2Worlds |
| Setting | Value |
|---|---|
| Simulator | Genesis 0.4.5 |
| Renderer | LuisaRender, commit d176fd90 |
| Runtime | Python 3.12.13; PyTorch 2.11.0; CUDA 12.8 |
| GPUs | NVIDIA GeForce RTX 4090 D, 24 GB each |
| CPUs | AMD EPYC 7H12 (128 cores total) |
| RAM | 512 GiB |
| Method | Tokens (M) | Time (min) | Time excl. GPU simulation (min) |
|---|---|---|---|
| Text2Sim | 29.1 | 101.4 | 58.0 |
| GPT-5.6-Sol | 24.8 | 60.9 | 41.9 |
| Qwen-3.8-Max | 23.9 | 95.1 | 55.8 |
| Claude-Opus-5 | 28.7 | 58.7 | 40.6 |
| Code2Worlds | 27.26 | 113.0 | 59.8 |