LogicEnvGen: Task-Logic Driven Generation of Diverse Simulated Environments for Embodied AI
Authors: Jianan Wang, Siyang Zhang, Bin Li, Juan Chen, Jingtao Qi, Zhuo Zhang, Chen Qian
Organizations: College of Computer Science and Technology, National University of Defense Technology · Intelligent Game and Decision Lab (IGDL), Beijing · School of Artifical Intelligence, Shanghai Jiao Tong University
Simulated environments play an essential role in embodied AI, functionally analogous to test cases in software engineering. However, existing environment generation methods often emphasize visual realism (e.g., object diversity and layout coherence), overlooking a crucial aspect: logical diversity from the testing perspective. This limits the comprehensive evaluation of embodied agent adaptability and planning robustness across distinct simulated environments. To bridge this gap, we propose LogicEnvGen, a novel method driven by Large Language Models (LLMs) that adopts a top-down paradigm to generate logically diverse simulated environments as test cases for agents. Given an agent task, LogicEnvGen first analyzes its execution logic to construct decision-tree-structured behavior plans and then synthesizes a set of logical trajectories. Subsequently, it adopts a heuristic algorithm to refine the trajectory set, reducing redundant simulation. For each logical trajectory, which represents a potential task situation, LogicEnvGen correspondingly instantiates a concrete simulated environment. Furthermore, we introduce LogicEnvEval, a novel benchmark for simulated environment generation, with four quantitative metrics. Experimental results verify the lack of logical diversity in baselines and demonstrate that LogicEnvGen achieves 1.08-2.67x greater diversity, significantly improving the performance in revealing agent faults by 3.34%-72.00%.
Figures & tables
Figure 1: Simulated environments generated by Holodeck Yang et al. (2024c) and LogicEnvGen. In contrast, LogicEnvGen generates more logically diverse test cases based on the task description, enabling more comprehensive simulation.
Figure 2: Overview. Phase ①: Behavior Plan Derivation, decomposes the given task into independent subtasks, identifies uncertain environment factors impacting subtask execution, and generates decision-tree-structured behavior plan for each subtask. Phase ②: Logical Trajectory Collection, traverses and combines decision paths across trees to synthesize distinct logical trajectories for the entire task, each representing a potential task situation. Phase ③: Simulated Environment Construction, instantiates a concrete and physically plausible environment for each situation through three stages: floor plan design, environment object selection and arrangement.
Figure 3: Algorithm example. Logical trajectories reduce from 12 to 4 (3 blue, 1 purple).
Figure 4: LogicEnvEval Benchmark. (a) Task composition. SN: number of subtasks. TN: number of task trajectories (minimal). (b) Environment type distribution of all subtasks. (c) Action step distribution of correct policies. (d) Action type diversity in correct policies.
Dimension
Physical Constraints
Floor Plan
Adjacent rooms must not overlap.
Doors and windows must be embedded in the walls, with doors placed on the floor and windows positioned above it.
Entity
Objects must not collide with one another and must remain within room boundaries.
Relation
The positions and directions of objects must satisfy the specified spatial relations.
Table 1: Three dimensions of Physics Pass Rate.
Model
Method
Physics Pass Rate (%)
LogCov
SceVR
Fault Detection Rate (%)
ST
Floor Plan
Entity
Relation
Avg
(%)
(%)
CFactuals
Unreach
LBranch
Avg
(min)
DeepSeek-v3.2
CoT
98.00
26.66
19.66
48.11
73.08
84.48
70.00
64.00
64.00
66.00
35.86
IFG
89.82
11.16
7.70
36.23
91.78
90.25
90.00
90.00
88.00
89.33
98.21
Holodeck
100.00
100.00
100.00
100.00
37.07
74.00
24.00
22.00
28.00
24.67
14.71
I-Design
-
86.00
18.00
52.00
37.07
64.00
20.00
18.00
24.00
20.67
13.93
LogicEnvGen
100.00
100.00
100.00
100.00
98.97
92.40
94.00
90.00
94.00
92.67
51.42
Table 2: The comparative experiment results on the LogicEnvEval benchmark.
Method
PhyPR
LogCov
SceVR
FauDR
ST
(%)
(%)
(%)
(%)
(min)
CoT
45.26
96.67
86.67
84.33
12.29
IFG
34.29
98.33
93.43
95.00
19.28
Holodeck
100.00
92.00
79.00
72.67
9.37
I-Design
49.50
92.00
66.00
60.33
8.74
LogicEnvGen
100.00
99.67
94.62
95.67
11.78
Table 3: The comparative experiment results on the EMMOE-100 benchmark.
Method
PhyPR(%)
LogCov(%)
SceVR(%)
FauDR(%)
JI
W/O DBP
100.00
93.25
89.97
90.67
0.33
LogicEnvGen
100.00
98.97
92.40
92.67
0.88
Table 4: Results of ablation study on the decision-tree-structured behavior plan (DBP).
Method
LogCov(%)
JI
TC
SceVR(%)
FauDR(%)
Exhaustive
98.97
0.92
O( 2N )
92.59
92.00
W/O MTSA
98.97
0.27
-
91.78
94.00
LogicEnvGen
98.97
0.88
O(N)
92.40
92.67
Table 5: Results of comparative and ablation study on the Minimal Trajectory Selection Algorithm (MTSA). TC denotes Time Complexity. N denotes the number of trajectories generated by the Cartesian product.
Figure 5: Results of ablation study on the constraint-based arrangement solving (CAS).
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Acc_TD(%)
Acc_UFI(%)
Acc_BPG(%)
DeepSeek-V3
96.00±0.00
94.56±0.36
95.04±0.36
Gemini-2.5-Flash
97.33±0.00
94.00±0.00
94.00±0.00
Qwen2.5-72B
89.47±0.73
87.07±1.01
87.07±1.01
Appendix
Table 6: Quantitative result of LLM stability and accuracy across the three steps of Behavior Plan Derivation. “Acc_TD”, “Acc_UFI”, and “Acc_BPG” denote the accuracy of Task Decomposition, Uncertain Factor Identification, and Behavior Plan Generation, respectively.
Task name
Clean Living Room
Task description
You are in the living room. There are a sofa and a table with a wet mop leaning against it. There may be a book or a toy on the floor. The room has two toy boxes - a red one for dolls and a white one for other toys. There is a dirty mark on the floor, and you may find a pack of wet wipes on the table. Your task: Please check the living room floor, if there are items on floor, put them where they belong (e.g., book should be placed on the sofa, toy should be placed in the toy box). Clean the mark, preferably with the wet wipes on the table (if available). If there aren’t any wipes, you can use the wet mop instead.
Decision-tree-structured behavior plan
[ { "There is a toy on the floor?": { "YES": { "What is the type of the toy on the floor?": { "doll": "Place the toy in the red box.", "other types": "Place the toy in the white box." }, "NO": "Do nothing." } }, { "There is a book on the floor?": { "YES": "Place the book on the sofa.", "NO": "Do nothing." } }, { "There is a wet wipe on the table?": { "YES": "Clean stain with the wet wipe.", "NO": "Clean stain with the wet mop." } } ]
Appendix
Table 7: Task description and decision-tree-structured behavior plan of the example Clean Living Room .
Figure 6: Correct policy (Behavior Tree) of the example Clean Living Room .
Figure 7: Faulty policy (Behavior Tree) of the example Clean Living Room , which belongs to the Counterfactuals category. Reason: The agent does not pick up the toy (non-doll) before placing it into the white toy box.
Figure 8: Faulty policy (Behavior Tree) of the example Clean Living Room , which belongs to the Unreachable category. Reason: The agent picks up the toy (doll) but does not place it into the red toy box.
Figure 9: Faulty policy (Behavior Tree) of the example Clean Living Room , which belongs to the Lackbranch category. Reason: The agent fails to plan for situations where wet wipes are absent.
Figure 10: Qualitative comparison on two LogicEnvEval tasks. While baseline methods suffer from physical implausibilities (e.g., floating objects), irrational layout or missing task-relevant items, LogicEnvGen generates environments with superior visual fidelity, layout coherence, and strict task alignment.
Figure 11: VLM evaluation of LogicEnvGen and four baselines. The results from both Gemini 3.1 Pro and InternVL3.5-38B consistently show LogicEnvGen outperforms four baselines across three criteria.
Figure 12: CLIP Score comparison over two scenario categories.
Figure 13: Comparative human evaluation of LogicEnvGen and four baselines across three criteria. The pie charts show the distribution of annotator preferences, showing both the percentage and the actual number of annotations favoring each method.
Method
PhyPR(%)
LogCov(%)
SceVR(%)
Fault Detection Rate (%)
CFactuals
Unreach
LBranch
Avg
Holodeck
100.00
37.07
74.00
24.00
22.00
28.00
24.67
Holodeck+Logic
100.00
87.92
90.26
76.00
84.00
78.00
79.33
LogicEnvGen
100.00
98.97
92.40
94.00
90.00
94.00
92.67
Appendix
Table 8: Attribution ablation results on the LogicEnvEval benchmark.
Checker: Valid
Checker: Invalid
Human Evaluator: Valid
89
0
Human Evaluator: Invalid
1
20
Appendix
Table 9: Confusion matrix of the validity checker against human evaluation.
Figure 14: Case 1. Agent task: Bring me a desk lamp from study if there is a pan in the kitchen. Otherwise, bring me a pillow from the bedroom.
Figure 15: Case 2. Agent task: The bedroom has a bed, possibly with a book or clock on it, along with a bookshelf and a desk. The bookshelf’s three levels hold novels, essays, and other books respectively. Clean items on the bed and put them into their designated storage areas (e.g., book on the shelf, phone on the desk).
Figure 16: Case 3. Agent task: The living room has a sofa, a tea table, and two toy boxes—the blue one for dolls and the white one for other toys. The floor is stained and there might be a sofa pillow or toy on it. Check the floor and put misplaced items (like a pillow or toy) in their proper places. Clean the stain, preferably with the wet wipes on the table (if available), otherwise use the wet mop by the wall.
Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, brittle to extend, and fundamentally limited in diversity. A promising direction is to replace manually crafted environments with LLM-simulated counterparts. However, this paradigm hinges on an unexamined core assumption: LLMs can accurately simulate environmental feedback. In practice, LLM-simulated environments suffer from hallucinations, logical inconsistencies, and silent state drift failures that corrupt agent reward signals and compound the construction costs that the paradigm was designed to eliminate. To address this gap, we propose EnvSimBench with four contributions: 1) We provide the first formal definition and operationalization of Environment Simulation Ability (EnvSim Ability) as a quantifiable research objective. 2) We construct EnvSimBench, a rigorous benchmark covering 400 samples across 167 diverse environments, equipped with verifiable labels and fine-grained difficulty stratification along three axes. 3) Systematic evaluations reveal that all state-of-the-art language models suffer from a universal state change cliff: they achieve near-perfect accuracy on tasks when the environment state remains invariant, yet fail catastrophically when multiple states need simultaneous updates. This finding exposes EnvSim Ability as a critical yet largely unaddressed capability gap. 4) We design a constraint-driven simulation pipeline that substantially reduces hallucination, boosts environment synthesis yield by 6.8%, and cuts costs by over 90%. Overall, EnvSimBench serves as both a diagnostic framework and a practical optimization path for reliable LLM-based environment simulation, establishing a foundation for scalable agent training. Code and data are available at https://github.com/cookieApril/EnvSimBench
Yi Liu, TingFeng Hui, Wei Zhang +4
Beijing University of Posts and Telecommunications · The Hong Kong University of Science and Technology · Chongqing University
LLM/VLM-based digital agents have advanced rapidly thanks to scalable sandboxes for coding, web navigation, and computer use, which provide rich interactive training grounds. In contrast, embodied agents still lack abundant, diverse, and automatically generated 3D environments for interactive learning. Existing embodied simulators rely on manually crafted scenes or procedural templates, while recent LLM-based 3D generation systems mainly produce static scenes rather than deployable environments with verifiable tasks and standard learning interfaces. We introduce SimWorld Studio, an open-source platform built on Unreal Engine 5 for generating evolving embodied learning environments. At its core is SimCoder, a tool/skill-augmented coding agent that writes and executes engine-level code to construct physically grounded 3D worlds from language/image instructions. SimCoder self-evolves by using verifier feedback (e.g., compilation errors, physics checks, VLM critiques) to revise environments and autonomously add reusable tools and skills to its library. Generated worlds are exported as Gym-style environments for embodied agent learning. SimWorld Studio further enables co-evolution between environment generation and embodied learning: agent performance feedback guides SimCoder to generate adaptive curricula near the learner's capability frontier, so that environments become increasingly challenging as the embodied agent improves. Three case studies on embodied navigation show that self-evolution improves generation reliability, generated environments substantially improve embodied agent performance that generalizes to unseen benchmarks, and co-evolution yields an 18-point success-rate gain over fixed-environment learning and a 40-point gain over an untrained agent.
We present EmbodiedGen V2, a generative 3D world engine for building executable policy-ready environments for embodied intelligence. Sim-ready 3D asset generation has advanced rapidly, yet assembling such assets into policy-ready task environments remains largely manual, limiting scalable closed-loop learning. EmbodiedGen V2 addresses this gap through a unified sim-ready representation that connects cross-simulator assets, interaction affordances, task-driven worlds, large-scale multi-room scenes, and stateful Vibe Coding into a generative, editable, and reusable simulation pipeline. The generated environments support manipulation, navigation, mobile manipulation, cross-simulator deployment, and embodied policy training. In evaluation, the asset pipeline achieves 96.5% human acceptance and 98.6% collision success, and 83.3% of task-driven worlds are directly usable for downstream simulation without manual modification. Online reinforcement learning with generated environments further improves simulation success from 9.7% to 79.8%, and transfers to real robots with task success increasing from 21.7% to 75.0%. These results establish EmbodiedGen V2 as scalable simulation infrastructure for training, evaluating, and deploying embodied policies.