LogicEnvGen: Task-Logic Driven Generation of Diverse Simulated Environments for Embodied AI
Authors: Jianan Wang, Siyang Zhang, Bin Li, Juan Chen, Jingtao Qi, Zhuo Zhang, Chen Qian
Organizations: College of Computer Science and Technology, National University of Defense Technology · Intelligent Game and Decision Lab (IGDL), Beijing · School of Artifical Intelligence, Shanghai Jiao Tong University
Simulated environments play an essential role in embodied AI, functionally analogous to test cases in software engineering. However, existing environment generation methods often emphasize visual realism (e.g., object diversity and layout coherence), overlooking a crucial aspect: logical diversity from the testing perspective. This limits the comprehensive evaluation of embodied agent adaptability and planning robustness across distinct simulated environments. To bridge this gap, we propose LogicEnvGen, a novel method driven by Large Language Models (LLMs) that adopts a top-down paradigm to generate logically diverse simulated environments as test cases for agents. Given an agent task, LogicEnvGen first analyzes its execution logic to construct decision-tree-structured behavior plans and then synthesizes a set of logical trajectories. Subsequently, it adopts a heuristic algorithm to refine the trajectory set, reducing redundant simulation. For each logical trajectory, which represents a potential task situation, LogicEnvGen correspondingly instantiates a concrete simulated environment. Furthermore, we introduce LogicEnvEval, a novel benchmark for simulated environment generation, with four quantitative metrics. Experimental results verify the lack of logical diversity in baselines and demonstrate that LogicEnvGen achieves 1.08-2.67x greater diversity, significantly improving the performance in revealing agent faults by 3.34%-72.00%.
Figures & tables
Figure 1: Simulated environments generated by Holodeck Yang et al. (2024c) and LogicEnvGen. In contrast, LogicEnvGen generates more logically diverse test cases based on the task description, enabling more comprehensive simulation.
Figure 2: Overview. Phase ①: Behavior Plan Derivation, decomposes the given task into independent subtasks, identifies uncertain environment factors impacting subtask execution, and generates decision-tree-structured behavior plan for each subtask. Phase ②: Logical Trajectory Collection, traverses and combines decision paths across trees to synthesize distinct logical trajectories for the entire task, each representing a potential task situation. Phase ③: Simulated Environment Construction, instantiates a concrete and physically plausible environment for each situation through three stages: floor plan design, environment object selection and arrangement.
Figure 3: Algorithm example. Logical trajectories reduce from 12 to 4 (3 blue, 1 purple).
Figure 4: LogicEnvEval Benchmark. (a) Task composition. SN: number of subtasks. TN: number of task trajectories (minimal). (b) Environment type distribution of all subtasks. (c) Action step distribution of correct policies. (d) Action type diversity in correct policies.
Dimension
Physical Constraints
Floor Plan
Adjacent rooms must not overlap.
Doors and windows must be embedded in the walls, with doors placed on the floor and windows positioned above it.
Entity
Objects must not collide with one another and must remain within room boundaries.
Relation
The positions and directions of objects must satisfy the specified spatial relations.
Table 1: Three dimensions of Physics Pass Rate.
Model
Method
Physics Pass Rate (%)
LogCov
SceVR
Fault Detection Rate (%)
ST
Floor Plan
Entity
Relation
Avg
(%)
(%)
CFactuals
Unreach
LBranch
Avg
(min)
DeepSeek-v3.2
CoT
98.00
26.66
19.66
48.11
73.08
84.48
70.00
64.00
64.00
66.00
35.86
IFG
89.82
11.16
7.70
36.23
91.78
90.25
90.00
90.00
88.00
89.33
98.21
Holodeck
100.00
100.00
100.00
100.00
37.07
74.00
24.00
22.00
28.00
24.67
14.71
I-Design
-
86.00
18.00
52.00
37.07
64.00
20.00
18.00
24.00
20.67
13.93
LogicEnvGen
100.00
100.00
100.00
100.00
98.97
92.40
94.00
90.00
94.00
92.67
51.42
Table 2: The comparative experiment results on the LogicEnvEval benchmark.
Method
PhyPR
LogCov
SceVR
FauDR
ST
(%)
(%)
(%)
(%)
(min)
CoT
45.26
96.67
86.67
84.33
12.29
IFG
34.29
98.33
93.43
95.00
19.28
Holodeck
100.00
92.00
79.00
72.67
9.37
I-Design
49.50
92.00
66.00
60.33
8.74
LogicEnvGen
100.00
99.67
94.62
95.67
11.78
Table 3: The comparative experiment results on the EMMOE-100 benchmark.
Method
PhyPR(%)
LogCov(%)
SceVR(%)
FauDR(%)
JI
W/O DBP
100.00
93.25
89.97
90.67
0.33
LogicEnvGen
100.00
98.97
92.40
92.67
0.88
Table 4: Results of ablation study on the decision-tree-structured behavior plan (DBP).
Method
LogCov(%)
JI
TC
SceVR(%)
FauDR(%)
Exhaustive
98.97
0.92
O( 2N )
92.59
92.00
W/O MTSA
98.97
0.27
-
91.78
94.00
LogicEnvGen
98.97
0.88
O(N)
92.40
92.67
Table 5: Results of comparative and ablation study on the Minimal Trajectory Selection Algorithm (MTSA). TC denotes Time Complexity. N denotes the number of trajectories generated by the Cartesian product.
Figure 5: Results of ablation study on the constraint-based arrangement solving (CAS).
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Acc_TD(%)
Acc_UFI(%)
Acc_BPG(%)
DeepSeek-V3
96.00±0.00
94.56±0.36
95.04±0.36
Gemini-2.5-Flash
97.33±0.00
94.00±0.00
94.00±0.00
Qwen2.5-72B
89.47±0.73
87.07±1.01
87.07±1.01
Appendix
Table 6: Quantitative result of LLM stability and accuracy across the three steps of Behavior Plan Derivation. “Acc_TD”, “Acc_UFI”, and “Acc_BPG” denote the accuracy of Task Decomposition, Uncertain Factor Identification, and Behavior Plan Generation, respectively.
Task name
Clean Living Room
Task description
You are in the living room. There are a sofa and a table with a wet mop leaning against it. There may be a book or a toy on the floor. The room has two toy boxes - a red one for dolls and a white one for other toys. There is a dirty mark on the floor, and you may find a pack of wet wipes on the table. Your task: Please check the living room floor, if there are items on floor, put them where they belong (e.g., book should be placed on the sofa, toy should be placed in the toy box). Clean the mark, preferably with the wet wipes on the table (if available). If there aren’t any wipes, you can use the wet mop instead.
Decision-tree-structured behavior plan
[ { "There is a toy on the floor?": { "YES": { "What is the type of the toy on the floor?": { "doll": "Place the toy in the red box.", "other types": "Place the toy in the white box." }, "NO": "Do nothing." } }, { "There is a book on the floor?": { "YES": "Place the book on the sofa.", "NO": "Do nothing." } }, { "There is a wet wipe on the table?": { "YES": "Clean stain with the wet wipe.", "NO": "Clean stain with the wet mop." } } ]
Appendix
Table 7: Task description and decision-tree-structured behavior plan of the example Clean Living Room .
Figure 6: Correct policy (Behavior Tree) of the example Clean Living Room .
Figure 7: Faulty policy (Behavior Tree) of the example Clean Living Room , which belongs to the Counterfactuals category. Reason: The agent does not pick up the toy (non-doll) before placing it into the white toy box.
Figure 8: Faulty policy (Behavior Tree) of the example Clean Living Room , which belongs to the Unreachable category. Reason: The agent picks up the toy (doll) but does not place it into the red toy box.
Figure 9: Faulty policy (Behavior Tree) of the example Clean Living Room , which belongs to the Lackbranch category. Reason: The agent fails to plan for situations where wet wipes are absent.
Figure 10: Qualitative comparison on two LogicEnvEval tasks. While baseline methods suffer from physical implausibilities (e.g., floating objects), irrational layout or missing task-relevant items, LogicEnvGen generates environments with superior visual fidelity, layout coherence, and strict task alignment.
Figure 11: VLM evaluation of LogicEnvGen and four baselines. The results from both Gemini 3.1 Pro and InternVL3.5-38B consistently show LogicEnvGen outperforms four baselines across three criteria.
Figure 12: CLIP Score comparison over two scenario categories.
Figure 13: Comparative human evaluation of LogicEnvGen and four baselines across three criteria. The pie charts show the distribution of annotator preferences, showing both the percentage and the actual number of annotations favoring each method.
Method
PhyPR(%)
LogCov(%)
SceVR(%)
Fault Detection Rate (%)
CFactuals
Unreach
LBranch
Avg
Holodeck
100.00
37.07
74.00
24.00
22.00
28.00
24.67
Holodeck+Logic
100.00
87.92
90.26
76.00
84.00
78.00
79.33
LogicEnvGen
100.00
98.97
92.40
94.00
90.00
94.00
92.67
Appendix
Table 8: Attribution ablation results on the LogicEnvEval benchmark.
Checker: Valid
Checker: Invalid
Human Evaluator: Valid
89
0
Human Evaluator: Invalid
1
20
Appendix
Table 9: Confusion matrix of the validity checker against human evaluation.
Figure 14: Case 1. Agent task: Bring me a desk lamp from study if there is a pan in the kitchen. Otherwise, bring me a pillow from the bedroom.
Figure 15: Case 2. Agent task: The bedroom has a bed, possibly with a book or clock on it, along with a bookshelf and a desk. The bookshelf’s three levels hold novels, essays, and other books respectively. Clean items on the bed and put them into their designated storage areas (e.g., book on the shelf, phone on the desk).
Figure 16: Case 3. Agent task: The living room has a sofa, a tea table, and two toy boxes—the blue one for dolls and the white one for other toys. The floor is stained and there might be a sofa pillow or toy on it. Check the floor and put misplaced items (like a pillow or toy) in their proper places. Clean the stain, preferably with the wet wipes on the table (if available), otherwise use the wet mop by the wall.