Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.
Figures & tables
Figure 1: Overview of policy refinement. 1. Generate the initial structure by querying an LLM with the task description. 2. Update parameters through imitation learning on the demonstrations. 3. Generate analysis code based on the desired behavior of the policy and its structure topology. 4. Collect tabular trajectories , recording observations, actions, and named latent values during policy rollouts. 5. Execute analyses and revise the structure based on the analysis results. The updated program returns to parameter fitting and another refinement cycle. 6. Reinforcement learning after the structure aligns with the demonstrations to go beyond the demonstration performance.
Figure 2: A sample speed-control subgraph and its corresponding PyTorch implementation. Solid nodes distinguish observations (gray), semantic latents (blue), action outputs (green), and trainable weights (orange); dashed nodes denote operation nodes (some collapsed due to space constraints). The returned speed_error exposes the latent value for further rollout analysis.
Figure 3: Example of off-track recovery structure refinement. (a) The steering structure before refinement (other nodes omitted); (b) The refined structure adds stronger steering when the race car is off-track; (c) The desired behavior of re-centering, corresponding test, and the result; (d) Recorded lateral offset and steering.
Method
Demo
LLM calls ↓
Resets ↓
Return ↑
Human demonstrations
-
-
-
871.0±12.8
MLP + IL
Yes
0
0
636.6±258.6
MLP + RL, low compute
Yes
0
400
591.8±330.5
MLP + RL, high compute
Yes
0
2000
850.2±101.4
Initial structure + IL
Yes
1
0
766.5±100.4
Initial structure + RL, low compute
Yes
1
400
657.0±37.3
Table 1: Car Racing performance and resource usage. Returns are mean ± standard deviation over ten random seeds. Resets include refinement and training interactions, excluding final evaluation.
Figure 4: Execution feedback improves policy refinement in Car Racing. Generated analysis avoids the large intermediate drop under direct table feedback. Shading shows standard deviations.
Figure 5: Door Opening progresses through three revisions. The numbers below show the mean and std of distance from handle to target across 10 seeds; lower is better.
Feedback
Input tokens (k) ↓
Output tokens (k) ↓
Trajectory table
110.1±1.6
17.9±1.3
Analysis results
26.0±3.4
37.5±2.9
Table 2: Generated analysis uses fewer input tokens and more output tokens than direct trajectory-table feedback. Values summarize six refinement rounds, with population standard deviations.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Strategy
MLES
Initial
Refined
Speed Control
speed during straights
✓
✓
✓
speed at track center
✓
✓
✓
slow at corners
✓
✓
✓
slow at sides / large offset
✓
✓
✓
slow when outside track
✓
Appendix
Table 3: Strategy categories identified by manual inspection of Car Racing policies.
Data Analysis
Iter 1
Iter 2
Iter 3
Sanity Checks
Weights stay constant
✓
✓
✓
Signs consistent with assumptions
✓
✓
✓
Summary statistics
✓
Actions
Output reconstruction
✓
✓
Appendix
Table 4: Generated diagnostics adapt to the policy across three refinement rounds. Checks cover computational consistency, action behavior, and task-specific control relationships.
Round
Slip feature
Fitted wq
Diagnosis and revision
0
Mean wheel reading
+0.1729
Positive coefficient violates the intended sign; replace wheel mean with a relative proxy.
1
uˉ−v
+3.8367
Positive coefficient persists; rectify the residual.
2
max(uˉ−v,0)
+6.6077
Sign anomaly persists; constrain the coefficient to be nonpositive.
3
p/(1+αp)
−0.00390
Compression and wq=−∣η∣ preserve the branch’s speed-reducing role after fitting.
Appendix
Table 5: Successive revisions align the slip proxy’s learned contribution with its intended speed-reducing role.
Figure 6: Recorded slip and target-speed values before refinement (left) and after three revisions (right), with shared axis limits.
Figure 7: Car Racing and Door Opening: continuous control from semantic observations.