Organizations: School of Intelligent Science and Technology, Nanjing University, China · National Key Laboratory for Novel Software Technology, School of Artificial Intelligence, Nanjing University, China
Model-based imitation learning (MBIL) improves real-environment interaction efficiency by optimizing policies on imagined rollouts from a learned world model. However, the gap between model-induced and real-environment occupancies makes policy learning sensitive to model error. Conservative MBIL mitigates model exploitation during policy optimization, but when real-environment interactions are collected by the same conservative policy, uncertain regions around the expert distribution remain insufficiently sampled. Generic uncertainty-driven exploration, on the other hand, may allocate interaction to novel but task-irrelevant dynamics. We propose REcoverability-CONditioned Exploration for Model-Based Imitation Learning (RECON). RECON separates conservative policy learning from active data collection by maintaining a main policy for task execution and an explorer for real-environment interaction. The explorer is optimized based on epistemic uncertainty conditioned on recoverability estimated from multi-step main-policy imagination, focusing data collection on unknown states from which the main policy can still return toward expert behavior. Experiments on locomotion, navigation and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness, indicating that RECON directs real-environment interaction toward recovery regions around the expert distribution that are underexplored by prior methods, and thereby learns a world model better suited for imitation.
Figures & tables
Figure 1: Conservative model use is not sufficient for model acquisition. A conservative main policy avoids uncertain dynamics, leaving recovery-relevant regions underexplored. RECON uses a separate explorer to collect uncertain but recoverable transitions.
Figure 2: RECON architecture on the DreamerV2 backbone . The conservative main policy uses uncertainty pessimistically, while a separate explorer uses recoverability-gated uncertainty for acquisition. Both update one world model; only the main policy is deployed.
Figure 3: Main experiments. Imitation learning across eight environments under matched demonstrations and real-interaction budgets. RECON changes training-time acquisition; only its conservative main policy is evaluated.
Environment
Clean
Noise
Delay
Impulse/scale
Hopper
905.8/ 909.1 (+3.3)
862.9/ 882.2 (+19.3)
408.5/ 596.2 (+187.7)
864.0/ 916.0 (+52.0)
Walker
421.2/ 593.9 (+172.7)
401.6/ 572.8 (+171.2)
133.0/ 184.9 (+51.9)
378.6/ 514.6 (+136.0)
Drawer
100.0 / 100.0 (0.0 pp)
100.0 /96.7 (-3.3 pp)
95.0/ 96.7 (+1.7 pp)
100.0 /92.6 (-7.4 pp)
Faucet
30.0/ 70.0 (+40.0 pp)
15.0/ 75.0 (+60.0 pp)
5.0/ 85.0 (+80.0 pp)
15.6/ 66.1 (+50.5 pp)
Handle
75.0 /70.0 (-5.0 pp)
67.5/ 77.5 (+10.0 pp)
80.0 / 80.0 (0.0 pp)
77.8/ 86.1 (+8.3 pp)
Hammer
0.0/ 2.5 (+2.5 pp)
8.3 /6.3 (-2.1 pp)
8.3/ 26.3 (+17.9 pp)
10.0 /3.9 (-6.1 pp)
Table 1: Robust deployment of the main policy, reported as CMIL/RECON means over three seeds (absolute Δ ). DMControl entries and deltas are returns; Maze and Meta-World entries are success percentages and deltas are percentage points (pp). Higher is better.
Figure 4: PointMaze visitation and acquisition. Left: State-visitation frequency of CMIL and RECON under the same training budget, with expert trajectories and representative real rollouts overlaid. Right: Acquisition allocation across expert-support, near-support, and far-OOD regions, defined by distance to the expert state cloud.
Figure 5: Recoverability ablation. “w/o recovery” removes the recoverability term and uses rte=Dψ(st,at)+βUθ(st,at) . Because recoverability gating rescales the effective uncertainty bonus in RECON, we use β=5 for the ungated variant for a fair comparison.
Figure 6: Design ablation. Left: alternative exploration signals. Middle: temporal placement of fixed-length collection bursts. Right: gate aggregation and the stop-gradient control.
Figure 7: Diagnostics of recoverability conditioning. Left: Held-out discounted latent open-loop prediction gap across different regions. Right: Predicted recoverability GH (rollout horizon H∈{5,15} ) versus empirical expert-tube re-entry.
Environment
Expert H15 gap
Replay H5 gap
Uncertainty
Hopper
1.229/ 0.979 (-20.3%)
0.397/ 0.272 (-31.5%)
2.8/2.8 (0%)
Walker
0.608/ 0.507 (-16.6%)
0.250/ 0.214 (-14.4%)
8.3/ 8.0 (-3.6%)
Drawer
1.537/ 1.219 (-20.7%)
0.499/ 0.394 (-21.0%)
4.0/ 3.3 (-17.5%)
Faucet
0.774/ 0.720 (-7.0%)
0.278/ 0.272 (-2.2%)
3.1 /3.2 (+3.2%)
Handle
1.072/ 0.817 (-23.8%)
0.401/ 0.330 (-17.7%)
3.1/ 2.8 (-9.7%)
Hammer
0.933 /0.934 (+0.2%)
0.349/ 0.343 (-1.6%)
3.2/ 2.9 (-8.3%)
Table 2: World-model diagnostics, reported as CMIL/RECON (relative change). Gaps are latent open-loop prediction errors defined in Appendix B.6 . Lower is better; uncertainty is in ×10−3 .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
πt={πe,πm,t∈W,otherwise,
Appendix
Algorithm 1 RECON training and data acquisition
Component
Hyperparameter
Value
Input/replay
observation; replay capacity
64×64 RGB; 5×106 transitions
sequence batch; sequence length
32; 50
random seed transitions; update ratio
5,000; 2 updates per 5 actions
World model
RSSM deterministic/stochastic size
400 / 360 (Gaussian)
transition ensemble; shared hidden size
10 heads; 400
encoder/decoder CNN depth
48 / 48
Appendix
Table 3: Shared RECON hyperparameters.
Figure 8: The eight evaluation environments. Policies receive the underlying 64×64 RGB observations.
Quantity
CMIL
RECON
Difference
World-model parameters
17.38M
17.38M
0
Total stored parameters
21.52M
25.46M
+3.94M (+18.3%)
Explorer share of RECON parameters
—
15.5%
—
FP32 storage for added parameters
—
15.8 MB
+15.8 MB
Serialized checkpoint
83 MiB
98 MiB
+15 MiB
Logged FPS, last 50 records
2.541±0.060
2.430±0.148
−4.4%
Appendix
Table 4: Compute overhead relative to CMIL.
Figure 9: Recoverability prediction and acquisition efficiency. (a) AUROC between the predicted recoverability score and empirical recovery within 15 real-environment steps as the imagination horizon varies. (b) Reduction in ensemble uncertainty on fixed main-policy anchor transitions. The acquisition rewards are CMIL: Dψ−10Uθ , V-MAIL: Dψ , U-only: Dψ+5Uθ , and RECON: Dψ+10UθG .
Figure 10: Disagreement ranking and held-out model error. H=15 open-loop prediction error across within-run uncertainty deciles on common held-out PointMaze spatial bins.
College of Connected Computing, Vanderbilt University, USA. · School of Computer Science, University of Waterloo, Canada. · School of Computer Science, The University of Sydney, Australia. +1