Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to unrealistic synthetic data and unstable policy optimization. Many existing methods primarily control rollouts using transition-level uncertainty. We propose \emph{in-distribution imagination} (IDI), a rollout control framework that estimates trajectory support in a learned representation space and adaptively truncates rollouts that leave the offline trajectory manifold. Combined with trajectory-regularized RL, an extension of entropy-regularized RL, IDI consistently improves performance in limited-data settings. Experiments show that trajectory support predicts rollout failure substantially better than transition-level uncertainty, highlighting the importance of trajectory-level rollout control in MBORL.
Figures & tables
Figure 1 : Conceptual overview of in-distribution imagination (IDI). (a) Conventional MBORL relies on transition-level reliability, but locally plausible transitions may still accumulate into unrealistic trajectories due to compounding model error. (b) IDI instead measures trajectory support in a learned latent trajectory space and truncates imagination when trajectories leave the offline manifold. This converts rollout control from local uncertainty estimation to trajectory-level support estimation.
Figure 2 : Normalized return learning curves on MuJoCo locomotion benchmarks under the limited-data setting. TRAC is trained using only 100 episodes, while the dashed line denotes TRAC trained on the full dataset ( 1000 episodes). Fixed-horizon and uncertainty-based truncation improve over the limited-data baseline, but their effectiveness varies across environments. IDI consistently achieves the strongest performance through history-aware rollout control.
Figure 3
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Method
HalfCheetah
Hopper
Walker2d
TRAC (1000 traj.)
93.2±1.1
91.5±2.3
88.7±1.9
TRAC (100 traj.)
67.4±3.6
61.2±5.4
58.7±4.2
Fixed H=1
74.5±2.8
67.9±4.7
65.1±3.6
Fixed H=3
81.7±2.5
74.3±3.8
71.8±3.1
Fixed H=5
84.6±2.1
79.8±3.1
75.2±2.7
Fixed H=10
78.9±8.7
72.1±10.3
69.4±9.1
Appendix
Table 2 : Final normalized return (mean ± std) under the limited-data setting.
Method
HalfCheetah
Hopper
Walker2d
Fixed H=5
0.67
0.61
0.55
Uncertainty truncation
0.74
0.68
0.64
IDI
0.80
0.80
0.77
Appendix
Table 3 : Fraction of lost performance recovered relative to full-data TRAC.