cs.AIOct 1, 2026

Calibration-risk routing for controlled world-model adaptation

Authors: Yifan Zhang, Liang Zheng

Organizations: Central South University

Abstract

Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL

    Jul 18, 2026Zhaohui WangModel-Based Reinforcement LearningProbabilistic Safety

  2. All Models are Wrong, Knowing Where is Useful: On Model Uncertainty in Reinforcement Learning

    May 31, 2026Bernd Frauenknecht, Devdutt Subhasish, Artur Eisele +2Model-Based Reinforcement LearningScalable Robot Learning

  3. Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

    Sep 10, 2026Fernando Palafox, David Fridovich-KeilModel-Based Reinforcement LearningIn-Context Learning