Calibration-risk routing for controlled world-model adaptation
Organizations: Central South University
Abstract
Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.
Figures & tables
| Name | Base task | Simulator target dynamics | Counted violation |
|---|---|---|---|
| Ant actuator | Ant-v5 | 8 active actions 4–7 disabled | or rear-wall hit |
| Gravity | HalfCheetah-v5 | vertical gravity | |
| Friction Walker | Walker2d-v5 | friction factor |
| Method | Simulator train | Target train | Validation | Final test |
|---|---|---|---|---|
| Calibration-routed MC-WM | 50,000 | 50,000 | episodes | 20 episodes |
| Residual MC-WM core | 50,000 | 50,000 | episodes | 20 episodes |
| Direct target MBPO / target SAC | 0 | 50,000 | episodes | 20 episodes |
| DARC | 50,000 | 50,000 | episodes | 20 episodes |
| Raw simulator SAC | 50,000 | 0 | episodes | 20 episodes |
| Ant actuator | HalfCheetah gravity | Walker friction | ||||
|---|---|---|---|---|---|---|
| Method | Return | Violations | Return | Violations | Return | Violations |
| Calibration-routed MC-WM | 4084.9 2105.9 | 1.39 1.19 | 4954.8 488.6 | 0.20 0.40 | 394.6 68.4 | 0.31 0.09 |
| MC-WM core | 1361.7 527.6 | 1.38 1.48 | 4609.9 602.0 | 2.48 5.31 | 434.4 247.8 | 24.98 59.24 |
| Direct target MBPO | 4011.5 3346.2 | 1.26 1.32 | 4827.5 1043.2 | 0.66 0.89 | 350.7 100.7 | 23.59 21.45 |
| Target SAC | 2159.2 461.1 | 3.95 1.63 | 3770.3 644.4 | 9.52 10.24 | 261.2 98.8 | 7.16 8.19 |
| DARC | 1862.2 534.7 | 1.80 2.38 | 1268.6 528.1 | 44.23 26.03 | 252.5 138.0 | 17.26 19.03 |
| Environment | Selected R/D | Exact match | Return difference (95% CI) | Violation difference (95% CI) |
|---|---|---|---|---|
| Ant actuator | 0/10 | 10/10 | 1529.9 [426.9, 2833.1] | 0.72 [-0.05, 1.60] |
| HalfCheetah gravity | 0/10 | 10/10 | 188.0 [-178.7, 591.7] | 0.08 [-0.11, 0.37] |
| Walker friction | 10/0 | 10/10 | -21.2 [-56.0, 8.7] | -0.04 [-0.19, 0.06] |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Unique meaning |
|---|---|
| ; | state space; action space |
| domain MDP, with denoting simulator or target | |
| ; | transition kernel; expected one-step reward in domain |
| ; | policy-return discount; initial-state distribution |
| policy with parameter vector | |
| discounted return of policy in MDP |
| Environment | Active coeff. | Support overlap | SINDy MSE | Additive MSE | Error 95th pct. | Mean confidence | Reject rate |
|---|---|---|---|---|---|---|---|
| Ant, actuator failure | 228.9 (50.5) | 0.24 [0.19, 0.30] | 0.34 (0.04) | 0.21 (0.06) | 4.32 (0.73) | 0.23 (0.08) | 0.0% (0.0) |
| HalfCheetah, soft gravity | 271.2 (58.9) | 0.26 [0.16, 0.37] | 0.98 (0.94) | 0.40 (0.05) | 5.24 (0.32) | 0.27 (0.02) | 35.1% (8.4) |
| Walker2d, friction | 210.5 (30.2) | 0.20 [0.15, 0.26] | 2.53 (1.07) | 0.56 (0.20) | 6.22 (1.01) | 0.26 (0.04) | 55.9% (4.5) |
| Environment | Ablation | Return difference (95% CI) | Violation difference (95% CI) |
|---|---|---|---|
| Ant actuator | No checkpoint selection | -428.8 [-650.1, -212.8] | -0.01 [-0.32, 0.33] |
| Ant actuator | No confidence weighting | 593.8 [-2024.7, 3881.6] | 2.28 [-0.38, 6.29] |
| Ant actuator | Fixed residual route | -2231.8 [-4461.5, -871.3] | -0.50 [-1.18, 0.17] |
| Ant actuator | Fixed residual route, neural only | -2302.4 [-4409.3, -901.8] | -0.57 [-1.23, 0.02] |
| Ant actuator | No predicate filtering | -162.2 [-2178.5, 1960.5] | 1.43 [-0.72, 4.96] |
| Ant actuator | No sparse branch | 0.0 [0.0, 0.0] | 0.00 [0.00, 0.00] |
| Environment | Return difference (95% CI) | Violation difference (95% CI) | Calibration-risk difference (95% CI) | State-parameter difference (95% CI) | Classification |
|---|---|---|---|---|---|
| Ant actuator | 29.7 [-546.1, 652.8] | -0.04 [-0.81, 0.72] | 0.629 [0.485, 0.796] | 1414 [1396, 1431] | Unresolved |
| HalfCheetah gravity | 156.3 [-192.8, 506.9] | 0.66 [-0.24, 2.18] | -0.103 [-0.170, -0.047] | 1347 [1308, 1401] | Unresolved |
| Walker friction | 45.4 [-56.3, 147.4] | 0.33 [-0.08, 0.99] | -0.007 [-0.010, -0.004] | 1386 [1337, 1440] | Unresolved |
| Component | Parameter | Value |
|---|---|---|
| Simulator ensemble | members / hidden layers / width | |
| Simulator ensemble | batch size / learning rate | |
| Residual split | fit / selection / calibration | |
| Model-family route | metric / tie rule | standardized joint RMSE / residual |
| Sparse regression | ridge / threshold | |
| Feature discovery | budget / minimum gain |