Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.
Figures & tables
Figure 1: MC-WM execution path. Initial target data fit a residual and a direct candidate; a disjoint calibration partition fixes the deployed family. Online target replay refits only the deployed state-and-reward model. Trust weights scale losses rather than physical rewards; a separate violation cost is described in Section 3.5 .
Name
Base task
Simulator → target dynamics
Counted violation
Ant actuator
Ant-v5
8 active → actions 4–7 disabled
∣vx∣>1.5 or rear-wall hit
Gravity
HalfCheetah-v5
vertical gravity −19.62→−9.81
qz>0.2
Friction Walker
Walker2d-v5
friction factor 0.3→1.0
s(0)>1.25
Table 1: Controlled simulator-to-target environments. Each reported constraint is identical in the simulator and target member of a pair.
Method
Simulator train
Target train
Validation
Final test
Calibration-routed MC-WM
50,000
50,000
8×5 episodes
20 episodes
Residual MC-WM core
50,000
50,000
8×5 episodes
20 episodes
Direct target MBPO / target SAC
0
50,000
8×5 episodes
20 episodes
DARC
50,000
50,000
8×5 episodes
20 episodes
Raw simulator SAC
50,000
0
8×5 episodes
20 episodes
Table 2: Registered per-seed environment access. Episode lengths determine the realized validation and test transition counts, which are recorded in the interaction ledger.
Ant actuator
HalfCheetah gravity
Walker friction
Method
Return ↑
Violations ↓
Return ↑
Violations ↓
Return ↑
Violations ↓
Calibration-routed MC-WM
4084.9 ± 2105.9
1.39 ± 1.19
4954.8 ± 488.6
0.20 ± 0.40
394.6 ± 68.4
0.31 ± 0.09
MC-WM core
1361.7 ± 527.6
1.38 ± 1.48
4609.9 ± 602.0
2.48 ± 5.31
434.4 ± 247.8
24.98 ± 59.24
Direct target MBPO
4011.5 ± 3346.2
1.26 ± 1.32
4827.5 ± 1043.2
0.66 ± 0.89
350.7 ± 100.7
23.59 ± 21.45
Target SAC
2159.2 ± 461.1
3.95 ± 1.63
3770.3 ± 644.4
9.52 ± 10.24
261.2 ± 98.8
7.16 ± 8.19
DARC
1862.2 ± 534.7
1.80 ± 2.38
1268.6 ± 528.1
44.23 ± 26.03
252.5 ± 138.0
17.26 ± 19.03
Table 3: Held-out target-domain outcomes over ten independently trained seeds (mean ± sample SD). Higher return and fewer violations are better. Test seeds are fixed across methods; target training, validation, and test access are reported separately in Appendix D .
Environment
Selected R/D
Exact match
Return difference (95% CI)
Violation difference (95% CI)
Ant actuator
0/10
10/10
1529.9 [426.9, 2833.1]
0.72 [-0.05, 1.60]
HalfCheetah gravity
0/10
10/10
188.0 [-178.7, 591.7]
0.08 [-0.11, 0.37]
Walker friction
10/0
10/10
-21.2 [-56.0, 8.7]
-0.04 [-0.19, 0.06]
Table 4: Exact route intervention over ten paired seeds. Selected R/D is the automatic residual/direct count; exact match compares the complete metric stream with the selected forced arm. Differences are automatic routing minus the unselected forced alternative, so positive return and negative violations favor the selector.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Unique meaning
S ; A
state space; action space
Mℓ
domain MDP, with ℓ∈{S,T} denoting simulator or target
Pℓ ; ρℓ
transition kernel; expected one-step reward in domain ℓ
γ ; p0
policy-return discount; initial-state distribution
πϕ
policy with parameter vector ϕ
JM(π)
discounted return of policy π in MDP M
Appendix
Table 5: Complete mathematical notation.
Environment
Active coeff.
Support overlap
SINDy MSE
Additive MSE
Error 95th pct.
Mean confidence
Reject rate
Ant, actuator failure
228.9 (50.5)
0.24 [0.19, 0.30]
0.34 (0.04)
0.21 (0.06)
4.32 (0.73)
0.23 (0.08)
0.0% (0.0)
HalfCheetah, soft gravity
271.2 (58.9)
0.26 [0.16, 0.37]
0.98 (0.94)
0.40 (0.05)
5.24 (0.32)
0.27 (0.02)
35.1% (8.4)
Walker2d, friction
210.5 (30.2)
0.20 [0.15, 0.26]
2.53 (1.07)
0.56 (0.20)
6.22 (1.01)
0.26 (0.04)
55.9% (4.5)
Appendix
Table 6: Per-run model diagnostics aggregated across training seeds. Parentheses contain cross-seed SD unless a range is shown.
Environment
Ablation
Return difference (95% CI)
Violation difference (95% CI)
Ant actuator
No checkpoint selection
-428.8 [-650.1, -212.8]
-0.01 [-0.32, 0.33]
Ant actuator
No confidence weighting
593.8 [-2024.7, 3881.6]
2.28 [-0.38, 6.29]
Ant actuator
Fixed residual route
-2231.8 [-4461.5, -871.3]
-0.50 [-1.18, 0.17]
Ant actuator
Fixed residual route, neural only
-2302.4 [-4409.3, -901.8]
-0.57 [-1.23, 0.02]
Ant actuator
No predicate filtering
-162.2 [-2178.5, 1960.5]
1.43 [-0.72, 4.96]
Ant actuator
No sparse branch
0.0 [0.0, 0.0]
0.00 [0.00, 0.00]
Appendix
Table 7: Paired component matrix for the frozen gap-routed predecessor. Differences are ablation minus predecessor; negative return therefore favors the predecessor, while negative violations favor the ablation. This matrix is not a source-identical ablation of the final calibration-risk selector.
Environment
Return difference (95% CI)
Violation difference (95% CI)
Calibration-risk difference (95% CI)
State-parameter difference (95% CI)
Classification
Ant actuator
29.7 [-546.1, 652.8]
-0.04 [-0.81, 0.72]
0.629 [0.485, 0.796]
1414 [1396, 1431]
Unresolved
HalfCheetah gravity
156.3 [-192.8, 506.9]
0.66 [-0.24, 2.18]
-0.103 [-0.170, -0.047]
1347 [1308, 1401]
Unresolved
Walker friction
45.4 [-56.3, 147.4]
0.33 [-0.08, 0.99]
-0.007 [-0.010, -0.004]
1386 [1337, 1440]
Unresolved
Appendix
Table 8: Paired dense-minus-sparse residual comparison under fixed residual deployment. Calibration risk is standardized transition error, and state parameters count trainable parameters in the state-correction branch.
Figure 2: Validation return and violations against cumulative target access. Lines are ten-seed means and bands are 20,000-resample percentile-bootstrap 95% confidence intervals for the cross-seed mean. The horizontal coordinate includes target training plus cumulative validation interactions. Validation outcomes do not enter gradient updates, but the registered candidate may use its validation trajectory for checkpoint restoration.
Figure 3: Actual target access by method and environment. Bars show cross-seed means for training, validation, and final-test transitions; error bars show the 20,000-resample percentile-bootstrap 95% CI of mean total access.
Component
Parameter
Value
Simulator ensemble
members / hidden layers / width
5/3/200
Simulator ensemble
batch size / learning rate
256/10−3
Residual split
fit / selection / calibration
0.8/0.1/0.1
Model-family route
metric / tie rule
standardized joint RMSE / residual
Sparse regression
ridge / threshold
0.05/0.05
Feature discovery
budget / minimum gain
100/10−5
Appendix
Table 9: Fixed MC-WM hyperparameters used by every calibration-routed candidate run.