Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
Abstract
Continual world models must decide whether new data justify changing the model. Fixed replay schedules and prediction-error triggers specify when to update, but neither reveals the value of an individual update: one deployment run cannot show how the same model would have performed at that moment had it held its parameters. We introduce the fork ledger, which branches a deployment stream at pre-registered decision points into matched update and hold continuations under common random numbers. It evaluates both continuations on the same episodes and records . Always applying one fixed update mechanism lowers return on all three simulated control tasks: CartPole (; checkpoint-bootstrap CI , against a converged return near ), Walker (; ) and Cheetah (; ). Divergence is an outcome of applying the update, so the estimand counts every attempted fork; restricted to the of that did not collapse, CartPole and Walker are unchanged in sign ( and ) and Cheetah becomes unresolved (; ). The task is the unit of inference: each contributes attempted forks over five pretrained checkpoints crossed with two drift directions. The ledger makes counterfactual utility observable for a fixed mechanism, allowing triggers to be judged by the updates they select rather than by surprise detection alone.