Production machine-learning models are derived artifacts of time-bounded training snapshots: a deployed model is a materialized view over a training cut that ages the instant it is built. A common response is to replace the fixed retraining cadence with an adaptive trigger -- a weighted staleness score that retrains when accumulated source risk crosses a threshold. We show this is the wrong lever, and identify the right one. First, an equivalence limit: any refresh trigger that is a static, strictly monotone function of a single shared global training-data age is operationally equivalent to a calibrated uniform age timer, so a global staleness budget, however elaborately it weights segments, sources, and sensitivities, carries no scheduling information a clock does not. The limit also shows how to escape it: refresh segments differentially, giving each its own age and refresh interval, which is meaningful when refresh cost is separable across segments (incremental training or per-segment models). We solve the resulting budget-allocation problem. In the frequent-refresh regime each segment's optimal refresh rate is proportional to the square root of its risk wjλj (weight times change rate), and the optimal policy never costs more than the uniform timer, beating it by a closed-form Cauchy-Schwarz "price of uniformity" that is zero for homogeneous workloads and grows with heterogeneity. In a discrete-event simulation with real Poisson change events, the optimal policy lowers realized weighted stale exposure by 8-29% relative to the uniform timer at matched refresh budget, winning on 86-100% of seeds; a naive exposure-threshold policy does not, showing the allocation is what helps; and the advantage survives 50% rate-estimation noise. The leverage in model refresh is not a better score but a better action.
Figures & tables
Figure 1: The equivalence limit (Theorem 1 ). NWSE(a) for a heterogeneous six-segment set is strictly monotone in the shared age, so each budget β maps through F−1 to a unique age τβ : thresholding the score is thresholding age in different units.
Figure 2: The optimal allocation (Theorem 2 ). In the frequent-refresh regime the exact numerical optimum obeys the closed-form square-root law xj⋆∝wjλj (log–log slope 1/2 ); shown for a 40-segment moderate-heterogeneity profile.
Heterogeneity
Sparse
Moderate
Frequent
Low
0.2
0.2
0.3
Moderate
7.5
7.1
8.4
High
11.4
12.9
14.0
Extreme
21.0
24.1
29.5
Table 1: Analytical cost saving of optimal differential refresh over the uniform timer (mean over 50 profiles, %). Uniform is optimal at low heterogeneity; the gap grows with dispersion of segment risk.
Heterogeneity
K
Opt. win (%)
Naive (%)
Opt. seeds (%)
Moderate
10
0 8.6
− 0.2
0 98
Moderate
20
0 7.7
−4.3
0 86
Moderate
40
0 8.6
−8.1
0 94
High
10
11.0
− 4.1
100
High
20
12.7
− 3.0
100
High
40
16.0
− 2.5
100
Table 2: Realized weighted stale exposure saving over the uniform timer in the discrete-event simulation, at matched refresh count (50 seeds). K is the uniform timer’s number of full refreshes. “Opt. win” is the optimal differential policy; “Naive” is the exposure-threshold policy; “seeds” is the fraction on which the optimal policy wins.
Figure 3: Realized saving over the uniform timer at matched refresh count (mean over the three budgets of Table 2 , error bars ± s.d.). The optimal differential policy wins and grows with heterogeneity; the naive exposure-threshold policy is near break-even or worse.
Heterogeneity
noise 0.00
0.10
0.25
0.50
Moderate
7.0
7.0
6.6
5.5
High
13.0
12.9
12.6
11.3
Extreme
24.3
24.3
24.0
23.2
Table 3: Analytical win (%) under multiplicative λ -estimation noise at the moderate budget (50 profiles, mean).
Organizations often have an incumbent predictive model in production when new data sources become available. Because historical training data lack the new features, a challenger model must be trained on a small but growing full-feature dataset. We study whether, and when, the organization should switch to the challenger. The decision is statistical and economic: the challenger's predictive performance improves as full-feature data accumulate, but repeated retraining is costly and delays benefits from deployment. We develop a framework linking learning-curve dynamics to model-switching economics. Under a standard power-law learning curve and finite data-collection horizon T, the optimal time to train and evaluate the challenger scales as T1/(1+α): learning-curve shape (through its learning speed α) is the primary theoretical determinant of when to stop experimenting; costs determine switching profitability. Even without knowing the learning curve, the operational problem is tractable: we show that any algorithm stopping on the T2/3 scale and making reliable switch/discard decisions achieves O(T2/3logT) regret relative to a full-foresight oracle. We propose a sequential evaluation algorithm that uses local learning-curve trends to anticipate improvement, and test it in a real-world credit-scoring study. Even with this local approximation, the algorithm theoretically and empirically achieves near-oracle performance. It is also more stable than greedy sequential evaluation algorithms, where noisy early estimates trigger premature discarding, or simple one-shot evaluation algorithms, which work only when their fixed evaluation time matches the (unknown in practice) theoretical timing scale. Our framework offers a step toward principled model governance when new data sources require costly collection, validation, and deployment.
Models can be retrained as new data arrive, but deploying every new version risks replacing a good policy with a worse one. We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression. Our offline meta-policy maximizes expected cumulative value subject to a budget on the expected number of updates that perform worse than the policies they replace. We estimate the value and risk of possible switches from historical learning trajectories, represent an update schedule as a path in a directed acyclic graph, and select a schedule using dynamic programming. A leading-order analysis identifies the signal-to-noise ratio of policy improvement as a key driver of update frequency, waiting times, and risk allocation: clearer improvements support earlier, more frequent updates, while noisier improvements call for longer waits or greater risk expenditure. Their asymptotic rates also reveal a diminishing marginal cost of achieving greater safety over time. Experiments on synthetic and clinical trial data illustrate the performance--risk tradeoff and compare our method with alternative baselines.
Wenbin Zhou, Michael Lingzhi Li, Shixiang Zhu
Heinz College of Information Systems and Public Policy, Carnegie Mellon University · Technology and Operations Management, Harvard Business School
Continual world models must decide whether new data justify changing the model. Fixed replay schedules and prediction-error triggers specify when to update, but neither reveals the value of an individual update: one deployment run cannot show how the same model would have performed at that moment had it held its parameters. We introduce the fork ledger, which branches a deployment stream at pre-registered decision points into matched update and hold continuations under common random numbers. It evaluates both continuations on the same episodes and records ΔR=Rupdate−Rhold. Always applying one fixed update mechanism lowers return on all three simulated control tasks: CartPole (−144.0; checkpoint-bootstrap 95% CI [−185.4,−116.1], against a converged return near 650), Walker (−82.8; [−101.1,−61.7]) and Cheetah (−18.6; [−29.0,−6.6]). Divergence is an outcome of applying the update, so the estimand counts every attempted fork; restricted to the 693 of 720 that did not collapse, CartPole and Walker are unchanged in sign (−113.4 and −82.1) and Cheetah becomes unresolved (−3.9; [−17.5,+13.0]). The task is the unit of inference: each contributes 240 attempted forks over five pretrained checkpoints crossed with two drift directions. The ledger makes counterfactual utility observable for a fixed mechanism, allowing triggers to be judged by the updates they select rather than by surprise detection alone.