Authors: Wenbin Zhou, Michael Lingzhi Li, Shixiang Zhu
Organizations: Heinz College of Information Systems and Public Policy, Carnegie Mellon University · Technology and Operations Management, Harvard Business School
Models can be retrained as new data arrive, but deploying every new version risks replacing a good policy with a worse one. We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression. Our offline meta-policy maximizes expected cumulative value subject to a budget on the expected number of updates that perform worse than the policies they replace. We estimate the value and risk of possible switches from historical learning trajectories, represent an update schedule as a path in a directed acyclic graph, and select a schedule using dynamic programming. A leading-order analysis identifies the signal-to-noise ratio of policy improvement as a key driver of update frequency, waiting times, and risk allocation: clearer improvements support earlier, more frequent updates, while noisier improvements call for longer waits or greater risk expenditure. Their asymptotic rates also reveal a diminishing marginal cost of achieving greater safety over time. Experiments on synthetic and clinical trial data illustrate the performance--risk tradeoff and compare our method with alternative baselines.
Figures & tables
Figure 1: The core tradeoff in deciding how often to update a policy. Updating slowly (left) leaves an outdated policy in place and loses value, while updating quickly (right) tracks improvements closely but risks deploying a policy worse than the one it replaces (a safety violation).
Figure 2: A deployment schedule and its path representation for T=4 . Left : The meta-policy updates at times 2 and 4 , retaining π2 at time 3 . Right : The corresponding path 1→2→4→5 deploys π1,π2,π4 , with node 5=T+1 serving as the terminal sink.
Figure 3: Phase diagrams of the schedules returned by Algorithm 1 on homogeneous value processes with constant drift μ and volatility σ ( T=512 , ϵ=0.35 , Δ=1×10−3 ). From left to right, the panels report the number of updates K , the mean per-update risk pˉ , the mean normalized pre-update interval preceding an update hˉ , and the expected deployed value J . Each cell is evaluated under the population rewards and safety costs and averaged over three simulated datasets. White-labeled contours show the leading-order theoretical predictions established in Section 4 .
Figure 4: Safety–value Pareto frontiers of the proposed meta-policy (Ours) and three baseline update rules, in three synthetic settings with low, medium, and high value-process SNR and on the International Stroke Trial (IST) data. The horizontal axis is the safety budget ϵ (log scale). The vertical axis is the objective of ( 1 ), normalized so that 0 and 1 correspond to never updating and to updating every period. All methods use the same estimated rewards and safety costs.
Figure 5: Update times in the optimal schedule selected by the proposed meta-policy as the safety budget ϵ varies, in the same four settings as Figure 4 . Each black dot indicates a scheduled normalized update time. The left dashed lines mark the smallest budget that admits one update.
An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause harm and must be taken offline, curtailing any future interaction. Imitating old behavior is safe, but excessive conservatism discourages exploration. How much behavior change is too much? We show how to use any safe reference policy as a probabilistic regulator for any optimized but untested policy. Conformal calibration on data from the safe policy determines how aggressively the new policy can act, while provably enforcing the user's declared risk tolerance. Unlike conservative optimization methods, we do not assume the user has identified the correct model class nor tuned any hyperparameters. Unlike previous conformal methods, our theory provides finite-sample guarantees even for non-monotonic bounded loss functions, and it introduces a new policy control setting. Our experiments on applications ranging from natural language question answering to biomolecular engineering show that safe exploration is not only possible from the first moment of deployment, but can also improve performance.
Drew Prinster, Clara Fannjiang, Ji Won Park +4
Prescient Design, Genentech, U.S.A. · Johns Hopkins University, Baltimore, U.S.A. · New York University, NYC, U.S.A. +1
Predictive models are often deployed through existing decision policies that stakeholders are reluctant to change unless a risk constraint requires intervention. We study risk-controlled post-processing: given a deterministic baseline policy, choose a new policy that maximizes agreement with the baseline subject to a chance constraint on a user-specified loss. At the population level, we show that the optimal policy has a threshold structure: it follows the baseline except on contexts where switching to the oracle fallback policy yields a large reduction in conditional violation risk. At the finite-sample level, given a fitted fallback policy and score, we develop a post-processing algorithm that uses calibration data to select a threshold. Leveraging tools from algorithmic stability and stochastic processes, we show that under regularity conditions, in the i.i.d. setting, the expected excess risk of the post-processed policy is O(logn/n). In the special case when an exact-safe fallback policy is available, the algorithm achieves precise expected risk control under exchangeability. In this setting, we also give high-probability near-optimality guarantees on the post-processed policy. Experiments on a COVID-19 radiograph diagnosis task, an LLM routing problem, and a synthetic multiclass decision task show that targeted post-processing can meet or nearly meet risk budgets while preserving substantially more agreement with the baseline than score-blind random mixing.
Meta-reinforcement learning (meta-RL) enables agents to adapt to unseen tasks with limited experience. Despite its promise, the application of meta-RL in real-world tasks is hindered by safety requirements, which have been underexplored in prior work. In this paper, we propose a safe meta-RL framework that explicitly accounts for safety during adaptation. Our key insight is to reason about safety in the information space, which captures both the physical state and the agent's belief over the underlying task. Within this space, we introduce a safety value function that measures the probability of the agent avoiding unsafe regions indefinitely. We show that this function satisfies a self-consistency condition and a Bellman equation, which make it learnable via meta-RL. Based on this formulation, we develop a safe meta-RL algorithm that learns the safety value function and leverages it for safety filtering and constrained policy optimization. Experiments on meta-RL benchmarks demonstrate the effectiveness of the proposed method.
Zeyang Li, Sunbochen Tang, Navid Azizan
Laboratory for Information and Decision Systems (LIDS), Massachusetts Institute of Technology, Cambridge, MA 02139, USA