stat.MESep 15, 2026

Conformal Policy Learning with Distribution-Free Safety Guarantees

Authors: Ying JinNaoki Egami

Organizations: Department of Statistics and Data Science, University of Pennsylvania · Department of Political Science & Statistics and Data Science Center, Massachusetts Institute of Technology

Abstract

Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of ``do no harm.'' In this paper, we propose \textit{conformal policy learning} (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs.

Explore similar work

CardsList
  1. Conformal Policy Control

    Mar 2, 2026Drew Prinster, Clara Fannjiang, Ji Won Park +4Conformal Risk ControlSafety Constraints

  2. Set-Valued Policy Learning

    May 19, 2026Laura Fuentes-Vicente, Mathieu Even, Gaëlle Dormion +3Treatment EffectReinforcement-Learning Policies