stat.MEApr 15, 2026

Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance

Authors: Carri W. Chan, Yi Han, Hannah Li, Benjamin L. Ranard

Organizations: Decision, Risk, and Operations, Columbia Business School · Department of Statistics, Columbia University · Division of Pulmonary, Allergy, and Critical Care Medicine, Department of Medicine, Vagelos College of Physicians and Surgeons, Columbia University

Abstract

AI tools increasingly drive targeted interventions in various service settings, including healthcare, education, and public services. Algorithms score individuals, trigger outreach to those above a threshold (e.g., high-risk or high-value), and encourage them to request service; then providers deliver service to those who request. Much of the work in this area has focused on improving predictive accuracy, implicitly assuming that better predictions lead to better outcomes. We show that predictions are only one component of a larger service system: when service capacity is limited and behavioral responses to outreach are probabilistic, system efficacy depends on operational forces that predictive accuracy does not capture. In such settings, the optimal score threshold must balance two effects: ensuring all capacity is filled (utilization) and, when capacity is constrained, preventing low-value requests from crowding out high-value ones (cannibalization). We characterize the optimal threshold and prove that thresholds based solely on predictive accuracy are generally suboptimal. Further, algorithm selection metrics such as AUC can be misaligned with operational performance: they weight all thresholds equally, while optimal deployment uses a subset of thresholds that depends on both capacity and compliance behavior. We introduce a new metric, Operational AUC (OpAUC), and show that it identifies the efficacy-optimal algorithm. Finally, we conduct a case study on sepsis early warning data that illustrates the magnitude of the improvement available from better threshold selection and shows that a predictor with lower AUC can achieve higher system efficacy under optimal deployment.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RISED: A Pre-Deployment Evaluation Framework for High-Stakes AI Decision-Support Systems, with Application to Healthcare

    May 13, 2026Rohith Reddy Bellibatlu, Manpreet Singh, Yash Jajoo +2Clinical Decision SupportPre-Deployment Safety Assessments

  2. Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing

    Jul 29, 2026Brandon Gower-Winter, Georg KremplClinical PredictionImplications

  3. Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

    May 21, 2026Naveen Raman, Santiago Cortes-Gomez, Mateo Dulce Rubio +2HealthcarePre-Deployment Safety Assessments