cs.LGOct 5, 2026

Multigroup Fairness and Omniprediction: Separations and Equivalences

Authors: Sílvia Casacuberta, Parikshit Gopalan, Varun Kanade, Omer Reingold, Konstantinos Stavropoulos, Pranay Tankala

Organizations: Stanford University · Apple · University of Oxford · Institute for Advanced Study · Harvard University

Abstract

Omniprediction is a learning guarantee which requires a single predictor to be competitive relative to the best hypothesis from a benchmark class for any loss chosen from a family of loss functions. Loss Outcome Indistinguishability (loss OI for short) is a stronger notion that implies omniprediction. It requires the predicted distribution on labels to be indistinguishable from the true distribution to tests that depend on the loss functions and the benchmark class. Multiaccuracy and multicalibration are multigroup fairness notions that generalize classical notions of calibration and accuracy in expectation. Most known learning algorithms for omniprediction (both for the standard notion and for strengthenings like loss OI) rely on some version of these multigroup fairness notions, or on an intermediate notion called calibrated multiaccuracy. We ask if this is necessary: Does omniprediction require some form of multigroup fairness? We show that the answer is no for (plain) omniprediction, and yes for loss OI. First, a sequence of works shows that multicalibration or calibrated multiaccuracy imply omniprediction. We rule out even a weak converse, by showing that omniprediction for proper losses does not imply even accuracy in expectation, a much weaker notion than any of calibration, multiaccuracy, or multicalibration. Second, prior work showed how to achieve loss OI from a combination of calibration and multiaccuracy. We show a converse: loss OI is equivalent to a form of calibrated multiaccuracy.

Explore similar work

Jun 18, 2026cs.LG

Optimal Deterministic Multicalibration and Omniprediction

A model is multicalibrated on a collection of group weights GG if it is calibrated -- i.e. unbiased even conditional on its prediction -- not just overall, but also after reweighting contexts by each g∈Gg \in G. It is a useful property for many downstream applications and is a basic desideratum of trustworthy machine learning. Before this work, all predictors known to attain the minimax-optimal O~(ε−3)\widetilde O(\varepsilon^{-3}) sample complexity rate for ε\varepsilon-multicalibration were randomized, while deterministic predictors were known only with substantially worse sample complexity. Whether randomization is necessary for optimal sample complexity in multicalibration was explicitly asked by [CLNR26] and implicitly in several prior works. We resolve this open problem by giving a minimax-optimal multicalibration algorithm that outputs a deterministic predictor. We then generalize the algorithm to produce optimal deterministic predictors that satisfy outcome indistinguishability (OI) with respect to finite or finitely covered collections of tests. As an application, this also gives deterministic omnipredictors and panpredictors with optimal sample complexity, resolving open problems posed by [OKK25] and [BHHLZ25].
Jul 31, 2026cs.LG

End-to-End Fairness Optimization with Fair Decision-Focused Learning

Many real-world systems rely on predictive models to inform decisions, and fairness concerns arise in both the prediction and decision stages. We introduce end-to-end fairness optimization (E2EFO) as a unifying framework that integrates fairness across the prediction-to-decision pipeline. We focus on resource allocation with group-based fairness: the prediction task estimates allocation impacts while limiting accuracy disparity across groups, and the decision task distributes those impacts equitably by optimizing a group-based alpha-fairness measure. Within this framework, we propose fair decision-focused learning (FDFL), a training paradigm that jointly accounts for prediction accuracy, prediction fairness, and decision regret -- the loss in decision fairness due to imperfect predictions. FDFL trains the predictor by gradient descent, combining the objective gradients through multi-task learning techniques. The core computational challenge is the decision Jacobian with respect to the predictor parameters: we derive exact closed-form formulas for a tractable class of fair allocation and apply a differentiable optimization layer in the general case. We further establish a finite-sample generalization bound for the scalarized FDFL objective. Numerical experiments on a healthcare-based single resource allocation and a synthetic multiple resource allocation illustrate the value of jointly accounting for prediction fairness and decision fairness in prediction-informed decision-making.
May 23, 2026stat.ML

Multicalibration Boosting: Theory, Convergence, and Transferability

Multicalibration extends classical calibration by requiring predictions to be unbiased over a rich collection of functions, encompassing both prediction slices and subpopulations. It has emerged as a powerful framework for fairness, robustness, and reliable prediction, yet the theoretical understanding of multicalibration boosting (MCBoost) remains fragmented and often relies on restrictive assumptions. In this work, we develop a unified and refined perspective on MCBoost that subsumes existing variants, including multiaccuracy, BatchGCP, and BatchMVP. We uncover several phenomena that provide new insights into its practical behavior: even highly accurate and flexible predictors can remain substantially miscalibrated; enforcing multicalibration introduces a calibration-risk trade-off; and early stopping plays a central role in controlling this trade-off. On the theoretical side, we establish a general framework for MCBoost under weaker and more realistic conditions. We show that the boosting iterates converge to a Bregman projection of the population-optimal predictor onto the cumulative span generated by the audit class, thereby explicitly characterizing the function space on which multicalibration is achieved. We further derive convergence rates under different smoothness assumptions, finite-sample guarantees, and principled stopping rules that ensure multicalibration at termination. Finally, we extend the theory of universal adaptability under covariate shift, providing more general transfer guarantees and clarifying when multicalibrated predictors generalize across domains. These results provide a more complete theoretical foundation and practical guidance for multicalibration boosting, positioning it as both a unifying framework and a reliable post-processing approach for modern predictive models.