Augmented Feature Boosting for Multicalibration
Organizations: Cornell University · Ben-Gurion University of the Negev
Abstract
Multicalibration requires a predictor's residuals to be unbiased not only globally, but also after conditioning on the predictor's own level sets and reweighting by a rich class of test functions. Standard boosting approaches in the distributional setting achieve this by repeatedly discretizing the predictor's range then auditing and repairing the resulting level sets. One consequence is that in practice, the algorithm's guarantees are sensitive to this parametrization of the rounding parameter. A natural theoretical question, then, is how to do discretization-free boosting which avoids this rounding within the boosting process itself. Here, we analyze an alternative feature-augmentation boosting paradigm inspired by Tax et al. (2026): at each round, a squared-loss oracle is called on hypotheses that receive the previous predictor's output as an additional feature, and only the final predictor is rounded to have a finite set of level sets to provide the multicalibration guarantee with respect to. We give a theoretical analysis of this procedure through the expressivity of the augmented hypothesis class, and show how the expressivity of this class yields a hierarchy of guarantees, including multiaccuracy, multicalibration, and the stronger notion of level-set multicalibration.
Figures & tables
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Notation | Meaning |
| Feature space and label space. Throughout the main analysis, labels lie in , and often in . | |
| Population distribution over examples . | |
| Empirical sample of independent examples from . When used inside empirical losses, also denotes the uniform distribution over this sample. | |
| Sample size. | |
| The set . | |
| , | Squared loss of a predictor under the population distribution and empirical sample , respectively. |