The following motif is common in spatiotemporal settings: we have a sequence of covariate and label pairs observed for a relatively short, recent time period. We have access to unlabeled covariates over a longer time period. Data is observed over many spatial locations. For instance, crop yield might be observed over a large geographical area for recent years, but weather data (which is informative about crop yield) is available for a much longer period. The goal is to estimate, at each spatial location, the expected label (e.g., crop yield) in the future and provide a valid confidence interval for this value. The observed time period alone is too short for reliable estimates. Imputing missing labels with machine learning can cause substantial bias. Prediction-powered inference (PPI) can correct for this bias, but it relies on an i.i.d. assumption that breaks under our expected temporal dependencies. Heteroskedasticity and autocorrelation consistent (HAC) procedures account for temporal correlation, but have not been adapted to cases where some labels are imputed. We provide reliable point estimates and confidence intervals given: short labeled time series (across spatial locations), a longer unlabeled time series, and an imperfect predictor of labels given covariates. We show our method outperforms natural alternatives.
Figures & tables
Δ
E[Δ^]
E[θ^PPI−θ]
Chunking
−0.606
−0.602
0.003
Interleaving
−0.080
−0.001
0.078
Table 1: For each row (using a chunked or interleaved split), we report true prediction bias (left column), estimated prediction bias (middle), and bias of the PPI estimate (right). In each column, we report the expected result by averaging over 105 replicates. In every case, the standard error has size strictly less than 3×10−3 ; see Table 3 in Appendix C for the full standard errors.
Less Extreme Weather
More Extreme Weather
Method
Scaled RMSE
Coverage
Scaled Half-Width
Scaled RMSE
Coverage
Scaled Half-Width
Classical
0.0503
30.02%
0.0194
0.0372
31.56%
0.0147
IWI
0.0399
21.19%
0.0088
0.0302
20.79%
0.0064
chunkPPI
0.0299
30.91%
0.0115
0.0218
32.88%
0.0087
HAC
0.0503
91.72%
0.0871
0.0372
88.00%
0.0593
PPITSAS (ours)
0.0299
93.45%
0.0527
0.0218
92.64%
0.0362
Table 2: For each method, we report the performance metrics in the less extreme weather setting (left) and the more extreme weather setting (right). The best performance for each metric, according to the goalposts defined in Section 5.1 , are shaded in green; in the event of a tie, all winners are shaded. We ignore scaled half-widths for coverages below 85%, and shade them in gray.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
DL={(Xt(s′),Yt(s′))}t∈TL,s′∈S
Appendix
Algorithm 1 PPITSAS . PPI for time-series across space
Wˉ←T1∑t=1TWt.
(13)
Appendix
Algorithm 2 HAC . Heteroskedasticity and autocorrelation consistent variance estimation [ 14 ]
E[Δ]
E[Δ^]
E[θ^PPI−θ]
Chunking
−0.6058(0.0020)
−0.6024(0.0022)
−0.0029(0.0011)
Interleaving
−0.0796(0.0008)
−0.0012(0.0002)
−0.0778(0.0007)
Appendix
Table 3: For each row (using a chunked or interleaved split of the labeled data), we report the true prediction bias (left column), estimated prediction bias (middle column), and bias of the PPI estimate (right column). In each column, we report the result averaged over 105 replicates. Monte Carlo standard errors are reported in parentheses.
Many applications require statistically valid inference across many related tasks, while using only a handful of high-quality labels per hypothesis. In AI evaluation, these tasks may correspond to model behaviors across prompts, subgroups, or hypotheses; in social science surveys, they may correspond to related questions, populations, or measurement conditions. Prediction-powered inference (PPI) uses abundant but inexpensive proxy measurements to improve inference from limited, ground-truth labels, but commonly used methods treat tasks independently and therefore fail to exploit shared structure across related tasks. This limitation is especially important in settings where only a small number of labels are available per task. To address this issue, we introduce a multi-task prediction-powered inference framework that uses labeled data from related tasks to improve power while preserving task-specific inference. Our methods exploit the shared structure in the proxy-ground-truth relationship through cross-task recalibration, while retaining within-task rectification and power tuning to construct accurate point estimates and confidence intervals. We prove that efficiency gains beyond power-tuned PPI are only possible when the proxy-ground-truth relationship contains nonlinear structure; affine cross-task recalibrations are asymptotically equivalent to using the original proxy. We complement our theoretical findings with experiments on synthetic and semi-synthetic datasets, as well as a case study auditing language models on election-related information during the 2024 U.S. presidential election. Using a large human-annotation study, we show that cross-task recalibration can substantially reduce confidence interval widths when labels are scarce.
Nicolas Emmenegger, Ellery Stahler, Chara Podimata
Prediction-Powered Inference (PPI) is a popular strategy for combining gold-standard and possibly noisy pseudo-labels to perform statistical estimation. Prior work has shown an asymptotic \enquote{free lunch} for PPI++, an adaptive form of PPI, showing that the \textit{asymptotic} variance of PPI++ is always less than or equal to the variance obtained from using gold-standard labels alone. Notably, this result holds \textit{regardless of the quality of the pseudo-labels}. In this work, we demystify this result by conducting an exact finite-sample analysis of the estimation error of PPI++ on the mean estimation problem. We give a \enquote{no free lunch} result, characterizing the settings (and sample sizes) where PPI++ has provably worse estimation error than using gold-standard labels alone. Specifically, PPI++ will outperform if and only if the correlation between pseudo- and gold-standard is above a certain level that depends on the number of labeled samples (n). In some cases our results simplify considerably: For Gaussian data, for instance, the correlation must be at least 1/n−2 in order to see improvement. More broadly, by providing exact non-asymptotic expressions for the variance of PPI++ under sample splitting, we aim to empower practitioners to transparently reason about the benefits of PPI++ in specific applications. In experiments, we illustrate that our theoretical findings hold on real-world datasets.
Pranav Mani, Peng Xu, Zachary C. Lipton +1
Abridge AI, San Francisco, CA, USA · Machine Learning Department, Carnegie Mellon University, Pittsburgh, PA, USA · Department of Computer Science, Johns Hopkins University, Baltimore, MD, USA
Post-deployment monitoring of healthcare AI requires statistically valid, label-efficient methods, but gold-standard labels from clinician chart review are expensive. Prediction-powered inference (PPI) and active statistical inference (ASI) reduce label cost by combining a small labeled sample with abundant model predictions, but both are restricted to a single predictor, a poor fit for modern clinical pipelines that have multiple predictors of differing cost and accuracy available at inference time. We propose Active Multiple-Prediction-Powered Inference (AM-PPI), which routes each instance to a cost-appropriate predictor subset, samples gold-standard labels in proportion to the chosen subset's residual uncertainty, and reweights predictions to minimize estimator variance, all under a single deployment-time budget. AM-PPI generalizes ASI to leverage multiple predictors and extends Multiple-PPI from global per-predictor allocation to per-instance adaptive routing. We derive closed-form Karush-Kuhn-Tucker (KKT) conditions for all three decisions and prove, via biconvexity and strong duality, that the resulting fixed point is a global optimum despite the joint problem being non-jointly-convex. We establish asymptotic normality with valid coverage, minimum-variance unbiasedness within the linear-prediction augmented inverse propensity weighted (AIPW) class, and a closed-form criterion identifying when multiple predictors help. On synthetic data and three healthcare monitoring tasks, AM-PPI produces 10 to 40 percent narrower confidence intervals (CIs) than single-predictor ASI in the budget regime where routing matters, and matches the better baseline elsewhere.
Nicholas Brawand, Nima Leclerc, Anhthy Ngo +4
The MITRE Corporation · Georgia Institute of Technology