Prediction-powered inference (PPI) and its power-tuned extension (PPI++) improve confidence intervals by combining a small gold-standard labeled sample with a large AI model's predictions. Its efficiency gain relies on low residual variance, which may not hold if the predictor is pre-trained on a different source domain. We propose Transfer Calibrated Prediction-Powered Inference (TC-PPI), adapting the source-domain predictor to the target domain using gold-standard samples through cross-fitting. This approach supports various adaptation methods, such as sparse linear calibration, LoRA, and fine-tuning. Our jointly tuned cross-fit estimator, Joint-TC-Cross-PPI++, maintains unbiasedness and is simultaneously at least as efficient as classical inference, PPI, and PPI++, thereby protecting against negative transfer. We provide high-dimensional MSE bounds for calibration and show empirical improvements over baseline methods across various real-world applications.
Figures & tables
Tabular (Mean)
Speech (WER)
Protein Language Models (Mean Mutational Fitness)
Method
Census Income n=200
Capital Bikeshare n=100
AfriSpeech-200 (Yoruba), n=500
AutoEval-ProteinGym n=500
Classical
(1.000, 1.000)
(1.000, 1.000)
(1.000, 1.000)
(1.000, 1.000)
Joint TC-Cross-PPI++
(0.581, 0.798 )
( 0.139 , 0.334 )
(0.890, 0.946 )
( 0.605 , 0.806 )
PPI++
(0.613, 0.803)
(0.211, 0.460)
(0.930, 0.963)
(0.695, 0.850)
TC-Cross-PPI++
( 0.578 , 0.806)
(0.140, 0.335)
( 0.886 , 0.950)
(0.607, 0.809)
WILDS (Mean)
LLM Arena (Bradley–Terry)
Table 1: Representative small-label results for different datasets and the inferential target in parentheses. MSE and interval width are normalized by classical inference at the same label budget.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Application
Population ( N+n )
Gold labels ( n )
Repetitions
K
Census mean / OLS
380,091
100–2000
200
5
Census healthcare
318,215
100–2000
100
5
Capital Bikeshare
8,734 2 2 2 Target-year (2012) population. The 8,645 source-year (2011) rows are used only to fit the source predictor.
50–500
100
5
ASR accent transfer
15,378
50–1000
250
5
Camelyon17 (WILDS)
119,958 3 3 3 Pooled held-out-hospital population ( 85,054 test plus 34,904 validation patches). The 335,996 training-hospital patches, including the in-distribution validation split, are used only to fit and select the source predictor.
25–500
200
5
PovertyMap (WILDS)
7,872 4 4 4 Fold A’s pooled OOD-country population (3,963 test plus 3,909 validation clusters). The 10,796 training- and in-distribution-validation-country clusters are used only to fit and select the source predictor.
25–500
200
5
Appendix
Table 2: Population size ( N+n ), gold-label range ( n ), repetition count, and cross-fitting fold count ( K ). N rows are unlabeled in each repetition. Parenthesized values give the LoRA variant’s count where it differs from the sparse-calibration default shown (if applicable).
Application
r
α
Dropout
Target modules
Epochs
LR
Batch (train/predict)
Max length
ProteinGym (BLAT)
1
16
0.05
query, value
100
10−3
256 / 256
—
AutoEval-ProteinGym
2
16
0.05
query, value
100
10−3
512 / 512
—
Chatbot Arena
2
16
0.05
c_attn
100
10−3
16 / 64
1024
AutoEval MT-Bench Arena
2
16
0.05
c_attn
100
10−3
16 / 64
1024
Appendix
Table 3: LoRA adapter hyperparameters by application. Target modules follow each backbone’s attention-layer naming ( query / value for ESM-2, c_attn for GPT-2-family models). Both arena applications share distilgpt2 as backbone.
Application
Smallest n
Largest n
Census Income
1.03 ( n=100 )
0.96 ( n=2000 )
Capital Bikeshare
0.61 ( n=50 )
0.41 ( n=500 )
Camelyon17 (WILDS)
1.07 ( n=25 )
0.60 ( n=500 )
PovertyMap (WILDS)
1.97 ( n=25 )
0.79 ( n=500 )
ProteinGym
0.04 ( n=100 )
0.03 ( n=1000 )
AfriSpeech-200 (Yoruba)
1.07 ( n=50 )
0.97 ( n=1000 )
Appendix
Table 4: Residual variance of the transfer-calibrated predictor relative to the source predictor, Var(Y−f^t)/Var(Y−fs) , on the gold sample and averaged over trials, at the smallest and largest label budgets. Values below one indicate calibration reduced the residual variance.
Application
n
Method
(MSE, CI width) ratios
ProteinGym (BLAT)
1000
Classical
(1.000, 1.000)
TC-Cross-PPI++
( 0.854 , 0.886 )
TC-Cross-PPI++-LoRA
(0.904, 0.937)
AutoEval-ProteinGym
500
Classical
(1.000, 1.000)
TC-Cross-PPI++
(0.423, 0.822)
TC-Cross-PPI++-LoRA
( 0.223 , 0.746 )
Appendix
Table 5: Sparse-linear versus LoRA calibration at one representative n per application (Figure 7 shows the full ProteinGym/AutoEval-ProteinGym sweep over n ; the AutoEval MT-Bench Arena sweep over n=100 – 750 is summarized in the text). Each cell reports (MSE ratio, CI width ratio), both normalized by classical inference. Bold marks the better of the two calibrated methods, independently for each of the two ratios.
Department of Statistics, University of Washington · Department of Statistics and Division of Biostatistics, Center for Targeted Machine Learning and Causal Inference, University of California, Berkeley