Weighted split conformal prediction reweights calibration scores by the likelihood ratio between the test and training covariate distributions and guarantees marginal coverage under covariate shift. We study its coverage conditional on the calibration data. An elementary argument, based on a single concentration inequality at a fixed population quantile, gives explicit training-conditional bounds without unspecified constants, and shows that the relevant scale is not the supremum of the likelihood ratio but a variance proxy built from the chi-squared divergence of the shift and from the average of the ratio over the part of the test population, of probability equal to the miscoverage level, where it is largest. A two-point lower bound shows that the root-m rate and the chi-squared contribution are intrinsic to the shift. Run at an explicitly inflated level, the weighted quantile becomes a deterministic PAC prediction set. We compare it with randomized rejection sampling and with importance-weighted learn-then-test and, through a certified choice of a clipping level for the likelihood ratio, map the regime in which each gives the narrower valid set. The analysis extends to estimated likelihood ratios and to tail functionals estimated from an unlabeled source sample, which yields a fully finite-sample certificate.
Figures & tables
(a) fraction of calibration sets with Pe>α
B
CVaRα
none
WQ
WQ-Bk
WQ-Bk clip
RS+CP
RS+CP clip
IW-HB
IW-EB
1.00
1.00
0.528
0.002
0.012
0.012
0.047
0.047
0.011
0.000
2.31
2.26
0.501
0.001
0.005
0.005
0.039
0.045
0.004
0.000
4.07
3.97
0.479
0.000
0.001
0.001
0.037
0.026
0.001
0.000
6.01
5.86
0.489
0.000
0.000
0.000
0.027
0.027
0.000
0.000
Table 1: Smooth shift, m=2000 , α=δ=0.05 , 1000 replications, exact population functionals. Widths are medians of 2τ relative to the uninflated set; clip levels in parentheses. WQ: tail-adaptive Bernstein inflation; WQ-Bk: Bentkus inflation; RS+CP: rejection sampling with Clopper–Pearson; IW-HB, IW-EB: importance-weighted learn-then-test with Hoeffding–Bentkus and empirical Bernstein p-values (Section 6.4 ).
estimator
nw
B
∥v−v∥L1
TV(Q,Q)
Δ
γm(δ)
fail (none)
fail (Bern.)
med. cov. (Bern.)
logistic
250
6.16
0.121
0.0579
0.0028
0.055
0.492
0.000
0.9554
logistic
1000
6.47
0.117
0.0580
0.0021
0.057
0.446
0.000
0.9577
logistic
4000
4.47
0.041
0.0202
0.0047
0.045
0.568
0.000
0.9433
kde
1000
36.47
0.686
0.1888
0.0032
0.196
0.522
0.000
1.0000
Table 2: Estimated weights (smooth shift, c=2 , m=2000 , α=0.1 , δ=0.05 , 500 replications). The Bernstein inflation uses B from the estimate; failures are counted at the nominal α (not at α+Δ ).
(a) fraction of calibration sets with Pe>α
dataset
tilt
B
CVaRα
none
WQ
WQ-Bk
WQ-Bk clip
RS+CP
RS+CP clip
IW-HB
IW-EB
Wine
original
57.6
30.8
0.532
0.000
0.000
0.000
0.024
0.004
0.000
0.000
Wine
B=4
4.0
2.0
0.496
0.000
0.001
0.001
0.043
0.028
0.000
0.000
Abalone
original
43.9
20.6
0.513
0.000
0.000
0.000
0.027
0.003
0.000
0.000
Abalone
B=4
4.0
2.5
0.474
0.000
0.002
0.002
0.038
0.042
0.000
0.000
Concrete
original
7.1
6.3
0.482
0.000
0.001
0.001
0.030
0.026
0.001
0.000
Table 3: UCI datasets with exponential-tilt covariate shift, α=0.1 , δ=0.05 , 1000 calibration draws from the pool; procedures and conventions as in Table 1 ; sample sizes in Appendix D . “ ∞ ” means that the set is all of Y for more than half of the draws.
shift
δ
m
CVaRα/B
WQ
WQ-Bk
RS
RS clip ( B′ )
WQ clip ( B′ )
WQ-Bk clip ( B′ )
CP-C
CP+TV
IW-HB
IW-EB
emp. oracle
rare, π=0.01 , B=21
0.05
2000
0.25
1.228
1.192
1.368
1.118 (1)
1.145 (4)
1.135 (5)
1.991
1.117
1.566
∞
1.059
rare, π=0.01 , B=21
0.05
5000
0.25
1.120
1.104
1.184
1.093 (1)
1.103 (8)
1.095 (10)
1.883
1.092
1.270
1.435 2%
1.039
rare, π=0.01 , B=21
0.05
20000
0.25
1.053
1.048
1.078
1.076 (21)
1.053 (21)
1.048 (21)
1.818
1.072
1.110
1.094
1.019
rare, π=0.01 , B=21
0.01
2000
0.25
1.331
1.272
1.715 27%
1.142 (1)
1.175 (3)
1.161 (3)
2.182
1.145
∞
∞
1.076
rare, π=0.01 , B=21
0.01
5000
0.25
1.158
1.139
1.275
1.106 (1)
1.117 (5)
1.110 (6)
1.960
1.106
1.367
∞
1.052
rare, π=0.01 , B=21
0.01
20000
0.25
1.068
1.061
1.114
1.080 (1)
1.068 (21)
1.061 (21)
1.847
1.080
1.145
1.136
1.027
Table 4: Widths of PAC sets relative to the uninflated weighted quantile ( α=0.05 , 600 replications; every entry has empirical failure fraction ≤δ up to Monte Carlo error, standard error ≈0.01 ; widths are medians over all draws with the set Y counted as infinite width, “ ∞ ” when more than half of the draws give Y , and a superscript gives the percentage of draws giving Y when it is positive). “WQ” is the tail-adaptive weighted quantile with the Bernstein inflation and “WQ-Bk” with the Bentkus inflation, “RS” rejection sampling with Clopper–Pearson; “clip” the certified-optimally clipped version of each with its clip level B′ in parentheses; “CP-C” the source Clopper–Pearson baseline of Park et al. [15] at level α/B , “CP+TV” source Clopper–Pearson at level α−TV(Q,P) , “IW-HB” and “IW-EB” importance-weighted learn-then-test [ 1 ] with the Hoeffding–Bentkus and the empirical Bernstein p-value. The narrowest certificate in each row is in bold; the empirical oracle is not a certificate.
WQ-Bern.
RS+CP coverage over 300 reruns
RS+CP width / WQ width
shift
set
coverage
min
max
IQR
min
median
max
tilt, c=2 , B=4
0
0.961
0.944
0.969
0.0045
0.922
0.972
1.048
tilt, c=2 , B=4
1
0.962
0.951
0.964
0.0034
0.943
0.979
1.010
rare, π=0.01 , B=21
0
0.964
0.933
0.988
0.0128
0.838
1.039
1.266
rare, π=0.01 , B=21
1
0.966
0.935
0.993
0.0129
0.832
1.051
1.372
Gaussian, c=1
0
0.975
0.927
0.996
0.0167
0.756
1.018
1.406
Table 5: Rejection sampling rerun 300 times on a fixed calibration set ( m=5000 , α=δ=0.05 ; two calibration sets per shift). Coverage is training-conditional coverage; widths are relative to the tail-adaptive Bernstein-inflated weighted quantile on the same set, whose coverage is in the third column.
Table 6: Certificates with functionals estimated from an independent unlabeled source sample of size ns (Proposition 16 , halves for selection and certification, empirical Bernstein bounds, δ′=δ/5 ), α=δ=0.05 , 150 replications with a fresh source sample each; widths relative to the uninflated weighted quantile, clip levels in parentheses, widths are medians over all replications with Y counted as infinite, “ ∞ ” when more than half are Y , superscripts as in Table 4 . “RS” is rejection sampling as published, which needs only B ; the last two columns are the certificates of Table 4 with exact functionals, on the same calibration draws. Every entry has empirical failure fraction ≤δ up to Monte Carlo error (standard error ≈0.018 ); the narrowest finite-sample certificate in each row is in bold.