Split conformal prediction uses prediction errors on held-out (calibration) data to determine how wide the prediction intervals should be. It guarantees distribution-free finite-sample coverage when these errors and the error at the target site are exchangeable. This assumption may fail under spatial dependence and nonrandom sampling geometry. Existing spatial methods use fitting residuals to remove the predictable part of spatial variation from calibration and target errors. However, the spatial variation that only the calibration residuals can predict remains in both the target and calibration errors, reducing the efficiency and stability of the interval. We address this by additionally conditioning on the calibration residuals sequentially, which scales to large networks through nearest-neighbour approximations. Under a correct working covariance and an elliptical residual law, the resulting interval has exact finite-sample coverage under any spatial design, and under further conditions it is asymptotically oracle efficient. We also bound coverage loss under covariance misspecification and develop a diagnostic that identifies regions at risk of undercoverage. In simulated data, our method produces narrower and more stable intervals than global and localized state-of-the-art alternatives. In a national PM2.5 application, it produces narrower intervals within the network and identifies regions at risk of coverage failure.
Figures & tables
Figure 1: The top panel shows the three fitting–calibration splits on fixed data in the two-cluster design. The table below reports sensitivity of naive split conformal, GSCP, LSCP (under different bandwidths h ), and the proposed approach to the fitting–calibration split. Each entry reports coverage; mean width (CVR) over 20,000 datasets for nominal 90% intervals. Under strong localization, LSCP can produce infinite intervals.
Figure 2: Distribution of rank of target score in the two-cluster example. The top row uses nearby calibration, and the bottom row uses distant calibration. The left column shows the calibration and target locations; the middle and right columns show the distribution of the rank of the target score among calibration scores under GSCP and the proposed whitened method. Dashed lines mark the expected bin height under a uniform rank distribution, and shaded regions contain ranks producing noncoverage.
μF∪C(s0)=−da⊤r+,1:n,νF∪C(s0)=d−2
Algorithm 1 Whitened spatial conformal prediction
Design
Naive
GSCP
LSCP
Whitened
Exponential, ϕ=0.5
.899/2.19(2.33)
.902/1.75(1.91)
.909/1.80(1.89)
.900/1.66(1.12)
Exponential, ϕ=1
.907/2.59(2.54)
.910/1.84(2.00)
.918/1.90(1.79)
.916/1.76(1.10)
Exponential, ϕ=2
.908/3.02(2.25)
.918/2.01(1.36)
.922/2.07(1.47)
.907/1.93(1.11)
Exponential, ϕ=6
.910/3.59(1.87)
.902/2.42(1.21)
.911/2.50(1.26)
.903/2.32(1.16)
Exponential, ϕ=8
.907/3.64(1.54)
.906/2.55(1.17)
.910/2.62(1.12)
.903/2.43(1.06)
Exponential, ϕ=30
.906/3.74(1.07)
.900/3.37(1.05)
.911/3.47(1.08)
.902/3.24(1.05)
Table 1: Efficiency and stability under random sampling. Each entry reports coverage/mean width, with the CVR from fixed-design counterparts in parentheses. Nominal coverage is 0.90 .
Figure 3: Spatial designs used to vary the predictive information contained in the calibration residuals. Blue circles denote fitting observations, orange triangles denote calibration observations, and red stars denote prediction sites. The halo designs assign progressively more observations near each target to calibration, while the lower panels represent transect sampling, boundary prediction, and a cluster monitoring network. Each title reports the available gain.
Design
Naive
GSCP
LSCP
Whitened
Available gain
Achieved gain
Halo 10
.900/3.01
.894/2.23
.907/2.30
.898/1.92
20.2
14.0
Halo 20
.900/3.03
.899/2.35
.910/2.41
.909/1.91
27.1
18.7
Halo 40
.904/3.12
.908/2.57
.916/2.62
.913/1.93
34.1
24.8
Transect
.904/2.95
.902/2.25
.912/2.30
.903/1.92
17.0
14.6
Boundary
.890/3.27
.894/2.48
.910/2.57
.899/1.97
26.5
20.3
Monitoring network
.905/3.23
.899/3.06
.917/3.14
.918/2.16
40.4
29.3
Table 2: Value of informative calibration sets. Method entries report coverage/mean width. The last two columns give the available gain and the achieved width reduction of whitening relative to GSCP, both in percent. Nominal coverage is 0.90 .
Figure 4: Designs used to distinguish nonrandom spatial assignment from failures of calibration–target comparability. Symbols are as in Figure 3 . The upper row shows disjoint-halves, quadrant, and clustered-calibration designs under a stationary process. The lower row shows distant-patch calibration, block holdout, and varying dependence range.
Design
Naive
GSCP
LSCP
Whitened
Predicted coverage
Role-dependent sampling
Disjoint halves
.897/3.41
.900/3.22
.910/3.24
.914/1.98
.908
Quadrant
.889/3.15
.895/2.73
.902/2.79
.907/1.93
.905
Clustered calibration
.903/3.21
.906/2.04
.923/2.18
.904/1.93
.901
Coverage stress tests
Distant patch
.914/3.34
.850/1.90
.907/∞
.908/1.86
.911
Table 3: Coverage beyond random sampling. Method entries report coverage/mean width. The last column gives the coverage predicted for the whitened interval from 200 separate diagnostic observations in the target region. Under distant patch, LSCP is unbounded at 25% of the targets, which we count as covered; its finite intervals have coverage/mean width .876/2.14 . Nominal coverage is 0.90 .
Scheme
Naive
GSCP
LSCP
Whitened
Random folds
.903/5.94
.887/5.16
.892/5.07
.898/4.85
Whole-state holdout
.868/5.44
.873/5.23
.873/5.10
.879/5.13
Table 4: Coverage and interval width in the PM 2.5 application. Each entry reports coverage/mean width in μg/m3 . Random-fold results are pooled across the twenty folds. Whole-state results are computed for each of the 49 regions and averaged with equal weight. Nominal coverage is 0.90 .
Figure 5: Empirical coverage and mean width by distance to the nearest fitting monitor under random folds at nominal coverage 0.90 . GSCP and LSCP reach the naive width as distance increases. Their smaller widths in the nearest quintile are accompanied by coverage seven to eight percentage points below nominal.
Tercile of κrel
Regions
Mean κrel
Estimated coverage
Realized coverage
Lowest
7
0.60
0.987
0.979
Middle
6
0.93
0.915
0.909
Highest
6
1.38
0.769
0.744
Table 5: Regional diagnostic under whole-state holdout. Regions are grouped into terciles of κrel , estimated from a diagnostic half of each region’s monitors. Coverage is measured on the disjoint evaluation half, and estimated coverage is G(ξ1−α/κrel) . Estimated and realized coverage are averaged over 20 random splits within each region and then over regions with equal weight, and regions with fewer than sixteen monitors are excluded. Nominal coverage is 0.90 .
Design
Fixed h=0.40
Selected h
Selected unbounded (%)
Median selected h
Exponential, ϕ=0.5
.909/1.80
.908/1.80
0.0
0.50
Exponential, ϕ=1
.918/1.90
.915/1.90
0.0
0.52
Exponential, ϕ=2
.922/2.07
.922/2.08
0.0
0.50
Exponential, ϕ=6
.911/2.50
.908/2.50
0.0
0.48
Exponential, ϕ=8
.910/2.62
.914/2.64
0.0
0.48
Exponential, ϕ=30
.911/3.47
.910/3.47
0.0
0.50
Table S1: LSCP with fixed bandwidth h=0.40 and selected bandwidth. Coverage/mean width are computed at nominal 0.90 . Unbounded intervals count as covering their target and are excluded from the mean width. h=0.40 is unbounded only under distant patch, at 24.8% of targets, where its entry is computed on the bounded intervals.
Factor
Ordering
Coverage
Mean width
Dense Cholesky
as supplied by the split
0.904
2.113
Dense Cholesky
coordinate sum
0.906
2.115
Dense Cholesky
maximin
0.904
2.113
Dense Cholesky
random
0.905
2.113
NNGP, K=25
as supplied by the split
0.905
2.116
Table S2: Effect of factorization and ordering in whitening. Coverage and mean width are computed at nominal 0.90 , and averaged over the nine Regime A designs.
Correlation
νt
Naive
GSCP
LSCP
Whitened
Exponential, ϕ=2
10
.899/3.27
.903/2.13
.909/2.19
.903/2.04
Exponential, ϕ=2
5
.906/3.45
.906/2.26
.915/2.34
.912/2.18
Gaussian, ϕ=6
10
.906/3.95
.901/2.10
.909/2.18
.904/2.02
Gaussian, ϕ=6
5
.902/4.17
.909/2.20
.914/2.26
.906/2.12
Table S3: Coverage/mean width at nominal coverage 0.90 under a spatial t process.
Online conformal prediction must balance fast adaptation to distribution shift against stable coverage: feedback-driven methods react quickly but become volatile, while strongly discounted Bayesian methods lag and inflate intervals at tight coverage. We introduce \textbf{State-Adaptive Bayesian Conformal Prediction (SA-BCP)}, which forms the predictive quantile as a gated convex combination of long-term temporal inertia and local spatial evidence from a kernel density estimate, controlled by a single interpretable evidence threshold K. We establish three results: (i) asymptotic marginal validity of the resulting intervals up to a gate-controlled bias that vanishes as spatial evidence accumulates (exact under recurrent states); (ii) a closed-form expression for the MSE-optimal threshold, KMSE∗=α(1−α)/MT, trading the coverage-indicator (Bernoulli) variance against the temporal structural bias MT; and (iii) a rolling-origin procedure for selecting K online -- consistent under stationarity, with O(TlogN) regret against the best fixed K and, for a segmented variant, a sublinear dynamic-regret bound under sublinearly many (BT=o(T)) threshold shifts. Across four financial-volatility and weather datasets, three target coverage levels, and eight baselines, SA-BCP attains at-or-above-nominal coverage in most settings while producing substantially sharper intervals -- up to roughly 3× lower Winkler score than discounted Bayesian CP at the tightest coverage -- and a coverage-matched audit confirms these efficiency gains are not an artifact of under-coverage. We disclose our principal limitation: a volatility-specialized CF-GARCH competitor remains more efficient on its home volatility-base series, though it does not transfer across domains.
Yu-Hsueh Fang, Chia-Yen Lee
Department of Information Management, National Taiwan University, No. 1, Sec. 4, Roosevelt Rd., Da’an Dist., Taipei, 106216, Taiwan.
Conformal prediction gives prediction intervals with finite-sample coverage when the data are exchangeable. Many time-indexed datasets are not exchangeable. They have seasons, recurring regimes, changing frequencies, or other forms of structured dependence. This paper studies a simple way to use that structure. We propose spectral adaptive conformal prediction, a method that forms weighted conformal quantiles using local spectral similarity and then updates the target miscoverage level online. The spectral weights choose calibration residuals that look relevant to the current test point. The adaptive update corrects the long-run miss rate when uncertainty changes over time. We give an approximate coverage result for the fixed spectral weighted quantile and a deterministic long-run calibration result for the adaptive update. Simulations with recurring regimes and slowly changing frequencies, together with three U.S. real-data examples, show that the hybrid method can improve on fixed spectral weighting, while also showing that spectral weighting must be monitored through effective sample size diagnostics.
Jeffery Opoku, David Banahene
aUniversity of Texas Rio Grande Valley, United States · bFlorida International University, United States
We study online conformal prediction in a partially observed panel: a new cross-section of peer outcomes is observed before each target outcome, target feedback may be intermittent or absent, and neither units nor rounds need be exchangeable. We propose Weighted Temporal Quantile Adjustment (W-TQA), which combines similarity weights learned from unit histories with an adaptive target-specific miscoverage level. We prove that neither target-peer mismatch nor coverage on unrevealed rounds is identifiable, so assumptions on cross-unit similarity and on the feedback mechanism cannot be avoided. We bound the past-conditional miscoverage in terms of this mismatch, quantify the cost of learning the weights under a profile-similarity assumption, and show that the implemented procedure, which falls back on the largest peer score, attains long-run average coverage under missing-completely-at-random feedback and a feasibility condition on the calibration panel. On synthetic panels and four real panels, including a genuinely asynchronous U.S.-to-Tokyo equity panel, W-TQA attains the highest coverage on the worst-covered target units on every panel; similarity weighting contributes most when target feedback is scarce, and temporal adaptation more as feedback accumulates.
Daohong Tu, Kay Giesecke
Department of Management Science and Engineering Stanford University