A mean shift between two data sources can be easy to detect but hard to remove without substantially changing their representations. We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in Rd, learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data. We derive the exact finite-sample minimax risk over all such maps, (d−k)E[1/(d+2J)] with J∼Pois(κ/2), where the budget allows deleting k directions and κ is the calibration signal-to-noise ratio. Projecting out the mean calibration difference attains it without knowing κ or the noise scale. This exposes a detection-repair gap: detecting the shift needs only κ≫d, whereas removing a fixed fraction of it at constant distortion needs κ≍d, as for estimating its direction. Standard linear concept erasers (MP, SAL, LEACE) remove the same calibration difference, so the formula gives, before fitting, exactly how much shift they leave on fresh data and how much calibration a target requires. The limit is robust: pairing keeps it exact for non-Gaussian shared content, the projection keeps its guarantee under anisotropic noise, and selective abstention cannot close the gap. On paired clinical and wearable sleep EEG, where differences between participants act as calibration noise, the formula predicts the device shift left in new participants, and more recordings per person soon stop helping. Together, these results tell whether a correction that falls short needs a better method, more recordings, or more participants.
Figures & tables
Figure 1: Detecting a shift is not enough to repair it. (a) Detection needs κ∝d and repair κ∝d , so the wedge of shifts that are detectable but not repairable widens with dimension. Detection is a test of “no shift” at the 5% significance level, and repair uses the best linear map that deletes one direction. (b) At d=256 , the squared length ∥X∥2 of the normalized mean calibration difference has expected value d+κ : noise d (grey) plus signal κ (blue). The test needs only the total to clear the dashed threshold, whereas editing leaves about the grey share of the bar (orange). The star marks the same setting in both panels. All values are exact.
Figure 2: What the frontier predicts for fitted erasers and for a real device shift. (a) Simulation at d=N=32 with θ=s2=1 : share of the shift left in the population against median distortion cost, over 256 calibrations ( ±1.96 Monte Carlo SE). No shared linear editor can go below the exact frontier (line). Methods on the line are calibration-limited: four times the budget lowers the frontier only from 0.49 to 0.44 . The orange arrow is the excess of one-step INLP predicted from its fitted kernel ( Equation 10 ), and the grey arrow is LEACE’s extra distortion at the same kernel. (b, c) BOAS: share of the shift left in held-out participants (points, ±1.96 SE over 20 random splits of the same participants, which understates the uncertainty from the choice of people) and its prediction from the calibration pool alone (lines). (b) Dotted levels are the limits as n→∞ . (c) At n=8 , the shift is detected (grey) long before the erasers remove it (blue).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Setting
Predicted
Observed
Cost-one residual
d=256 , κ=128
0.6652
0.6644±0.0009
Cost-one residual
d=256 , κ=512
0.3326
0.3325±0.0006
Eraser residual, Gaussian content
d=N=64
0.4961
0.4922±0.0067
Eraser residual, real content
N⋆=63 , target 0.50
0.5000
0.504,0.501,0.503 ( ±0.007 )
Eraser residual, real content
N⋆=190 , target 0.25
0.2495
0.251,0.247,0.248 ( ±0.004 )
INLP excess over MP
d=N=32 , k=1
0.181
0.186±0.012
Appendix
Table 1: Exact predictions against Monte Carlo estimates. Means carry ±1.96 Monte Carlo SE ( 8,192 draws for the frontier, 512 fits per eraser cell, 256 calibrations for INLP), and probabilities carry exact 95% binomial intervals. Real-content rows list the three pools at d=64 , each with one value shared by MP, SAL and LEACE. Device rows give the share of the population mean shift left in held-out participants, as mean ±1.96 SE over random splits of the same participants ( 20 for BOAS, 40 for Flex), and its prediction from the calibration pool ( Appendices N and N.1 ).
Figure 3: Detection, deletion, and certified action are different quantities. At d=256 , lines are exact formulas and points use 8,192 draws per strength. (a) Detection power, confidence-ball action probability, and the optimal cost-one expected residual. (b) The hard-budget frontier for three deletion dimensions.
Figure 4: Conditional improvement has an action-frequency cost. Exact selective frontiers at d=64 . Open circles mark the ungated p=1 limits, and diamonds the confidence-ball rule’s action probability and conditional mean residual. The dashed line is ε=0.5 . A diamond may lie above it, because the certificate controls joint bad action. The oracle curves fix κ , whereas the confidence-ball rule does not know it.
Figure 5: An average budget rewards adaptive deletion rank. Left: exact hard-budget, adjacent-rank mixing, and optimal average-budget risks at d=16 , κ=8 , with points from 131,072 shared draws. Right: at average cost four the optimal rule deletes one direction with probability 0.8 and all sixteen with probability 0.2 . Black points with exact 95% binomial intervals check these probabilities.
Figure 6: More frequent action under the same uniform joint-error guarantee. At d=256 : (a) exact action probabilities of the confidence-ball and finite-mixture rules with known scale (solid) and with the noise scale estimated from N=4 pairs (dash-dotted), Monte Carlo estimates on shared draws (points), and the power of the radial test (dotted). (b) The exact joint probability of acting and missing the target, which every rule keeps below α at all strengths. Thin lines repeat the known-scale rules at d=64 .
Figure 7: An illustration of the non-Gaussian-content extension. Lines are exact predictions, and points are fitted-map population residuals averaged over 512 calibration samples ( ±1.96 Monte Carlo SE). MP, SAL and LEACE share the residual at every fit, so each estimate is shown once. Markers distinguish the three pools of real represented content. Shift and noise are synthetic.
Erasers
Detection
Subtract Dˉ
P
n
Observed
Plug-in
At κˉ
Obs.
Pred.
Obs.
Pred.
1
1
0.960±0.006
0.967
0.969
0.20
0.07
48.28
42.27
1
8
0.858±0.013
0.874
0.885
0.23
0.12
8.97
8.33
1
128
0.690±0.033
0.757
0.783
0.23
0.13
3.94
3.78
4
1
0.900±0.010
0.904
0.905
0.27
0.16
11.57
10.57
4
8
0.653±0.016
0.663
0.669
0.51
0.57
2.19
2.08
Appendix
Table 2: BOAS, N2 epochs, d=118 , 85 participants. The eraser columns give the share of the population mean shift left in held-out participants after MP, SAL and LEACE, observed (mean ±1.96 SE over 20 random splits of the same participants) and predicted from the calibration pool by the plug-in two-level prediction and by the frontier formula at κˉ . Detection is the rejection probability of the level- 0.05 radial test. The last two columns give the separation left by subtracting Dˉ , relative to no correction.
Setting
Participants
d
RMS deviation
Max. deviation
Within 0.03
Cap
N2, d=118 (main)
85
118
0.016
0.067
95%
35
N2, no channel rule
97
118
0.021
0.064
88%
28
All stages
86
118
0.014
0.058
94%
36
N2, d=28
85
28
0.023
0.074
78%
9
N2, d=58
85
58
0.018
0.053
89%
17
N2, d=236
85
236
0.016
0.047
96%
67
Appendix
Table 3: BOAS: agreement between observed and predicted eraser residual over the 80 cells of the grid under other analysis choices. The last column is the one-participant cap θ2d/\trΣb .
d
N
k
Frontier
INLP excess
R-LACE excess
Observed ± SE
Predicted
Observed ± SE
16
8
1
0.6423
0.0683±0.0070
0.0737
0.0000±0.0000
16
8
4
0.5138
0.0036±0.0082
0.0030
0.0053±0.0079
16
16
1
0.4838
0.1236±0.0080
0.1399
0.0000±0.0000
16
16
4
0.3871
0.0083±0.0064
0.0049
−0.0088±0.0067
16
32
1
0.3216
0.1484±0.0071
0.1471
0.0000±0.0000
Appendix
Table 4: Complete controlled grid. The frontier is Rd∗(k,N) . Observed excess is the paired mean residual above MP plus random completion, with its Monte Carlo standard error over 256 calibrations. Predicted excess is the mean of gρK , which is below 2×10−8 for R-LACE in every row.
Distribution shift between training and deployment is a pervasive challenge for modern AI systems. In many cases, the target marginals of covariates and response are known or specified through population-level observations, boundary conditions, properties of simulator configurations, or alignment-time distributional constraints. Such knowledge may provide valuable side information for regression estimation. We study this problem in the multivariate linear regression setting with a stable conditional mean E[Y∣X] across source and target, and identify the hybrid-loss estimator, which jointly incorporates both target marginals, as a benchmark target-aware estimator. Its direct computation, however, requires solving a coupled nonlinear optimization that is expensive at scale. Our main contribution is to develop and evaluate two computationally tractable alternatives: a constrained moment-matching estimator and a two-stage estimator that augments ordinary least squares with a calibration step. For all three estimators, we derive and compare closed-form asymptotic mean squared errors, yielding conditions under which the tractable alternatives match or closely approximate the hybrid benchmark, and regimes in which they do not. Monte Carlo experiments across three controlled shift regimes validate the theoretical results, investigate the accuracy-runtime tradeoffs among the three estimators, and translate into guidance on estimator choice. In particular, the two-stage estimator nearly matches the hybrid benchmark in the high signal-to-noise regime at essentially no additional cost, providing theoretical grounding for empirical observations in nonlinear settings.
Zhewen Hou, Tian Zheng
Department of Statistics Columbia University New York, NY
In small-batch scientific deployments, labeled target outcomes may be too scarce for reliable shift estimation even when unlabeled target inputs are available. We address the complementary setting where the practitioner has a pre-specified label-shift correction from domain knowledge and asks whether incoming labeled outcomes support it. We show that the per-observation likelihood ratio between a label-shift-corrected predictive and the source predictive is a conditional e-value, so its running product is a nonnegative martingale and Ville's inequality yields an anytime-valid confirmation rule. The log martingale equals the cumulative negative log-predictive density (NLPD) gap between the source and the corrected predictive, converting routine model monitoring into a formal sequential test. Rejection means the incoming data support the posited correction relative to the source predictive, but it is not a precise estimate of the degree of shift. Closed forms are available for GP sources with Gaussian label-shift ratios. GP regression simulations validate Type I control, finite-sample power, miscalibration sensitivity, and the small-batch advantage of a reliable prior over label-based re-estimation.
Seungjin Choi
CROID Research and aSSIST University, Seoul, Korea.
We study the offline gap between deterministic calibration distance C and its fractional relaxation L for binary unit-weight sequences under total absolute-change cost. We sharpen the offline comparison C <= L + O(sqrt(T)) (Qiao and Zheng, 2024, Theorem 2) to the sharp worst-case order Theta(T^(1/3)). If Delta_T is the supremum of C - L over length-T inputs, then T^(1/3)/1000 <= Delta_T <= 41T^(1/3) for T >= 216. The upper bound holds for every input, while each T >= 216 has a rational lower-bound input. For every input with m distinct forecasts, C <= L + m, and the unrestricted-sample worst-case sparse order is Theta(m). For rational forecasts and accuracy, with binary-encoded multiplicities of separately assignable unit identities, a grid-free polynomial-bit-time procedure returns B <= L <= U, U - B < eta, and an exactly calibrated compact repair of cost at most U + m <= L + m + eta.
Zinan Wang, Xinhao Yang
University of Manchester · University of Southern California