Automatic relevance determination (ARD), the default tool for variable selection in Gaussian-process (GP) regression, ranks inputs by inverse lengthscales -- which measure how fast a function varies, not how much an input contributes to prediction -- and offers no calibrated rule for deciding which inputs to keep. The prediction-centred alternative, the derivative sensitivity νj=E[(∂f/∂xj)2], is available in closed form from a fitted GP, but turning it into a selection rule is harder than it looks: at a null input the estimator is a degenerate quadratic form, so Wald and Bernstein-von Mises cutoffs are anti-conservative, and the natural residual bootstrap is mis-scaled. We show that a studentized multiplier bootstrap of the GP derivative process repairs both, prove its validity through an invariance principle for quadratic forms, and obtain asymptotic family-wise and false-discovery-rate control across inputs. Over 100 replications the rule controls FDR wherever inputs are truly null, while uncalibrated derivative rankings breach the target by up to 2x and a Bernstein-von Mises cutoff by 2.2x; at matched FDR it loses no power; it holds under a Matérn kernel and input correlation up to 0.99; on real data with planted and authentic null inputs it admits 5-12x fewer spurious inputs; it costs 5-18% of the GP fit; and a block-averaged variant retains validity at cost linear in n.
Figures & tables
GP
νj closed form
test νj=0
Null/Boot.
Sim.
Wycoff et al. (2021)
✓
✓
–
–
–
De Lozzo and Marrel (2016)
✓
✓
per-input
–
–
Deng et al. (2022)
✓
✓
BvM interval
–
–
Liu et al. (2023)
KRR
–
per-derivative
boot.
–
Paananen et al. (2019)
✓
ranking
–
–
–
Dai et al. (2023)
additive
–
knockoffs
–
FDR
Table 1: Closest prior work. “Null”: analysis of the degenerate null; “Boot.”: a valid bootstrap for it; “Sim.”: simultaneous FWER/FDR control over coordinates.
DPS-BY (ours)
Paananen–KL (uncal.)
ARD elbow
Benchmark
FDR
Power
FDR
Power
FDR
Power
Friedman
0.063 (.010)
1.00 (.00)
0.194 (.015)
1.00 (.00)
0.175 (.014)
1.00 (.00)
misranking
0.110 (.019)
0.85 (.02)
0.414 (.025)
0.91 (.02)
0.284 (.028)
0.86 (.02)
borehole
0.031 (.007)
0.66 (.01)
0.097 (.009)
0.74 (.01)
0.088 (.009)
0.73 (.01)
Table 2: FDR / power at q=0.2 , 100 replications (Monte-Carlo SE in parentheses; n=300 ). Boldface: breach of q . DPS-BY controls FDR throughout; the uncalibrated rankings, including faithful Paananen–KL, breach by 2× on misranking.
Friedman
misranking
Rule
FDR
Power
FDR
Power
Wald/BvM + BY
0.30 (.01)
1.00
0.45 (.03)
0.92
Permutation + BY
0.03 (.01)
1.00
0.08 (.02)
0.67
DPS-BY (ours)
0.06 (.01)
1.00
0.11 (.02)
0.85
Table 3: The obvious routes to a cutoff, 100 replications, q=0.2 (MC SE in parentheses). The Wald/BvM rule is anti-conservative for the structural reason in Proposition 2 ; the permutation test is valid but underpowered and 4× costlier. DPS-BY from Table 2 for reference.