Signals that predict whether a chain-of-thought (CoT) trace is correct are compared by AUC, but deploying one requires a threshold with a guarantee. We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces. The central observation is validity by abstention: an (α,δ)-valid procedure that issues a certificate with probability Pfire bounds the failure probability of an issued certificate only by δ/Pfire, so a certificate that rarely fires can be valid and wrong every time it is used. In a simulation with known risk the standard certificate fails in at most 0.3% of calibration draws but in up to 69% of those in which it fires. A certification floor and a lattice condition for Benjamini-Hochberg conformal selection explain why certificates abstain at these budgets, and the data bear them out: the standard certificate returns nothing or a large accepted set, and an unreadable residual-stream probe buys two to three times the coverage of the readable signals, an edge a cross-fitted reconstruction cannot recover linearly from the readable features. We then give a floor-started fixed-sequence certificate, valid without monotonicity assumptions, that covers more than the Bonferroni certificate on every model-signal pair and raises coverage at the non-vacuous target 0.75π0 from 0.05 to 0.16, although the floor keeps absolute coverage small. Finally, a certificate cannot see what matters after deployment: under benchmark shift the error among accepted traces tracks the new task's base error, and under best-of-n selection against the verifier it rises past the target while the empirical failure frequency stays below δ, because abstention absorbs the failures.
Figures & tables
Figure 1: Certification floor (Lemma 1 ): minimum number of calibration traces above the threshold for the CP certificate to certify α ( δ=0.10 , G=20 ) when a fraction ε of them are errors. Shaded: calibration-set sizes in this study (67–158). Analytic; no data.
Confidence
J-lens
Probe
α=π^0
0.75π^0
α=π^0
0.75π^0
α=π^0
0.75π^0
Model
ncal
π^0
π0held
Cov.
Pfire
R^/α
Cov.
Pfire
Cov.
Pfire
R^/α
Cov.
Pfire
Cov.
Pfire
R^/α
Cov.
Pfire
Qwen3.5-0.8B
158
0.64
0.68
0.22
0.578
0.83
0.01
0.045
0.08
0.251
0.91
0.00
0.013
0.49
0.963
0.79
0.06
0.380
Qwen3.5-4B
154
0.16
0.16
0.39
0.571
0.63
0.11
0.220
0.18
0.239
0.79
0.02
0.045
0.72
0.934
0.53
0.39
0.618
OLMo-3-7B
67
0.34
0.34
0.33
0.533
0.67
0.10
0.192
0.08
0.160
0.90
0.01
0.021
0.34
0.538
0.68
0.09
0.208
Qwen3-1.7B
78
0.21
0.22
0.09
0.115
0.92
0.01
0.012
0.04
0.054
1.07
0.00
0.006
0.21
0.293
0.72
0.05
0.094
Table 1: Rule ( 1 ) on GSM8K, δ=0.10 , 1,000 replicates per cell. ncal : calibration problems; π^0 : training-partition base error (sets the targets); π0held : base error of the remaining problems. Per signal: mean coverage, Pfire and R^/α (error among accepted test traces pooled over firing replicates, relative to the target; – if fewer than 50 replicates fire) at α=π^0 , and coverage and Pfire at 0.75π^0 . Test AUCs: Appendix Table 6 .
CP, Eq. ( 1 )
FS-floor
AUC
n
α
Cov.
Pfire
c
Cov.
Pfire
Pfail
0.82
154
0.75π0
0.321
0.55
4
0.635
0.85
0.053
0.82
154
0.5π0
0.051
0.09
2
0.272
0.49
0.039
0.71
154
0.75π0
0.052
0.11
3
0.231
0.40
0.040
0.71
154
0.5π0
0.003
0.01
1.5
0.045
0.12
0.022
0.76
78
0.75π0
0.061
0.12
3
0.262
0.40
0.043
Table 2: FS-floor (cross-fitted grid, c chosen before calibration by simulation) against the Bonferroni CP rule ( 1 ), δ=0.10 . Top: Gaussian score model with known risk at the given AUC and calibration size n , 20,000 draws; Pfail=P(R(λ^)>α) . Bottom: means over 7 models × 3 signals, 1,000 replicates each; ∗ median of the per-cell choices of c .
Figure 2: Calibration budget curves in the Gaussian score model ( π0=0.16 , α=0.75π0 , δ=0.10 , 2,000 draws per point): firing probability and mean coverage of FS-floor ( c=2 , solid) and rule ( 1 ) (dashed) against the number of calibration problems, for four values of the verifier’s AUC. Shaded: the budgets in this study.
Figure 3: Best-of- n selection against each verifier signal. Certificate panels : error among accepted BoN-selected test traces, pooled over firing replicates, relative to α=π0 ; thresholds calibrated on unpressured traces. Signals that fire in fewer than 2% of 1,000 replicates are omitted. Accuracy panels : change in accuracy of the selected trace relative to n=1 , with 95% bootstrap bands over problems.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Held-out AUC
Coverage, α=π0
Model
Probe
Readable
Recon.
Resid.
Resid. (in-sample)
Probe
Readable
Recon.
Qwen3.5-0.8B
0.812
0.700
0.738
0.721
0.621
0.714
0.485
0.590
Qwen3.5-4B
0.846
0.719
0.743
0.694
0.614
0.774
0.516
0.566
OLMo-3-7B
0.758
0.712
0.707
0.649
0.513
0.346
0.205
0.189
Qwen3-1.7B
0.767
0.618
0.567
0.697
0.539
0.217
0.067
0.016
Qwen2.5-7B-It
0.650
0.680
0.554
0.595
0.440
0.182
0.209
0.143
Appendix
Table 3: Cross-fitted probe reconstruction. Held-out AUC (mean over 5 train/held-out splits by problem) of the probe, the readable logistic regression, a ridge reconstruction of the probe score from the 119 readable features (fit on out-of-fold probe scores of the training problems), and the residual probe − reconstruction; all regularization constants chosen by grouped cross-validation on the training problems. “Resid. (in-sample)”: the residual when the ridge is instead fit on in-sample probe scores, which are nearly separated. Right: mean coverage of the CP certificate at α=π0 on held-out problems with each score as the signal.
Figure 4: Mean coverage at the three relative targets for each model and principal signal on GSM8K: rule ( 1 ) (dashed, open markers) and FS-floor with the cross-fitted grid and simulation-chosen c (solid). FS-floor covers more in all 21 pairs at all three targets, and the gap is largest where the certificate fires at all.
Figure 5: Validity by abstention in the Gaussian score model (the settings of Tables 2 and 7 ). Each point is one procedure in one setting: the probability that an issued certificate is wrong, against the firing probability. Proposition 3 bounds the former by min{1,δ/Pfire} (dotted). Rule ( 1 ) sits far below δ where it fires often and far above it where it rarely fires; FS-floor sits near δ ; the data-dependent walk of § 4.1 has no bound.
Figure 6: GSM8K-calibrated certificates on other benchmarks (Section 6 ; every cell in Table 10 ): excess of the error among accepted traces over the target, against how much harder the new task is. Filled: mean coverage above 2% . On the dotted diagonal the accepted set is no better than accepting everything.
Figure 7: Best-of- n : mean coverage of the naive arm (solid; threshold calibrated on unpressured traces) and of the selection-conditional arm (dashed; calibration traces selected by the same best-of- n rule), for the signals that fire in at least 2% of replicates. The correction restores exchangeability, but its coverage collapses with n except for the probe on Qwen3.5-4B, where best-of- n lowers the policy’s error below the target.
Bound
Mean cov.
Cov. ≥ 10%
Beats CP
Clopper–Pearson
0.211
14/21
–
Hoeffding–Bentkus
0.156
12/21
0/21
Betting mixture
0.087
6/21
0/21
Data-dep. walk †
0.435
21/21
–
Appendix
Table 4: Bounds in the same selective construction (same grid, Bonferroni split, data and replicates) at α=π0 ; 7 models × 3 signals. † No validity proof (see text).
Model
Target
α
Non-vac.
Confidence
J-lens
Logit lens
Surface
Probe
Qwen3.5-0.8B
0.5π0
0.322
yes
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.002 (0.02)
0.75π0
0.483
yes
0.007 (0.05)
0.001 (0.01)
0.000 (0.00)
0.000 (0.00)
0.063 (0.38)
π0
0.644
yes
0.218 (0.58)
0.080 (0.25)
0.051 (0.13)
0.000 (0.00)
0.489 (0.96)
0.05
0.050
yes
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.10
0.100
yes
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.20
0.200
yes
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
0.000 (0.00)
Appendix
Table 5: Mean coverage of the CP certificate (in parentheses: Pfire ) for every model, target and signal; GSM8K, δ=0.10 , 1,000 replicates. “Non-vac.”: target below the base error rate of the non-training problems.
Confidence
J-lens
Probe
BH: no rejection
Model
AUC
excl.
incl.
AUC
excl.
incl.
AUC
excl.
incl.
n0
min. rej.
Conf.
J-lens
Probe
Qwen3.5-0.8B
0.70
0.218
0.213
0.66
0.080
0.077
0.81
0.489
0.486
107.8
2/119
0.09
0.21
0.01
Qwen3.5-4B
0.73
0.389
0.390
0.65
0.176
0.172
0.82
0.721
0.715
23.4
31/116
0.87
0.97
0.57
OLMo-3-7B
0.75
0.326
0.313
0.63
0.081
0.081
0.76
0.338
0.323
23.4
6/50
0.41
0.69
0.28
Qwen3-1.7B
0.66
0.094
0.088
0.61
0.042
0.047
0.76
0.212
0.216
17.9
15/58
0.96
0.96
0.76
Qwen2.5-7B-It
0.60
0.149
0.160
0.62
0.156
0.157
0.68
0.232
0.232
12.9
18/54
0.98
0.90
0.91
Appendix
Table 6: Left: test AUC of the frozen scorer, and coverage of rule ( 1 ) (“excl.”, boundary trace excluded) and of the boundary-included variant (“incl.”), α=π0 . Right: BH conformal selection at level α=π0 : mean number of calibration errors n0 , the lattice minimum ⌈m/(α(n0+1))⌉ of Proposition 2 relative to the test size m , and the fraction of replicates with no rejection.
π0
AUC
ncal
α
Procedure
Pfire
P(R>α)
P(R>α∣fire)
med. R/α
p90 R/α
Mean cov.
0.16
0.82
154
0.75π0
CP, Eq. (1)
0.558
0.0014
0.003
1.06
1.13
0.322
0.16
0.82
154
0.75π0
Data-dep. walk
0.985
0.0319
0.032
1.06
1.24
0.631
0.16
0.82
154
0.5π0
CP, Eq. (1)
0.089
0.0019
0.021
1.04
1.15
0.049
0.16
0.82
154
0.5π0
Data-dep. walk
0.770
0.0368
0.048
1.11
1.34
0.331
0.16
0.71
154
0.75π0
CP, Eq. (1)
0.112
0.0018
0.016
1.04
1.14
0.054
0.16
0.71
154
0.75π0
Data-dep. walk
0.724
0.0374
0.052
1.09
1.20
0.306
Appendix
Table 7: Gaussian score model with known selective risk; 20,000 calibration draws per row, δ=0.10 . P(R>α) is the quantity bounded by (α,δ) -validity; P(R>α∣fire) is the probability that an issued certificate is wrong; the next two columns give the median and 90th percentile of R(λ^)/α among failing draws, i.e. how far a wrong certificate overshoots its target.
α=π0
α=0.75π0
Model
Signal
CP
In-s.
c=1
1.5
2
3
4
Sim.
CP
In-s.
c=1
1.5
2
3
4
Sim.
Qwen3.5-0.8B
Conf.
0.218
0.113
0.037
0.134
0.134
0.217
0.291
0.291
0.007
0.023
0.012
0.012
0.021
0.034
0.043
0.045
J-lens
0.080
0.076
0.007
0.030
0.030
0.094
0.134
0.134
0.001
0.011
0.002
0.002
0.007
0.011
0.013
0.013
Probe
0.489
0.278
0.319
0.502
0.502
0.616
0.652
0.652
0.063
0.064
0.108
0.108
0.151
0.180
0.219
0.219
Qwen3.5-4B
Conf.
0.389
0.466
0.241
0.367
0.464
0.650
0.701
0.701
0.113
0.280
0.131
0.194
0.253
0.366
0.313
0.365
J-lens
0.176
0.223
0.018
0.100
0.166
0.269
0.336
0.336
0.024
0.074
0.014
0.045
0.067
0.091
0.097
0.098
Appendix
Table 8: FS-floor on real data, mean coverage for each (model, signal) pair; GSM8K, δ=0.10 , 1,000 replicates. CP: rule ( 1 ). In-s.: FS-floor with the grid from in-sample training scores, c=2 . c=1,…,4 : cross-fitted grid at fixed c . Sim.: cross-fitted grid with c chosen per cell by the pre-calibration simulation.
Rule ( 1 )
FS-floor wins vs.
Predicted Pfire
Target
G=20
G=10
G=5
Floor start
FS-floor
G=5
Floor start
Spearman
MAE
π0
0.211 / 0.33
0.242 / 0.37
0.259 / 0.40
0.186 / 0.50
0.390 / 0.50
21/21
21/21
0.72
0.154
0.75π0
0.051 / 0.10
0.067 / 0.13
0.072 / 0.14
0.101 / 0.28
0.163 / 0.28
21/21
21/21
0.70
0.104
0.5π0
0.004 / 0.01
0.007 / 0.01
0.010 / 0.02
0.024 / 0.07
0.033 / 0.07
20/21
21/21
0.72
0.046
Appendix
Table 9: FS-floor against two cheap baselines on the real cells (means over 7 models × 3 signals, 1,000 replicates; each entry is mean coverage / Pfire ). G : number of Bonferroni grid points in rule ( 1 ) ( G=20 in the paper). Floor start: the single fixed threshold at j0 tested at level δ , i.e. FS-floor without the walk. Wins: pairs where FS-floor covers more than the baseline. The last two columns compare the Pfire predicted before calibration by the simulation that chooses c with the realized Pfire across the 21 pairs (Spearman correlation and mean absolute error).
Confidence
J-lens
Probe
Model
Benchmark
πnew
α
Cov.
R^
Cov.
R^
Cov.
R^
Qwen3.5-0.8B
ARC-C
0.77
0.64
0.00
0.61
0.06
0.83
0.27
0.75
BBH-logic
0.61
0.64
0.07
0.55
0.05
0.59
0.09
0.60
CSQA
0.79
0.64
0.00
0.01
0.05
0.86
0.27
0.76
StrategyQA
0.47
0.64
0.00
0.49
0.10
0.47
0.13
0.44
Qwen3.5-4B
ARC-C
0.07
0.16
0.11
0.02
0.27
0.07
0.60
0.06
Appendix
Table 10: Every shift cell (Section 6 ): GSM8K-trained scorer and GSM8K-calibrated threshold ( α : training-partition base error on GSM8K) applied to the same model’s traces on another benchmark. πnew : base error rate on the new benchmark. Per signal: mean coverage and R^ , the error among accepted traces pooled over firing replicates; bold where R^>α . 1,000 replicates per cell.
R^/α , naive
Corrected cov.
Δ acc
Model
Signal
Pfire
n=1
n=8
n=32
n=1
n=32
n=32
from peak
Qwen3.5-4B
confidence
0.031
0.92
1.14
1.25
0.019
0.002
-0.004
-0.022
J-lens
0.015
–
–
–
0.008
0.000
+0.001
-0.016
probe
0.136
0.70
0.89
0.69
0.087
0.222
+0.038
+0.000
logit lens
0.014
–
–
–
0.008
0.000
-0.017
-0.032
surface
0.002
–
–
–
0.001
0.000
-0.040
-0.040
Appendix
Table 11: Best-of- n selection (Section 7 ) per model and signal. Pfire : firing probability of the naive arm (independent of n ). R^/α : error among accepted BoN-selected test traces relative to the target, naive arm (– where fewer than 2% of replicates fire). Corrected: mean coverage of the selection-conditional arm at n=1 and n=32 . Δ acc: change in accuracy of the selected trace from n=1 to n=32 , and from its peak over n to n=32 .
Model
Traces
Acc.
Conf.
J-lens
Logit lens
Surface
Random
Probe
J − Conf.
J − Surf.
Probe − J
Qwen3.5-0.8B
1244
0.33
0.714
0.713
0.661
0.531
0.641
0.842
-0.001
+0.182 ∗
+0.129 ∗
Qwen3.5-4B
1365
0.84
0.715
0.707
0.669
0.515
0.553
0.847
-0.009
+0.191 ∗
+0.140 ∗
OLMo-3-7B
295
0.66
0.748
0.642
0.595
0.613
0.594
0.770
-0.106 ∗
+0.029
+0.128 ∗
Qwen3-1.7B
373
0.78
0.675
0.679
0.635
0.612
0.604
0.810
+0.005
+0.067
+0.131 ∗
Qwen2.5-7B-It
323
0.81
0.584
0.648
0.522
0.602
0.559
0.668
+0.065
+0.046
+0.020
Qwen3-4B
1387
0.91
0.610
0.607
0.551
0.549
0.625
0.751
-0.003
+0.057
+0.144 ∗
Appendix
Table 12: GSM8K out-of-fold AUC (5-fold, grouped by problem, logistic regression) of each signal, and paired bootstrap differences ( ∗ : two-sided p<0.05 , uncorrected). “Random”: J-lens features with a norm-matched random transport. These are cross-validated AUCs on all traces and differ from the frozen-scorer test AUCs of Table 6 .
Department of Mathematics and Statistics McGill University Montréal, QC, Canada · Department of Mathematics and Industrial Engineering Polytechnique de Montréal Montréal, QC, Canada