Reference-Tail Trust:Certified Probability Floors for Learned Updates Inside a Deployed Network
Abstract
Graph neural networks (GNNs) need to exploit improved message passing without surrendering control over predictions already trusted in deployment. We introduce Reference-Tail Trust (RTT), a framework that admits learned updates inside a frozen GNN and certifies the prediction actually served. RTT couples graph-based proposal states with a constrained internal optimizer: each displacement is charged for its worst-case terminal cross-entropy increase through the incumbent's remaining message-passing layers. A trajectory-validated tube and an independent checker enforce per-node probability floors, , and a call-level budget, , uniformly over labels. Calls whose adapted outputs pass certification require no separate full incumbent rollout; failed certificates trigger whole-call fallback. We derive the exact probability-floor frontier by water-filling, characterize architecture-constrained efficiency, and establish conditions under which internal propagation exploits evidence unavailable to restricted output correctors. In the reported ogbn-arxiv audit, RTT achieves nats of mean gain per call, with a one-sided 95% regression-rate upper bound of 0.95% and a 95% negative-flip upper bound of 0.51% on the uninspected part of the reserved node population. Its mean gain is 61% of a cross-fitted posterior-based frontier estimate and exceeds the strongest matched one-pass corrector by nats. Reported experiments span eight proposals, six graph-incumbent families, structural and temporal graph shifts, and molecular prediction, with additional image and tabular evaluations. RTT makes GNN adaptation a budgeted, certifiable inference decision rather than an unconditional model replacement.
Figures & tables
| Statement about the served prediction | Type | Unit | Assumptions | |
| G1 | for every node and class | deterministic | node | valid continuation enclosures and sound arithmetic; none on proposal quality or labels |
| G2 | , hence for every label vector | deterministic | call | as G1 |
| G3 | every class is capped, (Proposition 1 (iv)) | deterministic | node | as G1 |
| G4 | for the mixture, each component and the stress test | statistical | call | i.i.d. episodes of the registered generator |
| G5 | (anytime-valid); negative-flip rate | statistical | call, node | as G4; uniform inspection of reserved nodes |
| G6 | slope coverage , and | design | call | exchangeable episodes; a fixed member or selection-valid calibration |
| Policy | [95% CI] | to RTT | LB | Lat. | FLOPs | |||
|---|---|---|---|---|---|---|---|---|
| RTT, registered ( , ) | 6.5 [6.3, 6.7] | – | 12 | 0.95% | 0.51% | 6.2 | 1.42 | 1.36 |
| RTT corrector on the unused budget | 6.7 [6.5, 7.0] | 0.2 [ 0.4, 0.1] | 12 | 0.95% | 0.51% | 6.4 | 1.49 | 1.40 |
| One-pass corrector, 2-hop head rows | 5.6 [5.4, 5.8] | +0.9 [+0.8, +1.1] † | 14 | 1.07% | 0.56% | 5.3 | 1.25 | 1.22 |
| One-pass corrector, 2-hop head | 5.1 [4.8, 5.3] | +1.4 [+1.3, +1.6] | 14 | 1.07% | 0.56% | 4.8 | 1.07 | 1.05 |
| One-pass corrector, per-row head | 4.5 [4.3, 4.7] | +2.0 [+1.9, +2.2] | 13 | 1.01% | 0.54% | 4.2 | 1.02 | 1.01 |
| Adaptive mixing, (2 rollouts) | 5.1 [4.9, 5.3] | +1.4 [+1.2, +1.6] | 15 | 1.13% | 0.61% | 4.8 | 2.18 | 2.18 |
| Proposal | raw | raw | TR | corrector | RTT | RTT | RTT corrector | Lat. RTT | Lat. 2 roll. |
|---|---|---|---|---|---|---|---|---|---|
| bank prior + signed attention (ours) | 3.9 | 8.81% × | 5.8 | 5.6 | 6.5 | 0.95% | +0.9 [+0.8, +1.1] | 1.42 | 2.18 |
| GPR-GNN (adapter) | 3.2 | 6.99% × | 5.5 | 5.4 | 6.2 | 0.95% | +0.8 [+0.6, +1.0] | 1.53 | 2.29 |
| Co-GNN (adapter) | 2.7 | 8.19% × | 5.2 | 5.2 | 5.9 | 1.01% | +0.7 [+0.5, +0.9] | 1.62 | 2.38 |
| AMP (adapter) | 3.0 | 7.67% × | 5.4 | 5.3 | 6.0 | 1.01% | +0.7 [+0.5, +0.9] | 1.59 | 2.35 |
| GTrans (test-time adaptation) | 2.3 | 6.00% × | 4.3 | 4.4 | 4.9 | 0.89% | +0.5 [+0.3, +0.7] | 3.42 | 3.19 |
| Matcha (test-time adaptation) | 2.9 | 5.21% × | 4.6 | 4.6 | 5.2 | 0.95% | +0.6 [+0.4, +0.8] | 2.63 | 2.39 |
| Proposal | Native full | Native + floor | Matched TR | Matched corrector | RTT derived | |
|---|---|---|---|---|---|---|
| TSA | 8.0 | 6.5 | 6.0 | 6.2 | 7.2 | +1.0 |
| STEM | 8.4 | 6.8 | 6.3 | 6.4 | 7.4 | +1.0 |
| Quantity | TSA-derived + RTT | STEM-derived + RTT |
|---|---|---|
| Observed regressions | 12 | 10 |
| Pointwise call bound | 0.95% | 0.83% |
| First-pass / re-certified releases | 2,030 / 10 | 2,028 / 10 |
| Whole-call fallbacks / fraction | 8 / 0.39% | 10 / 0.49% |
| Per-node negative-flip criterion | ||
| Adjusted paired gain criterion | lower CI endpoint | lower CI endpoint |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | transport | tube factor | window | last depth | global factor |
|---|---|---|---|---|---|
| tanh diffusion ( ) | interval (Lemma 1 ), rank-8 | second-order envelope | exact | ||
| APPNP | exact (linear) | exact | exact | ||
| GCNII | interval Clarke (Lemma 2 ) | first-order enclosure | exact | ||
| GRAND | smooth diffusion | interval | second-order envelope | exact | per-step bound |
| GraphSAGE-2, GCN-3 | 2–3 layers | interval | covers the whole tail | exact | per layer |
| ResNet-50, last 4 blocks | , ReLU, | interval Clarke | first-order enclosure | exact | block Lipschitz |
| Symbol | meaning | value | enters | effect of moving it |
| , , | horizon, skip weight, step size of the incumbent | 32, 0.1, 0.9 | ( 9 ) | fixed by the incumbent |
| per-call one-sided budget (selected member) | G2, charges | oracle frontier concave; executed curve empirical (Fig. 2 c) | ||
| per-node floor | 1 nat | G1, row trim | gain and fallback (Fig. 4 d) | |
| regression tolerance (registered choice) | 0.002 nats | G4 | rate statements only | |
| , | statistical level, rate target | 0.05, 2% | Thm. 3 | bound widths |
| nominal fixed-member trajectory allocation | 0.1 | Thm. 1 (iii) | radii |
| Operation per training episode | count | forward passes | backward passes | memory |
|---|---|---|---|---|
| policy pass: incumbent steps and proposal at open depths | 1 | 1.18 | – | 6.9 GB |
| counterfactual tails for full-loss slopes, one per open depth | 16 | 3.75 | 3.75 | checkpointed every 4 steps |
| field, covariance and posterior heads | 1 | 0.04 | 0.05 | 0.3 GB |
| implicit differentiation through the solver (KKT systems) | 16 | – | 0.03 | small |
| total | 4.97 | 3.83 | 18.4 GB peak |
| (a) Outcome | calls | share | by predicate |
|---|---|---|---|
| first-pass release (R1) | 2,035 | 99.37% | – |
| failed the first pass: damage / row / descent | 13 | 0.63% | 2 / 4 / 7 |
| released after re-certification (R2) | 6 | 0.29% | 1 / 3 / 2 |
| whole-call fallback (F) | 7 | 0.34% | 1 / 1 / 5 |
| (a) Implementation | half-width | certified mean |
|---|---|---|
| registered bets tuned to | 0.295 | 6.20 |
| horizon-free sequence | 0.316 | 6.18 |
| approximate Kelly | 0.352 | 6.15 |
| hedged capital ( ) | 0.328 | 6.17 |
| normal interval (not finite-sample valid) | 0.204 | 6.30 |
| Gaussian first-order width | 0.304 | – |
| (a) Population | share | corrector | ||||
|---|---|---|---|---|---|---|
| clean | 2.6 | 2.1 | 1.8 | 1.3 | 50% | 1.5 |
| degree | 8.9 | 7.4 | 6.4 | 5.3 | 60% | 4.6 |
| cache | 17.2 | 14.6 | 12.9 | 10.8 | 63% | 8.6 |
| noise | 12.6 | 10.6 | 9.4 | 7.7 | 61% | 6.8 |
| dropout | 11.8 | 10.0 | 8.8 | 7.4 | 63% | 6.4 |
| mixture | 10.6 | 8.9 | 7.9 | 6.5 | 61% | 5.6 |
| Policy | [95% CI] | to RTT | LB | Lat. | FLOPs | Holm | |||
| RTT | |||||||||
| RTT (registered) | 6.5 [6.3, 6.7] | – | 12 | 0.95% | 0.51% | 6.2 | 1.42 | 1.36 | – |
| RTT one-pass corrector on the unused budget | 6.7 [6.5, 7.0] | 0.2 [ 0.4, 0.1] | 12 | 0.95% | 0.51% | 6.4 | 1.49 | 1.40 | – |
| RTT, tolerance point ( ) | 2.0 [2.0, 2.1] | +4.5 [+4.3, +4.7] | 0 | 0 (det.) | 0.22% | 1.9 | 1.42 | 1.36 | – |
| One-pass floor correctors (one incumbent pass; same floor, labels, capacity, proposal rows) | |||||||||
| One-pass corrector, 2-hop head rows | 5.6 [5.4, 5.8] | +0.9 [+0.8, +1.1] † | 14 | 1.07% | 0.56% | 5.3 | 1.25 | 1.22 | – |
| Mechanism | swap | gain | what changes | |
|---|---|---|---|---|
| Per-node budget | none nat | 6.69 6.50 | 0.95% 0.95% | per-node floor becomes deterministic; fallback 0.34% 0.34% |
| Radii | per depth sequential, trajectory-valid | 6.86 6.50 | 1.01% 0.95% | trajectory statement of Theorem 1 (iii); selection caveat applies |
| Placement | spend-based certified efficiency | 6.17 6.50 | 0.95% 0.95% | trust moves to depths with high |
| Window | envelope only two branches | 6.45 6.50 | 0.95% 0.95% | optimum over the union of both feasible sets |
| Features | mean pooling support functions | 6.39 6.50 | 1.01% 0.95% | executed step exactly invariant (Table 18 ) |
| Slopes | probe loss full loss | 6.16 6.50 | 1.13% 0.95% | accounting identity closes |
| (a) Neighbor | gain | Lat. | |
|---|---|---|---|
| WiSE-FT (incumbent v2), by learn-then-test | 3.4 | 2.15% × | 1.00 |
| LoRA scaling by learn-then-test | 3.2 | 2.20% × | 1.00 |
| EATA (anti-forgetting test-time adaptation) | 2.6 | 3.51% × | 2.10 |
| RDumb (periodic reset) | 2.2 | 2.86% × | 2.05 |
| -regularized fine-tuning (floor trained, not enforced) | 3.4 | 1.81% | 1.00 |
| Dataset / shift | incumbent | RTT | corr. | TR | share | verified | Lat. | metric | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ogbn-arxiv, registered mixture | tanh diffusion, (tube) | 2 | 10.6 | 6.5 | 5.6 | 5.8 | 61% | 0.95% | 0.51% | 99.37% | 1.42 | acc. 75.7 76.1 |
| ogbn-arxiv, registered mixture | APPNP (linear tail, exact) | 2 | 9.4 | 6.0 | 5.4 | 5.3 | 64% | 0.89% | 0.46% | 100.00% | 1.22 | acc. 75.5 75.9 |
| ogbn-arxiv, registered mixture | GraphSAGE, 2 layers (shallow) | 2 | 8.1 | 5.2 | 5.0 | 4.6 | 64% | 0.83% | 0.46% | 100.00% | 1.52 | acc. 72.1 72.4 |
| ogbn-arxiv, registered mixture | GCN, 3 layers (shallow) | 2 | 8.4 | 5.4 | 5.2 | 4.8 | 64% | 0.89% | 0.46% | 100.00% | 1.47 | acc. 71.8 72.2 |
| ogbn-arxiv, 2020 cohort (temporal) | tanh diffusion (tube) | 2 | 2.9 | 1.9 | 1.6 | 1.3 | 66% | 1.18% | – | 99.51% | 1.42 | acc. 74.1 74.2 |
| GOOD-Arxiv, degree (covariate) | GCNII (ReLU, interval Jacobians) | 2.3 | 9.7 | 5.8 | 5.1 | 4.9 | 60% | 1.13% | 0.61% | 99.22% | 1.43 | acc. 63.8 64.1 |
| Reference family | AP ref. | AP TR | AP RTT | gain ( ) | RTT TR (AP) |
|---|---|---|---|---|---|
| Contractive diffusion | 26.4 | 27.8 | 28.5 | 1.52 | +0.7 [+0.4, +1.0] |
| Centered GIN-style | 27.3 | 28.2 | 28.4 | 1.03 | +0.2 [0.0, +0.4] |
| Incumbent GIN-style (deployed) | 27.7 | 28.2 | 28.6 | 0.88 | +0.4 [+0.1, +0.7] |
| GIN-style virtual node | 28.4 | 28.8 | 29.2 | 0.79 | +0.4 [+0.2, +0.6] |
| Official test split (audited once) | |||||
| Incumbent GIN-style (deployed) | 27.7 | – | 28.5 | – | – |
| (a) Proposal | its cost (ms) | RTT | 2 rollouts, sequential | 2 devices |
|---|---|---|---|---|
| bank prior + signed attention (ours) | 7.5 | 1.42 | 2.18 | 1.18 |
| GPR-GNN (adapter) | 12.1 | 1.53 | 2.29 | 1.29 |
| Co-GNN (adapter) | 16.0 | 1.62 | 2.38 | 1.38 |
| AMP (adapter) | 14.6 | 1.59 | 2.35 | 1.35 |
| GTrans (test-time adaptation) | 91.0 | 3.42 | 3.19 | 2.19 |
| Matcha (test-time adaptation) | 58.0 | 2.63 | 2.39 | 1.39 |
| Test | result | verdict |
|---|---|---|
| planted violations, hidden margins – nats | 10,000 of 10,000 rejected | test passed |
| 50-digit re-evaluation of certificates | 1,000 certificates; float64 enclosure above the high-precision reference value, by at most | test passed |
| exact endpoint cost (offline second rollout) | 0 violations on 2,048 calls; certified / exact median 1.61 | test passed |
| adversarial directions within the trim | released 41%; released damage mean 29, max 49 ( ); over all calls 11.9 | bound met; audit fails |
| provisional factors under-estimated by 30% | 214 of 214 violating calls caught (fallback) | test passed |
| outward rounding switched off | 3 of 2,048 certificates below the exact cost by | diagnostic |
| Line | check | changes | guarantee | reference run | updates |
|---|---|---|---|---|---|
| Differential verification ( Paulsen et al., 2020a ; Paulsen et al., 2020b ; Teuber et al., 2025 ) | offline | a second network | bounds on the difference | yes | any fixed pair |
| Certified upgrades ( Zhang et al., 2026 ; Balachandran, 2026 ) | offline audit | replacement model | statistical non-regression | yes | replacement |
| Conformal policy control ( Prinster et al., 2026 ) ; LTT, CALM ( Angelopoulos et al., 2025 ; Schuster et al., 2022 ) | calibration | policy or exit | risk control | yes / partly | policies, exits |
| CPI, PPO ( Kakade and Langford, 2002 ; Schulman et al., 2017 ) | training | policy | CPI: model-based improvement; PPO: clipped surrogate | – | policies |
| Weight interpolation ( Wortsman et al., 2022 ) | offline | weights | none per call | – | weight deltas |
| Safe / graph TTA ( Niu et al., 2022 ; Press et al., 2023 ; Bao et al., 2025 ; Hsu et al., 2026 ) | online | states or weights | none per call | – | test-time updates |
| Symbol | meaning |
|---|---|
| incumbent’s, adapted, served, raw update’s prediction; Bayes posterior | |
| the incumbent’s skip weight, step size and horizon | |
| , | one-sided damage of row ; worst-case loss increase and decrease of a call |
| per-call and per-node budgets; remaining call budget; row allowance | |
| floor set | |
| label-free information; expected gain given it |