The Signed Geometry of One-Shot Recourse: On-Path Validity and the Signed-Curvature Criterion
Organizations: Google
Abstract
Closed-form recourse moves a rejected user along the unit gradient of the classifier score by the promised distance , at which the linearized score reaches zero. We ask when this one-shot step succeeds and what additional model queries change. To leading order the step ends on the favorable side exactly when the path curvature is nonnegative. Across 80 shallow models, the fraction of rejected users whose step ends there and the fraction with correlate at , although on Fashion-MNIST the first falls below the second by 8.2 points on average. No rule that uses only the score value and gradient can be valid for every score with path curvature bounded by without overshooting some by order . When the curvature is also Lipschitz and the step is short, one evaluation of at the promised point attains the minimax rate among deterministic one-query rules that know the curvature bound and its Lipschitz constant, and split-conformal calibration makes such a rule reach the first crossing or abstain with probability at least . Training with an asymmetric curvature penalty lets 99-100% of paths cross within the promised step on undershoot-prone shallow data, at about 4-22 times the overshoot of symmetric penalties (Fashion-MNIST, COMPAS). Because and depend on how the score is scaled, part of this gain can be a longer promised step, and at matched validity a smaller audit of briefly trained models finds no uniform advantage over tuned inflation. Where a per-user line search along the ray is affordable, it is exact to grid resolution and preferable.
Figures & tables
| Dataset | Method | Bal. Acc. (%) | (nearest) | (on-path) | Validity (%) |
|---|---|---|---|---|---|
| COMPAS | Unregularized | ||||
| COMPAS | Spectral Norm | ||||
| COMPAS | 1-Lipschitz GP | ||||
| COMPAS | Global Hutchinson | ||||
| COMPAS | MW-Hutchinson | ||||
| German | Unregularized |
| Queries | Rule | Val. % | Oversh. | |
|---|---|---|---|---|
| none | alpha-1 | 80 | 76.6 | 0.041 |
| none | tuned inflation | 80 | 98.2 | 0.061 |
| 1 HVP | signed-quadratic | 80 | 28.9 | 0.005 |
| 1 HVP | conformal-quadratic | 80 | 95.5 | 0.025 |
| 1 forward | conformal-probe | 80 | 96.0 | 0.0016 |
| 189 forward | ray line search | 80 | 100.0 | 0.000 |
| Data | med. | screen pass | p90 stress | slab cont. | crit. err. pass/reject |
|---|---|---|---|---|---|
| COMPAS | 0.018 | 93.8% | 75% | 99.1% | 0.021 / 0.000 |
| German | 0.024 | 100% | 97% | 97.9% | 0.000 / – |
| Adult | 0.61 | 68.2% | 60% | 99.3% | 0.000 / 0.035 |
| COMPAS (MW) | 0.008 | 98.5% | 89% | 99.3% | 0.041 / 0.000 |
| CelebA depth (unreg) | 138 | 7.8% | 3.9% | 99.6% | 0.123 / 0.423 |
| CelebA depth (Asym) | 42 | 10.5% | 7.7% | 99.4% | 0.171 / 0.366 |
| Suite (driver) | Data, epochs, seeds | Weights ; curvature target |
|---|---|---|
| Main comparison, main-paper Table 1 ( exp.py ) | COMPAS, German, Adult (20k); 50 epochs; seeds 0–9 | COMPAS/German/Adult: gradient penalty 0.5/0.05/0.5; MW-Hutchinson and Global Hutchinson 0.2/0.05/0.05; Spectral Norm none |
| Five-seed suite: signed-curvature criterion and common cohort, signed-quadratic and conformal rules, endpoint and clipping audits, score gauge, sign isolation, deployment shift, mask audit ( run_sig.py ) | COMPAS and German 50 epochs, Adult (8k) 30, F-MNIST 15; seeds 0–4 | MW / Global / Asym : COMPAS 0.2 / 0.2 / (2, 0.1); German 0.05 / 0.05 / (1, 0.05); Adult 2.0 / 1.0 / (10, 0.05); F-MNIST 0.1 / 0.1 / (10, 0.05). Sign isolation: twin at the Asym ; sign-blind at , , and times the Asym |
| Adult penalty sweep ( final_adult.py ) | Adult (8k); 30 epochs; seeds 0–4 | MW and Global ; Asym , |
| FaiR-N comparison ( fairn2.py , fairn_adult.py ) | COMPAS 50 epochs, seed 0 (sweep) and seeds 0–4; Adult (8k) 30 epochs, seeds 0–1; plain accuracy | FaiR-N adds times the squared difference of the two groups’ mean in each minibatch: (COMPAS sweep), (COMPAS, five seeds), (Adult); Asym on COMPAS, on Adult |
| COMPAS penalty suite ( compas_asym.py ) | COMPAS; 50 epochs; seeds 0–2; ray metrics on the first 200 rejected test points in split order; plain (not balanced) accuracy | MW 0.2; Asym , |
| F-MNIST, Table S3 ( fm_final.py ) | 15 epochs; seeds 0–4; unweighted loss | MW 0.1; Global 0.1; Asym (10, 0.05) |
| Dataset | Best sym. validity | Asym validity | Asym acc. | Asym overshoot |
|---|---|---|---|---|
| Adult (sweep, , ) | 70.9% (MW) | 99.4–100.0% | 81.4–81.5% | 0.038–0.092 |
| F-MNIST | 74.7% (MW) | 99.9% | 88.5% | 0.22 |
| CIFAR-10 ( , ten seeds) | 95.4% (Global) | 91.6% | 92.0% | 1.003 |
| Method | Acc. | Under | Over | Validity |
|---|---|---|---|---|
| Unreg | 88.40 | 0.0082 | 0.0516 | |
| MW-Hutch | 88.92 | 0.0065 | 0.0653 | |
| Global | 89.05 | 0.0108 | 0.0546 | |
| Asym | 88.53 | 0.0017 | 0.2203 |
| Dataset | Method | Probe val. % | Probe overshoot | HVP val. % | HVP overshoot | Overshoot ratio |
|---|---|---|---|---|---|---|
| COMPAS | asym | 94.5 | 0.0020 | 94.8 | 0.0120 | 6 |
| COMPAS | glob | 95.3 | 0.0000 | 93.3 | 0.0058 | 133 |
| COMPAS | mw | 94.1 | 0.0001 | 95.3 | 0.0077 | 70 |
| COMPAS | unreg | 95.7 | 0.0023 | 95.2 | 0.0203 | 9 |
| German | asym | 95.5 | 0.0003 | 94.8 | 0.0027 | 10 |
| German | glob | 99.0 | 0.0000 | 97.6 | 0.0035 | 127 |
| Inference-query budget | Rule | Validity % | Overshoot |
|---|---|---|---|
| none (training-time only) | asymmetric | 99.6 | 0.103 |
| validation only (closed form) | tuned inflation | 98.2 | 0.061 |
| one forward eval, no HVP | conformal-probe | 96.0 | 0.0016 |
| one HVP | conformal-quadratic | 95.5 | 0.025 |
| per-user ray search | ray line search (oracle) | 100.0 | 0.000 |
| Setting | models | |
|---|---|---|
| Shallow common cohort (COMPAS/German/Adult/F-MNIST) | 0.985 | 80 |
| Adult only ( suite) | 0.996 | 20 |
| one-shot | signed-quad | conformal-quad | tuned inflation | |||||
|---|---|---|---|---|---|---|---|---|
| Data | ||||||||
| COMPAS | 0.867 | 0.035 | 0.447 | 0.004 | 0.946 | 0.011 | 0.970 | 0.071 |
| Adult ( ) | 0.456 | 0.011 | 0.322 | 0.001 | 0.954 | 0.027 | 0.983 | 0.023 |
| F-MNIST | 0.739 | 0.098 | 0.157 | 0.015 | 0.959 | 0.060 | 0.975 | 0.130 |
| German | 1.000 | 0.020 | 0.229 | 0.000 | 0.960 | 0.003 | 1.000 | 0.020 |
| Setting | median reduction |
|---|---|
| COMPAS | 78% |
| F-MNIST | 40% |
| Adult (full UCI) | 32% |
| pooled (above three, ) | 54% |
| Rule | HVP | queries | ||
|---|---|---|---|---|
| one-shot ( ) | 0 | 0 | 0.766 | 0.041 |
| tuned inflation | 0 | 0 | 0.982 | 0.061 |
| signed-quadratic | 1 | 0 | 0.289 | 0.005 |
| conformal-quadratic | 1 | 0 | 0.955 | 0.025 |
| conformal-quadratic (Mondrian) | 1 | 0 | 0.951 | 0.020 |
| ray line search | 0 | many | 1.000 | 0.000 |
| Dataset / shift | Rule | Src % | Tgt % | Drop% | Tgt | |
|---|---|---|---|---|---|---|
| COMPAS (cov., priors) | VT inflation | |||||
| conformal stale | ||||||
| conformal recal. | ||||||
| asym. training | ||||||
| COMPAS (subpop.) | VT inflation | |||||
| conformal stale |
| Method | Acc. | |||||
|---|---|---|---|---|---|---|
| Unreg | ||||||
| MW | ||||||
| Global | ||||||
| Asym |
| Method | ||||||
|---|---|---|---|---|---|---|
| Unreg | ||||||
| MW | ||||||
| Global | ||||||
| Asym |
| Method | Acc. | ||||
|---|---|---|---|---|---|
| Unreg | 94.5 | 81.0 | 85.3 | 3.487 | 1.102 |
| MW | 93.3 | 92.5 | 98.1 | 1.635 | 3.129 |
| Global | 93.1 | 95.4 | 98.6 | 1.299 | 4.644 |
| Asym | 92.0 | 91.6 | 95.3 | 0.549 | 1.003 |
| Data | Meth. | S100 | |||||||
|---|---|---|---|---|---|---|---|---|---|
| compas | asym | 1.000 | 0.139 | 1.000 | 0.139 | 1.000 | 0.139 | 3 | |
| compas | glob | 1.033 | 0.040 | 1.033 | 0.040 | 1.033 | 0.040 | 3 | |
| compas | mw | 1.033 | 0.039 | 1.033 | 0.039 | 1.033 | 0.039 | 3 | |
| compas | unreg | 1.033 | 0.041 | 1.033 | 0.041 | 1.033 | 0.041 | 3 | |
| german | asym | 1.017 | 0.031 | 1.017 | 0.031 | 1.017 | 0.031 | 3 | |
| german | glob | 1.033 | 0.038 | 1.050 | 0.056 | 1.050 | 0.056 | 3 |
| Data | Target | Meth. | Sel. | Test | Gap | Under | Over | |
|---|---|---|---|---|---|---|---|---|
| compas | 0.90 | asym | 1.000 | -0.013 | 0.000 | 0.038 | 128–128 | |
| compas | 0.95 | asym | 1.003 | -0.005 | 0.000 | 0.042 | 128–128 | |
| compas | 0.99 | asym | 1.007 | 0.000 | 0.000 | 0.046 | 128–128 | |
| compas | 0.90 | glob | 1.017 | 0.000 | 0.001 | 0.021 | 128–128 | |
| compas | 0.95 | glob | 1.020 | 0.005 | 0.000 | 0.025 | 128–128 | |
| compas | 0.99 | glob | 1.032 | 0.003 | 0.000 | 0.039 | 128–128 |
| Data | Target | Group/method/rule | Valid. | Over | Bin. cost | Eff. cost |
|---|---|---|---|---|---|---|
| compas | 0.95 | Asym asym VT | 0.984 | 0.038 | 0.116 | 0.039 |
| compas | 0.95 | non-Asym glob VT | 0.979 | 0.025 | 0.129 | 0.027 |
| compas | 0.99 | Asym asym VT | 0.997 | 0.042 | 0.055 | 0.042 |
| compas | 0.99 | non-Asym unreg VT | 0.997 | 0.036 | 0.049 | 0.036 |
| compas | 100% | Asym asym VT | 1.000 | 0.046 | 0.046 | 0.046 |
| compas | 100% | non-Asym mw CI | 1.000 | 0.059 | 0.059 | 0.059 |
| Data | Method | Mask | Dims | Valid. | Under | Over | |
|---|---|---|---|---|---|---|---|
| compas | unreg | protected_only | 6 | 97.3 | 0.013 | 0.113 | 470 |
| compas | unreg | strict | 0 | – | – | – | 470 |
| compas | mw | protected_only | 6 | 99.5 | 0.000 | 0.052 | 473 |
| compas | mw | strict | 0 | – | – | – | 473 |
| compas | glob | protected_only | 6 | 99.5 | 0.000 | 0.040 | 476 |
| compas | glob | strict | 0 | – | – | – | 476 |
| Data | Cost | Best | Cost | |
|---|---|---|---|---|
| compas | binary | asym VT | 1.007 | 0.046 |
| compas | effort | unreg VT | 1.017 | 0.024 |
| fmnist | binary | unreg VT | 1.078 | 0.294 |
| fmnist | effort | unreg VT | 1.050 | 0.153 |
| german | binary | asym VT | 1.020 | 0.023 |
| german | effort | asym VT | 1.007 | 0.011 |
| Attempt | Intended mechanism | Observed outcome |
|---|---|---|
| Calibrated curvature target | Replace the fixed target of the asymmetric penalty by a per-point target , meant to clear the alpha-1 boundary with less excess curvature. | On COMPAS, Adult, and F-MNIST every calibrated point of the two-seed frontier run has lower validity and higher overshoot than some fixed- point, and all six three-seed configurations cost more than the fixed- penalty in both travel-error costs at . |
| Steepness penalty | Add times the mean of rejected minibatch points to the asymmetric penalty ( ) to change the fixed-scale tradeoff. | For every , accuracy falls from / / at to / / (COMPAS/Adult/F-MNIST), and at least one of three seeds rejects no test point. |
| Adult symmetric penalty sweep | Give MW/Global Hutchinson penalties larger regularization budgets on an undershoot-dominated dataset. | Validity remains far below the asymmetric penalty even when the diagnostic gap shrinks. |
| Result (section) | Code | Result, protocol and other files |
|---|---|---|
| Score-gauge audit (S1) | score_gauge_audit.py | results/score_gauge/ |
| Probe bounds and collocation cancellation (S1) | verify_probe_escape.py : exact recovery on quadratic profiles, the bound, the equivalence on trained models, the symbolic cubic case, and non-polynomial and kinked profiles | |
| Segment-safe fallback (S1) | endpoint_segment_fallback.py recomputes the upper-tail fallback and checks the branch invariant | endpoint-audit files (S2 row) |
| Drift screen and ambiguity mass (S1) | verify_drift_screen.py ; float64 recomputation and grid scan: verify_drift_float64.py | criteria: DRIFT_SCREEN_FALSIFIER.md |
| Theorem 6.3 and Corollary Pointwise sharpness of the probe. (S1) | verify_one_query_lower_bound.py ; verify_oracle_lattice.py ; their result files check exact transcripts, regularity, first roots, finite margins and the separate parameter scalings | criteria: ONE_QUERY_LOWER_BOUND_FALSIFIER.md ; ORACLE_LATTICE_FALSIFIER.md |
| Adaptive radii (conjecture) and Theorem Many Queries: an Adaptivity Separation (S1) | verify_adaptive_q_query_bounds.py , which also asserts that no matching adaptive converse is claimed; its summaries are rechecked by claims_manifest.py | results/theory_audit/adaptive_q_query_bounds.csv ; criteria: ADAPTIVE_QUERY_LOWER_BOUND_FALSIFIER.md ; written argument for the conjecture, not part of this supplement: ADAPTIVE_Q_QUERY_FINDING.md |
| Checklist area | Where documented |
|---|---|
| Problem and assumptions | Main-paper definitions, theorem statements, and limitations. |
| Algorithms and training | Main setup, S3.2, and the scripts of the code archive. |
| Data and preprocessing | Main setup, S3.2, the data notes and instructions of the code archive. |
| Evaluation metrics | Main definitions and S2. |
| Compute and runtime | Main-paper checklist and the runtime and progress output of the scripts. |
| Randomness and seeds | Main setup, S3.2, the seed columns of the result files, and the commands of the code archive. |