When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions
Organizations: Centro Nacional de Inteligencia Artificial (CENIA) · Barcelona Supercomputing Center (BSC) · Pontificia Universidad Cat´olica de Chile
Abstract
A valid curvature upper bound need not justify either a robustness certificate or an intervention on an intrinsic predictor property. We demonstrate this distinction for a last-layer relative-flatness proxy used in both settings. First, empirical-risk stationarity does not eliminate pointwise first-order loss terms: at a finite global empirical-risk minimum, the retained certificate expression underestimates a loss increase by over . We derive a globally valid, gauge-invariant feature-space repair. Second, common-row softmax shifts preserve predictions and the exact contraction while making the proxy unbounded. Even standard reference-class choices double it on average relative to the centered representation. For a single fixed-feature example with at least three classes, scalar retuning generically cannot align the induced probability updates. Row centering gives the orbit-minimized bound and restores value and full-model gradient invariance under this symmetry. Across 45 paired one-step tests on algorithmic and image models, amplified shifts separate raw-regularized predictors while quotient-regularized predictors remain aligned. Long-horizon CIFAR-10 experiments show substantial, reversible suppression of generalization, while evidence for selective delay after memorization is less consistent. Together, these results show that validity as a curvature upper bound does not by itself justify either inversion into a robustness certificate or differentiation into an intrinsic training intervention.
Figures & tables
| Actual raw (%) | Linear/raw | Bound/actual | |
|---|---|---|---|
| 100.00 | 1060.27 | 1.039 | |
| 100.00 | 106.03 | 1.393 | |
| 100.00 | 10.60 | 4.908 | |
| 100.00 | 3.53 | 12.581 | |
| 98.34 | 1.06 | 38.012 |