Same Loss, Different Gradients
Organizations: Nanjing Normal University · Nanjing University of Chinese Medicine
Abstract
Differentiable learning typically assumes that the scalar objective evaluated in the forward pass and the gradient supplied to the optimizer in the backward pass describe the same mathematical object. We show that this correspondence can fail when probabilistic objectives rely on finite special-function recurrences, custom backward rules, and numerical clipping. In high-dimensional von Mises-Fisher learning, real numerical implementations can produce identical forward scores and losses at the same learning state while supplying different gradients and following different optimization trajectories. We characterize the structure of this mismatch in finite-start Bessel recurrence and show that classwise radial mismatch can compose through probabilities into a locally nonconservative update field. Evaluating the accuracy of special-function values and derivatives separately is therefore insufficient to characterize the realized learning objective. Motivated by this observation, we introduce AR/FR, a fixed-depth analytic realization that constructs a potential and its derivative jointly, ensuring forward-backward coherence by construction. We establish a uniform cubic-order error bound relative to the exact Bessel ratio over the entire nonnegative concentration axis and propagate this guarantee to learning scores and objectives. As representation dimension increases, the original finite recurrence becomes sequentially deeper, whereas the worst-case AR/FR error guarantee tightens cubically, jointly providing coherence, certified fidelity, and fixed-depth computation. These results suggest that a differentiable numerical primitive is defined by both the values it realizes and the derivatives it actually supplies to the optimizer; together, they constitute the numerical realization of the learning algorithm.
Figures & tables
| Dataset | IF | Metric | Original recurrence | Consistent recurrence | AR/FR |
|---|---|---|---|---|---|
| CIFAR-10-LT | 10 | Acc. | |||
| AUROC | |||||
| CIFAR-10-LT | 50 | Acc. | |||
| AUROC | |||||
| CIFAR-10-LT | 100 | Acc. | |||
| AUROC |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| View | Max score | Loss | Max | Gradient relative | Gradient max |
|---|---|---|---|---|---|
| 2 | |||||
| 3 |
| Raw | Sign | Phase | Scaled | Limit | Rel. error | ||
|---|---|---|---|---|---|---|---|
| 63 | 128 | linear | |||||
| 64 | 130 | inverse | |||||
| 255 | 512 | linear | |||||
| 256 | 514 | inverse |
| Comparison | Reference | Maximum discrepancy |
|---|---|---|
| Finite-potential score | Exact score | |
| AR/FR same-state score | Exact score | |
| Clipped backward radial factor | Exact Bessel ratio | |
| Forward–backward coherence defect | Finite-forward derivative |
| Dimension | CUSF | AR/FR | AR/FR certificate |
|---|---|---|---|
| 64 | |||
| 128 | |||
| 256 | |||
| 512 | |||
| 1024 | |||
| 2048 |
| Dimension | Log-Miller | CUSF | AR/FR |
|---|---|---|---|
| 64 | |||
| 128 | |||
| 256 | |||
| 512 | |||
| 1024 | |||
| 2048 |
| Numerical realization | Runtime (ms) | Peak increment (MiB) |
|---|---|---|
| Original recurrence | 266.40 | 504.89 |
| Consistent recurrence | 412.80 | 1755.42 |
| Log-domain Miller | 120.28 | 504.89 |
| AR/FR | 4.90 | 508.12 |
| Structure-aware AR/FR | 0.206 | 32.00 |
| Dataset | IF | Metric | Original recurrence | Consistent recurrence | AR/FR |
|---|---|---|---|---|---|
| CIFAR-10-LT | 10 | Acc. | |||
| AUROC | |||||
| FPR95 | |||||
| CIFAR-10-LT | 50 | Acc. | |||
| AUROC | |||||
| FPR95 |