Differentiable learning typically assumes that the scalar objective evaluated in the forward pass and the gradient supplied to the optimizer in the backward pass describe the same mathematical object. We show that this correspondence can fail when probabilistic objectives rely on finite special-function recurrences, custom backward rules, and numerical clipping. In high-dimensional von Mises-Fisher learning, real numerical implementations can produce identical forward scores and losses at the same learning state while supplying different gradients and following different optimization trajectories. We characterize the structure of this mismatch in finite-start Bessel recurrence and show that classwise radial mismatch can compose through probabilities into a locally nonconservative update field. Evaluating the accuracy of special-function values and derivatives separately is therefore insufficient to characterize the realized learning objective. Motivated by this observation, we introduce AR/FR, a fixed-depth analytic realization that constructs a potential and its derivative jointly, ensuring forward-backward coherence by construction. We establish a uniform cubic-order error bound relative to the exact Bessel ratio over the entire nonnegative concentration axis and propagate this guarantee to learning scores and objectives. As representation dimension increases, the original finite recurrence becomes sequentially deeper, whereas the worst-case AR/FR error guarantee tightens cubically, jointly providing coherence, certified fidelity, and fixed-depth computation. These results suggest that a differentiable numerical primitive is defined by both the values it realizes and the derivatives it actually supplies to the optimizer; together, they constitute the numerical realization of the learning algorithm.
Figures & tables
Figure 1: Same loss, different gradients: diagnosis and paired realization. (a) A common forward objective can supply different updates when backward rules differ. (b) The coherence defect compares the supplied backward with the derivative of that forward. (c) AR/FR constructs the potential and derivative together, combining coherence with a uniform fidelity bound.
Figure 2: Local geometry of coherent and mismatched update fields. Both panels share the same scalar forward surface. A matched backward yields a conservative gradient field with symmetric Jacobian, whereas a mismatched backward can introduce a local rotational component and asymmetric cross-derivatives. The figure is an analytic illustration.
Figure 3: Same-state gradient mismatch becomes trajectory separation. Original and consistent recurrence share random or intermediate-training starting states and augmented inputs. Left: local gradient gaps with bitwise-identical forward values. Right: parameter and feature distances after separate updates.
Figure 4: Concentration and depth jointly determine coherence. Colour shows log10(∣gsupplied−Φ′∣/∣Φ′∣) , with a 10−16 display floor. The diamond locates captured PATT queries at p=512 and M/ν=2 .
Figure 5: Cubic error decay. Blue: sampled maximum error. Orange: the all-axis certificate ν−3 .
Dataset
IF
Metric
Original recurrence
Consistent recurrence
AR/FR
CIFAR-10-LT
10
Acc. ↑
92.04±0.17
92.12±0.18
92.16±0.19
AUROC ↑
93.07±0.13
92.30±0.27
93.43±0.46
CIFAR-10-LT
50
Acc. ↑
87.30±0.29
87.58±0.35
87.42±0.15
AUROC ↑
90.18±2.46
90.95±1.57
91.02±1.13
CIFAR-10-LT
100
Acc. ↑
83.82±0.21
84.67±0.41
84.73±0.31
AUROC ↑
90.66±0.13
89.97±0.39
89.95±0.34
Table 1: Classification and OOD detection under different numerical realizations. CIFAR results report mean ± standard deviation over three runs. Metrics are percentages; IF denotes the imbalance factor. Full FPR95 results appear in Appendix G.1 . Bold and underline indicate the best and second-best values within each setting.
Figure 6: Local runtime–memory trade-off.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: ProCo coherence defects depend on the visited state. Absolute differences compare the supplied radial rule with FP64 autodiff at M=2ν . Boxes show median and interquartile range; whiskers span extrema. (a) CIFAR queries, (b) ImageNet queries, and (c) iNaturalist class-bank concentrations. The dashed diagnostic scale is 10−6 ; values below 10−16 share a display floor.
View
Max score ∣Δ∣
Loss ∣Δ∣
Max ∣Dclip∣
Gradient relative ℓ2
Gradient max ∣Δ∣
2
0
0
2.55×10−3
2.56×10−3
1.40×10−5
3
0
0
2.55×10−3
2.56×10−3
1.35×10−5
Appendix
Table 2: Changing only the backward rule reveals a PATT coherence defect. Original and consistent recurrence use the same frozen CIFAR100-LT state and two augmented views (seed 3407). Score and loss columns report absolute differences; gradient columns compare the complete adjusted cross-entropy feature gradients. Dclip=Φ′−g=−d is the signed radial residual.
ν
p
Raw R
Sign
Phase
Scaled
Limit
Rel. error
63
128
1.645×105
+
linear
1.645×10−4
1.645×10−4
1.2×10−11
64
130
6.208×10−6
−
inverse
6.208×103
6.208×103
1.3×10−11
255
512
1.020×104
+
linear
1.020×10−5
1.020×10−5
3.2×10−9
256
514
9.856×10−5
−
inverse
9.856×104
9.856×104
3.3×10−9
Appendix
Table 3: Finite-recurrence asymptotics follow the predicted parity. At x=109 and M=2ν , odd and even orders are compared with their analytic limits using 100-digit arithmetic. The scaled quantity is R/x for odd ν and xR for even ν ; sign refers to R−Rν .
Comparison
Reference
Maximum discrepancy
Finite-potential score
Exact score
3.32×10−3
AR/FR same-state score
Exact score
1.17×10−9
Clipped backward radial factor
Exact Bessel ratio
2.55×10−3
Forward–backward coherence defect
Finite-forward derivative
2.56×10−3
Appendix
Table 4: Score fidelity and backward coherence at captured PATT states. Maximum absolute discrepancies over six CIFAR100-LT views. Score and radial fidelity use the exact Bessel reference; coherence compares the supplied backward with the derivative of its own finite forward.
Dimension
CUSF
AR/FR
AR/FR certificate
64
3.37×10−12
2.74×10−6
3.36×10−5
128
2.38×10−11
3.29×10−7
4.00×10−6
256
2.29×10−11
4.03×10−8
4.88×10−7
512
1.85×10−11
4.99×10−9
6.03×10−8
1024
2.45×10−11
6.21×10−10
7.49×10−9
2048
8.87×10−11
7.74×10−11
9.34×10−10
Appendix
Table 5: Ratio fidelity across feature dimensions. Maximum absolute errors use 61 log-spaced concentrations x/ν∈[10−3,103] per dimension and a 60-digit Bessel reference. CUSF forms a ratio from adjacent log-Bessel values in CPU FP64; AR/FR is evaluated in 60-digit arithmetic. The certificate ν−3 covers the full nonnegative axis.
Dimension
Log-Miller
CUSF
AR/FR
64
3.36×10−4
1.89×10−12
2.76×10−6
128
8.29×10−5
8.47×10−12
3.30×10−7
256
2.06×10−5
2.31×10−11
4.04×10−8
512
5.13×10−6
9.72×10−12
5.00×10−9
1024
1.28×10−6
3.21×10−11
6.21×10−10
2048
2.11×10−7
9.67×10−11
7.74×10−11
Appendix
Table 6: Fidelity of unit-length potential differences. Maximum absolute errors in Φν(x+1)−Φν(x) use the common 61-point concentration grid at each dimension. Log-Miller retains the finite start; CUSF subtracts separately evaluated potentials; AR/FR uses its paired analytic potential. All errors use the high-precision Bessel reference.
Figure 8: Score and loss errors remain below their certificates as class count varies. The 15 fixed-state configurations combine five class counts with prior ratios 1, 100, and 1000. (a) The three maximum score-error curves coincide and are drawn once. (b) Mean cross-entropy error is shown separately for each prior ratio. Dashed lines denote the respective theoretical bounds. Class counts are equally spaced discrete settings.
Figure 9: A frozen PATT state exhibits mainly gradient-magnitude distortion. The first CIFAR-10 query is varied over a 25×25 slice along two principal feature directions, with class statistics fixed. Scores and losses agree bitwise at every location. (a–b) Negative gradients from the finite forward and supplied backward, using common contours and arrow scales. (c) Their difference magnified 400× ; each field displays 64 interior arrows. (d–e) Empirical cumulative distributions over all 625 projected-gradient pairs. Relative mismatch has median 0.2561% ; 1−cos resolves the small directional difference.
Figure 10: Classwise mismatch controls the local Jacobian obstruction. (a) Signed antisymmetry versus concentration ratio at 29 constructed states: lines show the analytic prediction and circles the differentiated field; the consistent field remains zero. (b) Gradient-gap norm and absolute antisymmetry across nine mismatch strengths, with predictions shown as lines and computed values as markers. All interventions retain forward loss log2 .
Figure 11: Concentration and gradient scaling delimit the coherence defect. (a) Absolute clipped defect over 130 conditions; dashed segments show the asymptote (ν+3/2)/x , and the grey interval marks saved PATT query radii. Values below 10−16 share a display floor. (b) Cumulative distributions for 256 frozen PATT queries in the full feature space with FP64 geometry. Relative differences use the consistent gradient norm before scaling and the supplied gradient norm after optimal per-query scaling.
Numerical realization
Runtime (ms)
Peak increment (MiB)
Original recurrence
266.40
504.89
Consistent recurrence
412.80
1755.42
Log-domain Miller
120.28
504.89
AR/FR
4.90
508.12
Structure-aware AR/FR
0.206
32.00
Appendix
Table 7: Local H100 forward–backward costs. Median runtime and incremental peak allocation under the protocol in Appendix F.3 .
Dataset
IF
Metric
Original recurrence
Consistent recurrence
AR/FR
CIFAR-10-LT
10
Acc. ↑
92.04±0.17
92.12±0.18
92.16±0.19
AUROC ↑
93.07±0.13
92.30±0.27
93.43±0.46
FPR95 ↓
26.20±1.19
29.52±1.06
25.52±1.01
CIFAR-10-LT
50
Acc. ↑
87.30±0.29
87.58±0.35
87.42±0.15
AUROC ↑
90.18±2.46
90.95±1.57
91.02±1.13
FPR95 ↓
34.51±6.96
32.84±1.94
32.23±2.89
Appendix
Table 8: Complete classification and OOD detection results. Metrics are percentages; CIFAR results report mean ± standard deviation over three runs. Bold and underline indicate the best and second-best values within each setting.
How accurate must a numerical approximation be within a learning system? Primitive error alone cannot answer this question: errors of the same magnitude can have very different consequences for losses, predictions, and gradients at different learning states. We study this question through the learning objective itself. The objective weights classwise numerical errors nonuniformly according to the current state, so the importance of an error depends not only on its magnitude but also on the class it affects and the weight that class receives. For softmax cross-entropy, we characterize this coupling between class weights and errors and derive the exact extrema of the signed loss change over pairings of fixed non-target probability and score-error multisets, with the target probability and target score error held fixed. Building on this structure, we establish finite-error guarantees that propagate primitive error to losses, probabilities, predictions, and feature gradients, then invert these guarantees to obtain a certified primitive tolerance for the current state under prescribed learning-level error requirements. We give a complete instantiation of the framework in high-dimensional von Mises-Fisher learning. Controlled interventions and a large collection of saved learning states show that identical primitive error can produce substantially different learning consequences, while certified numerical tolerances vary by orders of magnitude across states under the same learning-level requirements. These results show that the adequacy of a numerical approximation must be assessed in relation to the current learning state and the quantity to be preserved; numerical accuracy should itself be treated as part of the learning objective.
Ningkang Peng, Qianfeng Yu, Jingyang Mao +2
Nanjing Normal University · Nanjing University of Chinese Medicine
Learn-then-differentiate (LTD) estimates gradients by fitting a model to simulation outputs and differentiating it. We develop a unified framework explaining what LTD differentiates and how accurately it estimates gradients. For models with a weighted representation, LTD differentiates a learned representation of the underlying probability measure. We then show how accuracy guarantees for fitted models translate into guarantees for gradients and higher-order derivatives, with rates approaching the standard Monte Carlo rate under suitable smoothness conditions. The framework recovers established results for kernel regression, local polynomial regression, and kernel ridge regression, and yields further guarantees for multiple kernel learning and smooth neural networks. These results provide a common foundation for understanding and analyzing LTD across learning methods.
Nifei Lin, Qingkai Zhang, L. Jeff Hong
Research Institute for Interdisciplinary Sciences, School of Information Management and Engineering, Shanghai University of Finance and Economics, Shanghai 200433, China · Department of Decision Analytics and Operations, City University of Hong Kong, Hong Kong, China · Department of Industrial and Systems Engineering, University of Minnesota, Minneapolis, Minnesota 55455
Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory. We show that small forward-marginal error does not guarantee numerical stability. We construct a single smooth score field with arbitrarily small forward-marginal L2 error. The learned reverse-time process is nonexplosive, has moments of every order, and can be arbitrarily close to the exact reverse-time process in path-space total variation. Yet its Euler--Maruyama discretizations converge in probability while every positive moment diverges. Thus weak convergence can hold even though every Wasserstein distance Wp, p≥1, diverges. The same failure can occur within one fixed finite neural architecture. We construct a family of bounded, globally Lipschitz denoisers for which both the forward-marginal error and the path-space total variation distance tend to zero, while their Euler--Maruyama endpoints diverge in every Wp. For compactly supported data, we also give a simple positive result. Projecting the learned denoiser onto a known bounded closed convex set containing the support preserves pointwise accuracy, gives grid-uniform moment bounds, and yields Wasserstein convergence under mild local regularity. Experiments with a small fixed DiT-style network show large growth along rare numerical trajectories and its suppression by denoiser projection, while overall trajectory errors remain small.