How accurate must a numerical approximation be within a learning system? Primitive error alone cannot answer this question: errors of the same magnitude can have very different consequences for losses, predictions, and gradients at different learning states. We study this question through the learning objective itself. The objective weights classwise numerical errors nonuniformly according to the current state, so the importance of an error depends not only on its magnitude but also on the class it affects and the weight that class receives. For softmax cross-entropy, we characterize this coupling between class weights and errors and derive the exact extrema of the signed loss change over pairings of fixed non-target probability and score-error multisets, with the target probability and target score error held fixed. Building on this structure, we establish finite-error guarantees that propagate primitive error to losses, probabilities, predictions, and feature gradients, then invert these guarantees to obtain a certified primitive tolerance for the current state under prescribed learning-level error requirements. We give a complete instantiation of the framework in high-dimensional von Mises-Fisher learning. Controlled interventions and a large collection of saved learning states show that identical primitive error can produce substantially different learning consequences, while certified numerical tolerances vary by orders of magnitude across states under the same learning-level requirements. These results show that the adequacy of a numerical approximation must be assessed in relation to the current learning state and the quantity to be preserved; numerical accuracy should itself be treated as part of the learning objective.
Figures & tables
Figure 1: From primitive error bounds to state-conditioned numerical tolerance. A primitive error bound propagates through the learning objective. The same classwise error structure can have different downstream effects in different learning states. Inverting certified loss and gradient bounds under budgets ϵL and ϵg yields the certified primitive tolerance b⋆(s;ϵ) . The upper panels are schematic. The lower panel shows saved-state tolerances under the joint budget ϵL=ϵg=10−5 .
Figure 2: Learning consequences change while primitive error stays fixed. a,b, At fixed f/τ , changing temperature rescales feature-gradient error without changing the scalar numerical problem. c, A shared target-class bias changes objective weights at fixed geometry and temperature. Blue shows actual error; orange shows the geometry-only certificate in a,b and the objective-conditioned certificate in c. The gray curve in c is the unchanged geometry-only bound; dotted lines mark a common gradient budget.
Figure 3: Query-level tolerance varies under a common learning-error budget. a, Ordered joint tolerances at budget 10−5 . b, The same queries grouped by dimension and temperature; black bars denote medians. Colors indicate dimension: blue 128, orange 512, and purple 1024.
Budget
Quantity
5%
Median
95%
Full span
10−4
Loss
4.86×10−7
9.58×10−6
2.40×10−3
5.74
10−4
Joint
7.35×10−8
1.06×10−6
1.44×10−4
6.24
10−5
Loss
4.86×10−8
9.58×10−7
2.86×10−4
6.65
10−5
Joint
7.35×10−9
1.06×10−7
1.44×10−5
7.16
Table 1: Distribution of certified primitive tolerance under loss and joint budgets. Columns report the 5th percentile, median, 95th percentile, and full logarithmic span log10(bmax/bmin) .
Figure 4: Local sensitivity and classwise correspondence shape numerical requirements. a, Ratio of first-order to nonlinear tolerance at budget 10−5 ; the inset magnifies the region near one. Colors denote dimension as in Fig. 3 . b, Cumulative distributions of minimum and maximum absolute errors over five fixed assignments of non-target probabilities. Target probability, entropy, exact top-two margin, geometry, and primitive errors are held fixed. c, Fractions of queries for which all assignments satisfy the joint budget, assignments yield different decisions, or all assignments fail.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Summary
Loss
Feature gradient
Fold change: median
3.18
3.28
Fold change: maximum
1.06×107
1.16×107
Median absolute range
2.36×10−9
9.09×10−9
Maximum absolute range
1.78×10−7
8.55×10−7
Appendix
Table 2: Class assignments change learning errors at fixed probability summaries. Fold change is the largest error divided by the smallest error for each state; absolute range is their difference.
Budget
Loss
Gradient
Joint
10−4
0
0
0
10−5
0
0
0
10−6
0
0
0
10−7
56
1014
1001
10−8
1150
526
521
10−9
571
148
138
Appendix
Table 3: Actual numerical adequacy under class-weight rearrangement. Counts give the states for which the five assignments yield different adequacy decisions. The joint requirement applies the same budget to absolute loss error and Euclidean feature-gradient error.
Differentiable learning typically assumes that the scalar objective evaluated in the forward pass and the gradient supplied to the optimizer in the backward pass describe the same mathematical object. We show that this correspondence can fail when probabilistic objectives rely on finite special-function recurrences, custom backward rules, and numerical clipping. In high-dimensional von Mises-Fisher learning, real numerical implementations can produce identical forward scores and losses at the same learning state while supplying different gradients and following different optimization trajectories. We characterize the structure of this mismatch in finite-start Bessel recurrence and show that classwise radial mismatch can compose through probabilities into a locally nonconservative update field. Evaluating the accuracy of special-function values and derivatives separately is therefore insufficient to characterize the realized learning objective. Motivated by this observation, we introduce AR/FR, a fixed-depth analytic realization that constructs a potential and its derivative jointly, ensuring forward-backward coherence by construction. We establish a uniform cubic-order error bound relative to the exact Bessel ratio over the entire nonnegative concentration axis and propagate this guarantee to learning scores and objectives. As representation dimension increases, the original finite recurrence becomes sequentially deeper, whereas the worst-case AR/FR error guarantee tightens cubically, jointly providing coherence, certified fidelity, and fixed-depth computation. These results suggest that a differentiable numerical primitive is defined by both the values it realizes and the derivatives it actually supplies to the optimizer; together, they constitute the numerical realization of the learning algorithm.
Ningkang Peng, Xiaoqian Peng, Yifan He +4
Nanjing Normal University · Nanjing University of Chinese Medicine
We investigate limitations of learning tanh neural networks from point evaluations under finite-precision computations and Lp accuracy guarantees, building on Berner, Grohs, and Voigtländer (2023). Our approach is based on a novel construction of sharply localized bump functions via iterated tanh activations. Using this mechanism, we show that, in a finite-precision setting, no adaptive randomized algorithm based on m samples can achieve a convergence rate higher than the Monte Carlo rate O(m−1/p) in the Lp norm, unless the sampling budget grows exponentially with the size of the network parameters and architecture. The results reveal fundamental limitations imposed by finite precision on the learnability of classes containing localized bump functions, extending previous results for ReLU networks to the tanh setting.
Philipp Grohs, Matěj Trödler
Faculty of Mathematics, University of Vienna · RICAM, Austrian Academy of Sciences
Learn-then-differentiate (LTD) estimates gradients by fitting a model to simulation outputs and differentiating it. We develop a unified framework explaining what LTD differentiates and how accurately it estimates gradients. For models with a weighted representation, LTD differentiates a learned representation of the underlying probability measure. We then show how accuracy guarantees for fitted models translate into guarantees for gradients and higher-order derivatives, with rates approaching the standard Monte Carlo rate under suitable smoothness conditions. The framework recovers established results for kernel regression, local polynomial regression, and kernel ridge regression, and yields further guarantees for multiple kernel learning and smooth neural networks. These results provide a common foundation for understanding and analyzing LTD across learning methods.
Nifei Lin, Qingkai Zhang, L. Jeff Hong
Research Institute for Interdisciplinary Sciences, School of Information Management and Engineering, Shanghai University of Finance and Economics, Shanghai 200433, China · Department of Decision Analytics and Operations, City University of Hong Kong, Hong Kong, China · Department of Industrial and Systems Engineering, University of Minnesota, Minneapolis, Minnesota 55455