cs.LGSep 30, 2026

Same Loss, Different Gradients

Authors: Ningkang Peng, Xiaoqian Peng, Yifan He, Anjie Hu, Chao Tan, Peirong Ma, Yanhui Gu

Organizations: Nanjing Normal University · Nanjing University of Chinese Medicine

Abstract

Differentiable learning typically assumes that the scalar objective evaluated in the forward pass and the gradient supplied to the optimizer in the backward pass describe the same mathematical object. We show that this correspondence can fail when probabilistic objectives rely on finite special-function recurrences, custom backward rules, and numerical clipping. In high-dimensional von Mises-Fisher learning, real numerical implementations can produce identical forward scores and losses at the same learning state while supplying different gradients and following different optimization trajectories. We characterize the structure of this mismatch in finite-start Bessel recurrence and show that classwise radial mismatch can compose through probabilities into a locally nonconservative update field. Evaluating the accuracy of special-function values and derivatives separately is therefore insufficient to characterize the realized learning objective. Motivated by this observation, we introduce AR/FR, a fixed-depth analytic realization that constructs a potential and its derivative jointly, ensuring forward-backward coherence by construction. We establish a uniform cubic-order error bound relative to the exact Bessel ratio over the entire nonnegative concentration axis and propagate this guarantee to learning scores and objectives. As representation dimension increases, the original finite recurrence becomes sequentially deeper, whereas the worst-case AR/FR error guarantee tightens cubically, jointly providing coherence, certified fidelity, and fixed-depth computation. These results suggest that a differentiable numerical primitive is defined by both the values it realizes and the derivatives it actually supplies to the optimizer; together, they constitute the numerical realization of the learning algorithm.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 30, 2026cs.LG

How Accurate Is Accurate Enough?

How accurate must a numerical approximation be within a learning system? Primitive error alone cannot answer this question: errors of the same magnitude can have very different consequences for losses, predictions, and gradients at different learning states. We study this question through the learning objective itself. The objective weights classwise numerical errors nonuniformly according to the current state, so the importance of an error depends not only on its magnitude but also on the class it affects and the weight that class receives. For softmax cross-entropy, we characterize this coupling between class weights and errors and derive the exact extrema of the signed loss change over pairings of fixed non-target probability and score-error multisets, with the target probability and target score error held fixed. Building on this structure, we establish finite-error guarantees that propagate primitive error to losses, probabilities, predictions, and feature gradients, then invert these guarantees to obtain a certified primitive tolerance for the current state under prescribed learning-level error requirements. We give a complete instantiation of the framework in high-dimensional von Mises-Fisher learning. Controlled interventions and a large collection of saved learning states show that identical primitive error can produce substantially different learning consequences, while certified numerical tolerances vary by orders of magnitude across states under the same learning-level requirements. These results show that the adequacy of a numerical approximation must be assessed in relation to the current learning state and the quantity to be preserved; numerical accuracy should itself be treated as part of the learning objective.
Sep 30, 2026stat.CO

Learn-Then-Differentiate Gradient Estimation

Learn-then-differentiate (LTD) estimates gradients by fitting a model to simulation outputs and differentiating it. We develop a unified framework explaining what LTD differentiates and how accurately it estimates gradients. For models with a weighted representation, LTD differentiates a learned representation of the underlying probability measure. We then show how accuracy guarantees for fitted models translate into guarantees for gradients and higher-order derivatives, with rates approaching the standard Monte Carlo rate under suitable smoothness conditions. The framework recovers established results for kernel regression, local polynomial regression, and kernel ridge regression, and yields further guarantees for multiple kernel learning and smooth neural networks. These results provide a common foundation for understanding and analyzing LTD across learning methods.
Jul 9, 2026stat.ML

Score Accuracy Along the Forward Diffusion Does Not Certify Numerical Stability in Diffusion Sampling

Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory. We show that small forward-marginal error does not guarantee numerical stability. We construct a single smooth score field with arbitrarily small forward-marginal L2L^2 error. The learned reverse-time process is nonexplosive, has moments of every order, and can be arbitrarily close to the exact reverse-time process in path-space total variation. Yet its Euler--Maruyama discretizations converge in probability while every positive moment diverges. Thus weak convergence can hold even though every Wasserstein distance WpW_p, p≥1p\ge1, diverges. The same failure can occur within one fixed finite neural architecture. We construct a family of bounded, globally Lipschitz denoisers for which both the forward-marginal error and the path-space total variation distance tend to zero, while their Euler--Maruyama endpoints diverge in every WpW_p. For compactly supported data, we also give a simple positive result. Projecting the learned denoiser onto a known bounded closed convex set containing the support preserves pointwise accuracy, gives grid-uniform moment bounds, and yields Wasserstein convergence under mild local regularity. Experiments with a small fixed DiT-style network show large growth along rare numerical trajectories and its suppression by denoiser projection, while overall trajectory errors remain small.