Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample. For discrete variables, however, the pathwise identity cannot generally be exact for every differentiable function. We propose a general framework to construct finite-order exact pathwise gradient estimators for a range of common discrete variables such as Poisson. The estimator is the least-norm solution among all solutions that are unbiased for polynomials of degree at most. The resulting estimators preserve the hard forward sample, require no temperature tuning, and can be implemented in a few lines of codes. Against other admissible solutions, our estimator is unique and minimizes weight variance; in contrast, prior works use categorical variables or augmented representations to approximate non-categorical variables that induces excess variance and computations. To understand approximation bias for functions beyond the prescribed class, we also derive a non-asymptotic bias bound. In experiments our low order methods match or improve tuned baselines across linear, nonlinear and hierarchical latent-variable models, while out-speeding competitors in every runtime benchmark.
Figures & tables
Figure 1: Fini te-order pathwise gradients. Left : pathwise gradients are unbiased for all differentiable functions - this is too rigid for discrete variables, so we relax the unbiasedness requirement to low-order polynomials; Center : generally, there are many solutions to the finite-order unbiasedness requirement; we use the least-norm principle to select the order-D unbiased solution, avoiding excess weight variance; Right : as an example, Fini-2 and 3 for Poisson reduce to a few lines of code.
Figure 2: Linear Poisson VAE with 512 latent dimensions. Left : mean single-sample gradient error relative to the analytic gradient early (epoch 100) and partway through training (epoch 1,000), shown as dots with standard-deviation bars. Both means are near zero; Fini-2 has about one-fifth the standard deviation of tuned EAT-cubic. Right : validation ELBO over training, mean ± standard deviation across five matched seeds. Fini-2 follows the analytic-gradient trajectory more closely.
Dataset
Fini-3
Fini-2
EAT-cubic (0.1)
GSM (0.5)
EAT-sigmoid (0.5)
Score
Synthetic
−330.61±0.79
−330.60±0.75
−330.62±0.79
−331.94±0.95
−332.18±0.97
−333.23±1.51
Table 1: Synthetic POGLM validation ELBO (higher is better), mean ± std. over 30 seeds. Parentheses give the selected temperature τ . Boldface marks the three highest means, which differ by at most 0.02 .
Bernoulli Likelihoods
Gaussian Likelihoods
MNIST
Fashion-MNIST
Omniglot
MNIST
Fashion-MNIST
Omniglot
Fini-3
−118.60±0.26
−251.98±0.40
−129.38±0.07
595.48±3.72
133.74±2.80
403.14±2.46
Fini-2
−118.88±0.19
−252.26±0.18
−129.93±0.12
594.47±3.92
132.09±3.29
403.80±2.00
EAT-cubic
−120.06±0.17 (0.5)
−253.04±0.22 (0.5)
−130.76±0.07 (0.2)
595.66±3.60 (0.5)
132.42±1.70 (0.5)
400.84±2.05 (0.5)
ReinMax
−125.87±0.38 (1.1)
−260.56±0.57 (1.1)
−140.07±0.40 (1.0)
589.21±2.78 (1.0)
104.06±2.51 (1.1)
396.63±1.53 (1.0)
Table 2: Nonlinear Poisson-VAE training ELBO ↑ , mean ± standard deviation over five matched seeds. Columns pair each dataset with a Bernoulli or Gaussian observation likelihood. First is bold, second is underlined.
20 Newsgroups
RCV1
Model
Training method
Train ELBO ↑
Test ELBO ↑
PPL ↓
Train ELBO ↑
Test ELBO ↑
PPL ↓
Poisson DEF
Fini-3
−179.30±2.81
−199.31±2.04
660.7±7.6
−269.27±3.16
−267.64±3.73
379.3±5.4
Fini-2
−179.30±2.62
−198.98±1.92
659.7±6.7
−269.20±3.05
−267.60±3.73
381.9±4.6
EAT-cubic (0.2)
−180.94±2.59
−199.77±2.17
664.6±8.5
−271.92±3.18
−270.16±3.82
388.4±4.1
ReinMax (1.3)
−191.25±2.88
−210.01±2.37
638.1±10.5
−301.86±2.83
−300.05±3.27
384.4±4.8
Table 3: Poisson DEF results on train and test ELBO ( ↑ ) and heldout perplexity (PPL, ↓ ), mean ± standard deviation over five matched seeds. Best is bold; second is underlined.
Family
Latent dim.
Fini-2
Fini-3
ReinMax
binomial n=8,p=0.5
2
−174.51±0.45
−176.91±1.07
−179.67±0.67 (1.0)
16
−123.38±0.14
−123.11±0.10
−125.55±0.10 (1.0)
128
−118.73±0.22
−118.43±0.17
−121.45±0.15 (1.1)
negbin r=8,p=0.5
2
−177.34±0.15
−177.31±0.24
−182.58±0.48 (1.0)
16
−129.29±0.16
−129.25±0.20
−136.58±0.71 (1.0)
128
−122.58±0.06
−122.54±0.36
−124.41±0.24 (1.1)
Table 4: Count-VAE training ELBO, mean ± standard deviation over five independent matched seeds. Best is bold; second is underlined.
Figure 3: Negative-Binomial representation diagnostic. Native Fini-2 and Fini-3 have lower gradient variance than matched Gamma–Poisson counterparts; ReinMax has higher variance.
Linear VAE
Nonlinear VAE
DEF
Estimator
L4
A100
L4
A100
L4
A100
Fini-3
0.326±0.010
0.361±0.008
1.868±0.074
2.016±0.083
0.928±0.021
0.907±0.016
Fini-2
0.309±0.005
0.338±0.004
1.768±0.017
1.848±0.046
0.863±0.014
0.835±0.016
EAT-cubic
1.427±0.005
0.581±0.010
2.680±0.040
2.942±0.073
1.546±0.039
1.565±0.015
ReinMax
1.144±0.006
0.568±0.009
2.692±0.035
2.859±0.104
1.451±0.024
1.610±0.134
Table 5: Wall-clock(s) per epoch ( ↓ ), mean ± standard deviation over five repeats.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Local rate–distortion update directions in selected high-performing runs. Our estimators closely track the near-exact MVD reference; EAT-cubic is close on 20 Newsgroups but deviates more on RCV1, while ReinMax tilts toward lower distortion and higher rate.
Policy-gradient methods usually optimize expected return, but many real world applications care about distributional properties of returns: tail risk, outlier robustness, or best-of-K discovery. We introduce OrderGrad, a family of likelihood-ratio and reparameterization gradient estimators for order-statistic objectives. OrderGrad optimizes finite-sample L-statistics, i.e., weighted averages of sorted rewards or costs, recovering objectives such as VaR, CVaR, trimmed means, medians, and top-m/best-of-K criteria by changing only the rank weights. For any fixed sample size and rank-weight vector, OrderGrad provides an unbiased gradient estimator for the corresponding order-statistic objective. The method is implemented as a simple reward transformation that can then be used in an otherwise standard policy-gradient or reparameterized update. We study the resulting estimator's variance behavior and evaluate it on tasks where mean optimization is mismatched to the deployment objective, including LLM math post-training and other tasks. OrderGrad provides a unified, plug-and-play route to risk-averse, robust, and exploratory learning. Code: https://github.com/paavo5/ordergrad
Gradient estimation -- the task of computing the gradient of the expected value of a probabilistic program -- has diverse applications in scientific computing, but is notoriously difficult because of issues such as high-dimensional integration, discrete random choices, and complex stochastic dependencies. This article introduces gradient inference, a new approach to developing sound and efficient gradient estimators for probabilistic programs. Gradient inference rests on a formal reduction from a gradient estimation problem to a closely related probabilistic inference problem, whose solution can be differentiated to obtain a gradient estimator. This inference problem is obtained by applying two powerful statistical operations -- coupling and factorization -- to the input probabilistic program. Our reduction lets us leverage the rich toolkit of probabilistic inference algorithms to design novel gradient estimators that extend and improve upon existing methods. We introduce GradInf, a probabilistic programming system that facilitates the sound and automated implementation of gradient inference. GradInf is centered around programmable source-to-source transformations for coupling and factorizing higher-order probabilistic programs, whose soundness is proven in terms of a denotational semantics. Key to our development is the use of information-flow typing to allow random choices in a probabilistic program to be factored out and partially evaluated, which improves our ability to deploy sophisticated probabilistic inference algorithms. The resulting system offers practitioners a principled framework for designing gradient estimators. We apply GradInf to several challenging case studies, showing that it can express prominent gradient estimators from the literature and enables the construction of new state-of-the-art estimators that outperform the best existing baselines.
Gaurav Arya, Mathieu Huot, Moritz Schauer +2
Carnegie Mellon University, Pittsburgh, USA · Massachusetts Institute of Technology, Cambridge, USA · Chalmers University of Technology & University of Gothenburg, Gothenburg, Sweden +1
In policy gradient reinforcement learning, access to a differentiable model enables 1st-order gradient estimation that accelerates learning compared to relying solely on derivative-free 0th-order estimators. However, discontinuous dynamics cause bias and undermine the effectiveness of 1st-order estimators. Prior work addressed this bias by constructing a confidence interval around the REINFORCE 0th-order gradient estimator and using these bounds to detect discontinuities. However, the REINFORCE estimator is notoriously noisy, and we find that this method requires task-specific hyperparameter tuning and has low sample efficiency. This paper asks whether such bias is the primary obstacle and what minimal fixes suffice. First, we re-examine standard discontinuous settings from prior work and introduce DDCG, a lightweight test that switches estimators in nonsmooth regions; with a single hyperparameter, DDCG achieves robust performance and remains reliable with small samples. Second, on differentiable robotics control tasks, we present IVW-H, a per-step inverse-variance implementation that stabilizes variance without explicit discontinuity detection and yields strong results. Together, these findings indicate that while estimator switching improves robustness in controlled studies, careful variance control often dominates in practical deployments.