Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation
Organizations: Department of Applied Mathematics and Statistics, Johns Hopkins University
Abstract
Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample. For discrete variables, however, the pathwise identity cannot generally be exact for every differentiable function. We propose a general framework to construct finite-order exact pathwise gradient estimators for a range of common discrete variables such as Poisson. The estimator is the least-norm solution among all solutions that are unbiased for polynomials of degree at most. The resulting estimators preserve the hard forward sample, require no temperature tuning, and can be implemented in a few lines of codes. Against other admissible solutions, our estimator is unique and minimizes weight variance; in contrast, prior works use categorical variables or augmented representations to approximate non-categorical variables that induces excess variance and computations. To understand approximation bias for functions beyond the prescribed class, we also derive a non-asymptotic bias bound. In experiments our low order methods match or improve tuned baselines across linear, nonlinear and hierarchical latent-variable models, while out-speeding competitors in every runtime benchmark.
Figures & tables
| Dataset | Fini-3 | Fini-2 | EAT-cubic (0.1) | GSM (0.5) | EAT-sigmoid (0.5) | Score |
|---|---|---|---|---|---|---|
| Synthetic |
| Bernoulli Likelihoods | Gaussian Likelihoods | |||||
|---|---|---|---|---|---|---|
| MNIST | Fashion-MNIST | Omniglot | MNIST | Fashion-MNIST | Omniglot | |
| Fini-3 | ||||||
| Fini-2 | ||||||
| EAT-cubic | (0.5) | (0.5) | (0.2) | (0.5) | (0.5) | (0.5) |
| ReinMax | (1.1) | (1.1) | (1.0) | (1.0) | (1.1) | (1.0) |
| 20 Newsgroups | RCV1 | ||||||
|---|---|---|---|---|---|---|---|
| Model | Training method | Train ELBO | Test ELBO | PPL | Train ELBO | Test ELBO | PPL |
| Poisson DEF | Fini-3 | ||||||
| Fini-2 | |||||||
| EAT-cubic (0.2) | |||||||
| ReinMax (1.3) | |||||||
| Family | Latent dim. | Fini-2 | Fini-3 | ReinMax |
|---|---|---|---|---|
| binomial | 2 | (1.0) | ||
| 16 | (1.0) | |||
| 128 | (1.1) | |||
| negbin | 2 | (1.0) | ||
| 16 | (1.0) | |||
| 128 | (1.1) |
| Linear VAE | Nonlinear VAE | DEF | ||||
|---|---|---|---|---|---|---|
| Estimator | L4 | A100 | L4 | A100 | L4 | A100 |
| Fini-3 | ||||||
| Fini-2 | ||||||
| EAT-cubic | ||||||
| ReinMax | ||||||
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.