Systematic Hazard Sampling: Minimal-Variance Inference for Discrete Diffusion and Flow Models
Organizations: Department of EE KAIST Daejeon, South Korea
Abstract
Uniform-noise discrete diffusion and flow models generate sequences non-autoregressively through iterative, context-dependent token replacements. However, these models are typically formulated as time-inhomogeneous continuous- or discrete-time Markov chains (CTMC/DTMC), sampled using independent Bernoulli change decisions per discretization step. This induces Poisson-binomial variance in per-position jump counts that grows with the number of required edits, leading to the common under-editing (residual noise) and over-editing (cascading substitutions) failure modes that degrade sample quality. We identify this sampler-induced variance as an orthogonal source of degradation, distinct from model-side errors and addressable purely at inference time. We propose Systematic Hazard Sampling (SHS), a training-free, drop-in, and hyperparameter-free inference principle for any sampler that admits a stay-vs.-replace decomposition. SHS models per-token edits as events driven by cumulative hazard (CTMC) or jump mass (DTMC) and triggers an edit whenever this quantity exceeds unit-spaced thresholds with a single random phase per position. For any fixed cumulative mass, this preserves the expected jump count while achieving the minimum conditional variance possible among unbiased integer estimators (at most 1/4), without altering per-jump destination sampling. Experiments on four uniform-noise discrete diffusion and flow language models spanning ~110M to ~3B parameters show that SHS consistently improves sample quality across numbers of function evaluations (NFE). Code is available at https://github.com/Jang-seunghwan/Systematic-Hazard-Sampling.
Figures & tables
| UDLM | GIDD | |||||||||
| NFE | Standard | SHS | Standard | SHS | ||||||
| Entropy | Gen. PPL | Entropy | Gen. PPL | Entropy | Gen. PPL | Entropy | Gen. PPL | |||
| 4 | ||||||||||
| 8 | ||||||||||
| 16 | ||||||||||
| 32 | ||||||||||
| UDLM | GIDD | |||||||||||||||
| NFE | Standard | SHS | Standard | SHS | ||||||||||||
| D-1 | D-2 | D-3 | S-BLEU | D-1 | D-2 | D-3 | S-BLEU | D-1 | D-2 | D-3 | S-BLEU | D-1 | D-2 | D-3 | S-BLEU | |
| 4 | ||||||||||||||||
| 8 | ||||||||||||||||
| 16 | ||||||||||||||||
| 32 | ||||||||||||||||
| NFE | Standard observed | Standard prediction | SHS observed | SHS bound |
| 4 | ||||
| 16 | ||||
| 64 | ||||
| 128 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| NFE | Standard | Adaptive (top-1) | SHS |
| 4 | 6.758 / 936.3 | 8.248 / 1995.0 | 6.634 / 869.5 |
| 8 | 6.660 / 443.4 | 8.384 / 1483.8 | 6.536 / 417.2 |
| 16 | 6.612 / 221.4 | 8.459 / 1094.8 | 6.554 / 218.1 |
| 32 | 6.649 / 150.2 | 8.454 / 783.7 | 6.588 / 142.0 |
| 64 | 6.654 / 113.7 | 8.421 / 591.9 | 6.619 / 113.0 |
| 128 | 6.649 / 104.9 | 8.483 / 574.9 | 6.618 / 99.6 |
| Model | Standard | top- | top- + SHS |
| UDLM | 6.884 / 190.65 | 6.866 / 187.45 | 6.817 / 174.55 |
| GIDD | 7.350 / 111.63 | 7.299 / 97.56 | 7.206 / 78.23 |
| GIDD, | Temperature | Temperature + SHS | |
| 1.0 | 7.350 / 111.63 | 7.234 / 84.83 | |
| 0.99 | 7.301 / 99.89 | 7.182 / 76.09 | |
| 0.975 | 7.218 / 81.91 | 7.118 / 65.02 |
| D-1 | D-2 | D-3 | Self-BLEU-2 | |
| Real data | 0.148 | 0.626 | 0.903 | 0.062 |
| NFE 16 | 0.180 / 0.171 | 0.702 / 0.688 | 0.958 / 0.955 | 0.057 / 0.059 |
| NFE 128 | 0.168 / 0.153 | 0.677 / 0.656 | 0.946 / 0.941 | 0.053 / 0.056 |
| NFE 512 | 0.169 / 0.153 | 0.678 / 0.654 | 0.946 / 0.939 | 0.052 / 0.054 |
| Phase | Gen. PPL | paired vs. per-position | |
| Independent per position (SHS) | 174.248 | – | 5.0789 |
| Shared across positions | 213.477 | 0.0039 | 5.0868 |
| Redrawn at each unit interval | 173.711 | 0.799 | – |
| NFE | NFE | |||
| under-edit ( ) | over-edit ( ) | under-edit ( ) | over-edit ( ) | |
| NFE | |||
| 2 | |||
| 4 | |||
| 8 | |||
| 16 | |||
| 32 | |||
| 64 |