Organizations: SCSE, Macau University of Science and Technology · Department of Broadband Communication, Pengcheng Laboratory · Institute of Automation, Chinese Academy of Sciences (CASIA) · School of Artificial Intelligence, Sun Yat-sen University
Binary diffusion models typically require a large number of function evaluations (NFEs) to generate high-quality samples, making practical inference computationally expensive. Reducing NFEs while preserving sample quality without distillation or additional training remains a significant challenge. Existing binary diffusion models define a discrete one-step forward path and then derive the reverse posterior. In low-NFE settings requiring cross-step sampling, they approximate the true multi-step likelihood with a single-step likelihood transition, which severely degrades sample quality. To address this fundamental limitation and decouple the generative dynamics from fixed discrete time steps, we propose Bernoulli Flow Models (BFM). Rather than relying on sequential one-step Markov diffusion chains, BFM defines a unified continuous global Bernoulli probability flow path between data distributions and pure noise, from which we derive analytical closed-form posterior transitions over arbitrary time intervals. Consequently, reducing the inference NFE is no longer an approximation based on skipping discrete steps; it only requires re-evaluating the analytical posterior over a new time grid. This eliminates the structural training-inference mismatch inherent to discrete chains and yields self-consistent low-NFE sampling. Experiments show that BFM is highly robust to aggressive NFE reduction. On LSUN Churches 256x256, a BFM trained with 256 steps achieves an FID of 9.22 using only 16 sampling steps, whereas the state-of-the-art discrete baseline degrades to 204.10. BFM also remains competitive with continuous and discrete generative baselines under standard full-step inference. These results establish BFM as a theoretically rigorous, self-consistent, and practically effective framework for fast binary data generation.
Figures & tables
Figure 1 : Comparison of BFM and BLD under reduced inference steps and during training. (a) FID versus NFE on LSUN Churches ( 256×256 ), evaluated from the 100K-iteration checkpoints trained with 256 sampling steps using the linear probability path scheduler, where only the inference NFE is varied. (b) Training convergence comparison on LSUN Bedrooms ( 256×256 ). We visualize the FID scores from 100K to 400K training iterations ( ×103 ). Both BFM and BLD are trained under identical experimental settings using a linear scheduler. BFM reaches lower FID consistently throughout this range.
Figure 2 : Visual comparison of samples generated with different NFEs from the respective checkpoints of BLD and BFM, both trained with 256 sampling steps on LSUN Churches 256×256 . For each method, each row corresponds to a different NFE setting, while images in the same column share the same initial latent code.
Methods
Steps
LSUN-Bedrooms 256x256
LSUN-Churches 256x256
FFHQ 256x256
FID ↓
Prec. ↑
Recall ↑
FID ↓
Prec. ↑
Recall ↑
FID ↓
Prec. ↑
Recall ↑
Continuous Generative Models
StyleGAN [ 16 , 17 ]
-
2.35
0.59
0.48
3.86
0.60
0.43
4.16
0.71
0.46
LDM-4/8/4 [ 24 ]
200
2.95
0.66
0.48
4.02
0.64
0.52
4.98
0.73
0.50
Patch-DM [ 9 ]
50
6.04
0.56
0.44
5.49
0.62
0.53
10.02
0.68
0.44
VQ-LCMD [ 22 ]
200
4.16
0.72
0.40
4.99
0.75
0.42
7.25
0.72
0.46
Table 1 : Comparison of various methods for image generation on LSUN Bedrooms, LSUN Churches, and FFHQ. All images are of resolution 256×256 . BLD* denotes our reproduction of BLD using its official codebase under the same machine environment and training configuration as BFM. The reproduced BLD* and BFM results are evaluated from checkpoints trained for 800K iterations. Within each method category, the best and second-best results in each column are highlighted in bold and underline , respectively.
Figure 3 : Samples from BFM on LSUN-Bedrooms, LSUN-Churches and FFHQ datasets. All samples resolution are 256x256.
Figure 4 : Left: Binarized MNIST samples generated by BFM. Right: Comparison of FID scores. The best and second-best results are highlighted in bold and underline , respectively.
Figure 5 : Bernoulli probability path schedulers.
Scheduler
FID ↓
BFM
BLD
Linear
9.22
9.55
Half-cosine
9.41
9.53
Reverse half-cos
9.27
9.54
Table 2 : FID comparison of different Bernoulli probability path schedulers for BFM and BLD on LSUN Churches 256 × 256, evaluated at 100K training iterations.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Unconditional 256×256 samples on LSUN Bedrooms. Sampling temperatures κ are linearly interpolated from 0.5 to 1.0 (left to right).
Figure 7 : Unconditional 256×256 samples on LSUN Churches. Sampling temperatures κ are linearly interpolated from 0.5 to 1.0 (left to right).
Figure 8 : Unconditional 256×256 samples on FFHQ. Sampling temperatures κ are linearly interpolated from 0.5 to 1.0 (left to right).
Figure 9 : Comparison of 20×20 Ising lattice samples generated by BFM and ground-truth samples at temperatures 2.0 and 2.5.
Figure 10 : Comprehensive evaluation of Ising model generation across the continuous temperature spectrum (1.50 to 3.10). (a) and (b) show the visual evolution of spin lattices generated by BFM and the ground-truth Wolff cluster [ 33 ] algorithm, respectively. (c) Quantitative comparison of the absolute magnetization M(T) , demonstrating that BFM accurately captures the thermodynamic properties.
Component
LSUN & FFHQ
Binarized MNIST
Ising Model
Transformer layers
24
24
24
Attention heads
12
12
12
Embedding dimension
768
768
768
Sequence length (block size)
256
28
20
Training iterations
500K
500K
100K
Batch size
96
512
512
Appendix
Table 3 : Configuration and training details for BFM across different datasets.
Task
Representation
Batch Size
Sampling Steps
Avg. Time / Sample (s)
Peak GPU Memory (GB)
LSUN / FFHQ Image-256
Binary latent + BAE decoder
1
64
0.3837
1.0743
Binarized MNIST
Native binary space
1
64
0.3507
0.6482
Ising Model
Native binary space
1
64
0.3531
0.6480
Appendix
Table 4 : Single-sample inference resource usage. We report the average wall-clock latency and peak GPU memory for one generated sample using 64 sampling steps. The Image-256 setting is shared by LSUN and FFHQ since they use the same image resolution and binary latent format.
Flow-based generative modeling in continuous spaces exploit Tweedie's formula to express the denoiser (learned in training) as a score function (used in sampling). In contrast, this relation has been largely missing in the discrete setting where common approaches focus on learning discrete scores and rates. In this work we close this gap for discrete non-negative ordinal data by introducing Binomial flows. Our framework provides a simple recipe for training a discrete diffusion model which simultaneously denoises, samples, and estimates exact likelihoods. We verify our methodology on synthetic examples and obtain competitive results on real-world data sets.
Discrete diffusion models are a powerful class of generative models with strong performance across many domains. For efficiency, however, discrete diffusion typically parameterizes the generative (reverse) process with factorized distributions, which makes it difficult for the model to learn the target process in a small number of steps and necessitates a long, computationally expensive sampling procedure. To reduce the gap between the target and model distributions and enable few-step generation, we propose Forward-Learned Discrete Diffusion (FLDD), which introduces discrete diffusion with a learnable forward (noising) process. Rather than fixing a Markovian forward chain, we adopt a non-Markovian formulation with learnable marginal and posterior distributions. This allows the generative process to remain factorized while matching the target defined by the noising process. We train all parameters end-to-end under the standard variational objective. Experiments on various benchmarks show that, for a given number of sampling steps, our approach produces a higher quality samples than conventional discrete diffusion models using the same reverse parameterization.
Generative modeling over discrete structures underpins applications across deep learning, from biological sequence design and code generation to large language models, yet generation often remains sequential, relying on autoregressive decoding or iterative refinement. In this work, we introduce Coupling Models(Coupling Models), a one-step discrete generative model that learns a direct coupling between discrete sequences and Gaussian latents. Unlike recent distillation methods that compress a pretrained multi-step sampler into a few steps, Coupling Model trains a purpose-built decoder to invert this coupling and generate samples in a single step. The model also avoids complex continuous flows over the simplex and hand-specified data-to-noise couplings. Empirically,Coupling Model improves the strongest one-step baselines in each domain: it reduces LM1B text-generation perplexity by 33% at its lowest-perplexity operating point, Fly Brain enhancer-design FBD by 18%, and MNIST-Binary FID by 46%. These results suggest that effective one-step discrete generation depends strongly on how data and noise are coupled before decoding. Code is available at https://github.com/pengzhangzhi/Coupling-Models.
Fred Zhangzhi Peng, Avishek Joey Bose, Anru R. Zhang +1
Duke University · Imperial College London · AITHYRA