Organizations: SCSE, Macau University of Science and Technology · Department of Broadband Communication, Pengcheng Laboratory · Institute of Automation, Chinese Academy of Sciences (CASIA) · School of Artificial Intelligence, Sun Yat-sen University
Binary diffusion models typically require a large number of function evaluations (NFEs) to generate high-quality samples, making practical inference computationally expensive. Reducing NFEs while preserving sample quality without distillation or additional training remains a significant challenge. Existing binary diffusion models define a discrete one-step forward path and then derive the reverse posterior. In low-NFE settings requiring cross-step sampling, they approximate the true multi-step likelihood with a single-step likelihood transition, which severely degrades sample quality. To address this fundamental limitation and decouple the generative dynamics from fixed discrete time steps, we propose Bernoulli Flow Models (BFM). Rather than relying on sequential one-step Markov diffusion chains, BFM defines a unified continuous global Bernoulli probability flow path between data distributions and pure noise, from which we derive analytical closed-form posterior transitions over arbitrary time intervals. Consequently, reducing the inference NFE is no longer an approximation based on skipping discrete steps; it only requires re-evaluating the analytical posterior over a new time grid. This eliminates the structural training-inference mismatch inherent to discrete chains and yields self-consistent low-NFE sampling. Experiments show that BFM is highly robust to aggressive NFE reduction. On LSUN Churches 256x256, a BFM trained with 256 steps achieves an FID of 9.22 using only 16 sampling steps, whereas the state-of-the-art discrete baseline degrades to 204.10. BFM also remains competitive with continuous and discrete generative baselines under standard full-step inference. These results establish BFM as a theoretically rigorous, self-consistent, and practically effective framework for fast binary data generation.
Figures & tables
Figure 1 : Comparison of BFM and BLD under reduced inference steps and during training. (a) FID versus NFE on LSUN Churches ( 256×256 ), evaluated from the 100K-iteration checkpoints trained with 256 sampling steps using the linear probability path scheduler, where only the inference NFE is varied. (b) Training convergence comparison on LSUN Bedrooms ( 256×256 ). We visualize the FID scores from 100K to 400K training iterations ( ×103 ). Both BFM and BLD are trained under identical experimental settings using a linear scheduler. BFM reaches lower FID consistently throughout this range.
Figure 2 : Visual comparison of samples generated with different NFEs from the respective checkpoints of BLD and BFM, both trained with 256 sampling steps on LSUN Churches 256×256 . For each method, each row corresponds to a different NFE setting, while images in the same column share the same initial latent code.
Methods
Steps
LSUN-Bedrooms 256x256
LSUN-Churches 256x256
FFHQ 256x256
FID ↓
Prec. ↑
Recall ↑
FID ↓
Prec. ↑
Recall ↑
FID ↓
Prec. ↑
Recall ↑
Continuous Generative Models
StyleGAN [ 16 , 17 ]
-
2.35
0.59
0.48
3.86
0.60
0.43
4.16
0.71
0.46
LDM-4/8/4 [ 24 ]
200
2.95
0.66
0.48
4.02
0.64
0.52
4.98
0.73
0.50
Patch-DM [ 9 ]
50
6.04
0.56
0.44
5.49
0.62
0.53
10.02
0.68
0.44
VQ-LCMD [ 22 ]
200
4.16
0.72
0.40
4.99
0.75
0.42
7.25
0.72
0.46
Table 1 : Comparison of various methods for image generation on LSUN Bedrooms, LSUN Churches, and FFHQ. All images are of resolution 256×256 . BLD* denotes our reproduction of BLD using its official codebase under the same machine environment and training configuration as BFM. The reproduced BLD* and BFM results are evaluated from checkpoints trained for 800K iterations. Within each method category, the best and second-best results in each column are highlighted in bold and underline , respectively.
Figure 3 : Samples from BFM on LSUN-Bedrooms, LSUN-Churches and FFHQ datasets. All samples resolution are 256x256.
Figure 4 : Left: Binarized MNIST samples generated by BFM. Right: Comparison of FID scores. The best and second-best results are highlighted in bold and underline , respectively.
Figure 5 : Bernoulli probability path schedulers.
Scheduler
FID ↓
BFM
BLD
Linear
9.22
9.55
Half-cosine
9.41
9.53
Reverse half-cos
9.27
9.54
Table 2 : FID comparison of different Bernoulli probability path schedulers for BFM and BLD on LSUN Churches 256 × 256, evaluated at 100K training iterations.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Unconditional 256×256 samples on LSUN Bedrooms. Sampling temperatures κ are linearly interpolated from 0.5 to 1.0 (left to right).
Figure 7 : Unconditional 256×256 samples on LSUN Churches. Sampling temperatures κ are linearly interpolated from 0.5 to 1.0 (left to right).
Figure 8 : Unconditional 256×256 samples on FFHQ. Sampling temperatures κ are linearly interpolated from 0.5 to 1.0 (left to right).
Figure 9 : Comparison of 20×20 Ising lattice samples generated by BFM and ground-truth samples at temperatures 2.0 and 2.5.
Figure 10 : Comprehensive evaluation of Ising model generation across the continuous temperature spectrum (1.50 to 3.10). (a) and (b) show the visual evolution of spin lattices generated by BFM and the ground-truth Wolff cluster [ 33 ] algorithm, respectively. (c) Quantitative comparison of the absolute magnetization M(T) , demonstrating that BFM accurately captures the thermodynamic properties.
Component
LSUN & FFHQ
Binarized MNIST
Ising Model
Transformer layers
24
24
24
Attention heads
12
12
12
Embedding dimension
768
768
768
Sequence length (block size)
256
28
20
Training iterations
500K
500K
100K
Batch size
96
512
512
Appendix
Table 3 : Configuration and training details for BFM across different datasets.
Task
Representation
Batch Size
Sampling Steps
Avg. Time / Sample (s)
Peak GPU Memory (GB)
LSUN / FFHQ Image-256
Binary latent + BAE decoder
1
64
0.3837
1.0743
Binarized MNIST
Native binary space
1
64
0.3507
0.6482
Ising Model
Native binary space
1
64
0.3531
0.6480
Appendix
Table 4 : Single-sample inference resource usage. We report the average wall-clock latency and peak GPU memory for one generated sample using 64 sampling steps. The Image-256 setting is shared by LSUN and FFHQ since they use the same image resolution and binary latent format.