Organizations: Institute for Interdisciplinary Information Sciences, Tsinghua University · Department of Statistics and Data Science, Tsinghua University
In ecology, psychometrics, and the analysis of social and financial networks, binary matrices are often analyzed conditional on their observed row and column sums, which restricts the problem to a finite sample space of matrices with the same margins. Two fundamental problems are to count this space and to sample uniformly from it. Sequential importance sampling (SIS) addresses both with independent weighted samples and an unbiased count estimator, but its efficiency depends critically on the proposal distribution. Existing proposals are analytically designed, and their accuracy can vary substantially with the margins. We show that the ideal SIS proposal, under which every weight equals the count and the variance vanishes, is exactly the policy of a generative flow network (GFlowNet) with unit reward on every matrix that has the given margins. We therefore propose MarginFlow, a framework that turns the design of the proposal into a learning problem and amortizes it across margins by exploiting their self-similarity. Every partial matrix is itself an instance with reduced margins, so one set transformer that reads the remaining margins serves every margin. We train MarginFlow on a pool of 1904 margins and evaluate it zero-shot on 1190 held-out margins, synthetic and real, from 3×3 to 870×6. On 1187 of the 1190 margins it matches or beats the best of 31 analytically designed configurations, chosen post hoc for each margin, and its median effective sample fraction is 99.8%. On the 56 margins where that best loses more than one nat of effective sample size, MarginFlow wins every one and raises the median effective sample fraction from 10.3% to 94.1%.
Figures & tables
Margins
n
CDHL
Harrison– Miller
Post-hoc best
MarginFlow (ours)
W/T/L
All
1190
41.5
97.4
99.4
99.8
415/772/3
synthetic, new seeds
880
34.8
97.2
99.4
99.8
316/563/1
synthetic, interpolated
100
60.2
98.5
99.7
99.9
28/72/0
real tables
210
53.0
97.5
99.2
99.8
71/137/2
Post-hoc best above 98%
755
56.6
98.9
99.9
99.9
0/754/1
between 90 and 98%
182
42.6
93.8
95.9
99.7
164/18/0
Table 1: Median effective sample fraction (%) on the 1190 held-out margins, by source and by the post-hoc best of the 31 configurations. The Harrison–Miller column is the network before training. W/T/L counts margins where MarginFlow is above, within 0.02 nats of, or below the post-hoc best.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Without a(t)
With a(t)
Margins
n
untrained
trained
untrained
trained
Post-hoc best
W/T/L
All
1183
0.0
99.3
97.4
99.8
99.4
9/895/279
post-hoc best above 98%
755
0.1
99.7
98.9
99.9
99.9
0/684/71
between 90 and 98%
182
0.0
98.7
93.8
99.7
95.9
0/135/47
between 37 and 90%
196
0.0
95.4
71.7
99.0
76.8
5/75/116
below 37%
50
0.0
81.6
8.3
95.1
11.4
4/1/45
Appendix
Table 2: Median effective sample fraction (%) with and without the analytically designed term a(t) , by the tiers of Table 1 , on the 1183 test margins where both are evaluated. W/T/L compares the trained network without the term to the one with it.
Margins
CDHL
Harrison–Miller
MarginFlow (ours)
121×91
1.0×1011
10.5
1.02 (1.01–1.02)
151×121
1.1×1010
2.3
1.02 (1.01–1.08)
201×151
1.9×1019
116.7
1.06 (1.02–1.16)
251×201
5.1×1017
4.8
1.07 (1.01–1.51)
301×226
4.5×1029
3011
1.20 (1.08–1.53)
376×301
2.3×1027
13.6
1.21 (1.02–2.87)
Appendix
Table 3: Draws per effective sample on the six test margins of the family of Bezáková et al. (2012) , exact for every proposal. MarginFlow is the median of the three seeds with their range.
Counter
Standard deviation of the log count
Time per margin
MarginFlow, N=4000 , one GPU
0.0007 [0.0002, 0.0022]
8.5 s
CDHL, N=4000 , one core
0.017 [0.004, 0.065]
42 s
Curveball chain, one core
0.038 [0.001, 0.081]
600 s
Curveball chain, one core
0.015 [0.002, 0.051]
3600 s
Appendix
Table 4: Standard deviation of the log count on 18 held-out margins with known counts, as median and range over the margins, and wall-clock time per margin.
MarginFlow
CDHL
Chain, 10 min
Chain, 60 min
Margins
Size
logZ
error
sd
error
sd
error
sd
error
sd
power law 0.6
6×6
8.4
0.0000
0.0003
+0.0017
0.0039
+0.001
0.001
0.000
0.002
tight columns
16×14
51.6
0.0000
0.0004
+0.0010
0.0042
+0.005
0.017
-0.005
0.012
species by site
26×12
70.7
+0.0001
0.0014
+0.0091
0.0205
+0.001
0.010
-0.011
0.015
bimodal
24×12
81.9
-0.0005
0.0017
+0.0020
0.0179
-0.015
0.025
+0.007
0.013
Southern Women
18×14
85.2
+0.0001
0.0006
-0.0002
0.0067
-0.017
0.046
+0.004
0.015
Appendix
Table 5: Counting on the 18 margins with known counts. Mean error and standard deviation of the log count for MarginFlow and CDHL are over 16 repetitions of N=4000 draws, and for the chain over four seeds at 10 and at 60 minutes of one core.
Seconds per repetition
Effective draws per second
Margin
Size
Types per state
Harrison– Miller
CDHL
Uniform
Margin- Flow
Harrison– Miller
Margin- Flow
Saturated columns
12×10
15
0.03
0.03
0.03
0.04
143 000
94 500
Species by site
20×15
2050
2.6
2.6
1.6
0.57
1 320
6 980
Web of Life
35×29
44063
13.8
13.8
11.4
9.0
282
444
Power-law
40×20
6924
3.8
3.7
3.6
0.90
471
4 320
Power-law
40×40
14198
2.6
2.7
3.7
0.65
1 460
6 160
Appendix
Table 6: Seconds for one repetition of N=4000 draws on held-out margins across the size range, on one machine with eight CPU threads and one A100 for the network’s forward pass, and effective draws per second, the repetition’s effective sample size divided by its time. Types per state is the largest number of feasible row types at any state the draws visited. The cap of Section 4.1 is checked on a census of four draws, so a full run can exceed it.
Tenth percentile
Harrison–
Post-hoc
Margin-
Post-hoc
Margin-
Margins
n
CDHL
Miller
best
Flow
best
Flow
W/T/L
by evaluation
evaluated exactly
681
65.4
98.5
99.7
99.9
88.5
98.9
176/502/3
estimated from draws
509
12.3
92.6
98.3
99.8
35.4
95.8
239/270/0
synthetic, new seeds
Appendix
Table 7: Median effective sample fraction (%) on the 1190 held-out margins by evaluation and by family, with the tenth percentile of the post-hoc best and of MarginFlow. W/T/L is as in Table 1 .
Harrison–
Post-hoc
Margin-
Margins
Size
Density
CDHL
Miller
best
Configuration
Flow
Seeds
power law 1.4
840×6
dense
0.1
0.6
0.4
HM-G 1
94.0
96.5, 93.2, 94.0
extreme sums
840×12
dense
0.2
0.7
0.5
HM-C 1
57.6
61.1, 32.7, 57.6
power law 1.4
420×6
dense
0.1
2.7
1.4
ME 0.75
98.2
97.6, 98.2, 98.2
power law 1.4
560×8
dense
0.2
0.7
1.6
HM-C 1
94.0
97.7, 93.2, 94.0
bimodal (2)
600×20
dense
0.0
0.4
2.3
ME 1
48.8
74.1, 48.8, 45.9
Appendix
Table 8: The 56 held-out margins on which the post-hoc best of the 31 configurations keeps under 37% of its draws, ordered by that fraction. Effective sample fractions in %, MarginFlow as the median of its three seeds followed by the three values.
Learning matrix-valued distributions from high-dimensional and possibly incomplete training data is challenging: ambient-space generative modeling is computationally expensive and statistically fragile when the matrix dimension is large but the sample size is limited. We propose CoreFlow, a geometry-preserving low-rank flow model that learns shared row/column subspaces across the matrix distribution, and then trains a continuous normalizing flow only on the induced low-dimensional core. CoreFlow is designed for settings where shared low-rank matrix geometry is present, especially in high-dimensional limited-sample regimes. This separates shared matrix geometry from sample-specific variation, preserves matrix structure, and substantially improves training efficiency. The same framework also handles incomplete training matrices through masked Riemannian updates and iterative completion. Across real and synthetic benchmarks, CoreFlow substantially improves spectral and moment-level generation quality in few-sample regimes while remaining competitive in data-rich settings, even under compression to 9% of the ambient dimension and with up to 40% missing training entries.
Dongze Wu, Linglingzhi Zhu, Yao Xie
H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology, Atlanta, GA
Generative Marginalization Models (MaMs) have been recently introduced as efficient neural sampling models for any-order autoregressive modelling of discrete distributions. By learning both the marginal and conditional probabilities of a persistent-block Gibbs sampler, MaMs enable fast posterior evaluation with a single neural network forward pass. While prior work has considered MaMs to be distinct from Generative Flow Networks (GFlowNets), a well-established paradigm for inference in discrete stochastic models, we show that they are equivalent. Then, we also extend MaMs' sampling strategy to non-autoregressive generative processes. In particular, we describe an automatic criterion for full-state rejuvenation of the Gibbs sampler, derived from the Gelman-Rubin statistic, which plays a key role in speeding up learning convergence. Our experiments show that our method, called Particle GFlowNets, markedly accelerates training in large combinatorial spaces.
Tiago da Silva, Diego Mesquita, Salem Lahlou
MBZUAI · School of Applied Mathematics, Getulio Vargas Foundation
Flow matching models effectively represent complex distributions, yet estimating expectations of functions of their outputs remains challenging under limited sampling budgets. Independent sampling often yields high-variance estimates, especially when rare but high-impact outcomes dominate the expectation. We propose a non-IID sampling framework that jointly draws multiple samples to cover diverse, salient regions of a flow matching model's generative distribution. To balance diversity and quality, we introduce a score-based regularization for the diversity mechanism (SR), which uses the score function, i.e., the gradient of the log probability, to ensure samples are pushed apart within high-density regions of the data manifold, mitigating off-manifold drift. To enable unbiased estimation when desired, we further develop an approach for importance weighting of non-IID flow samples by learning a residual velocity field that reproduces the marginal distribution of the non-IID samples and by evolving importance weights along trajectories. Empirically, our method produces diverse, high-quality samples and accurate importance-weight estimates and debiased expectation estimates, advancing the reliable characterization of flow matching model outputs.
Xinshuang Liu, Runfa Blark Li, Shaoxiu Wei +1
University of California, San Diego · San Diego, CA, USA