Organizations: The University of Tokyo, Tokyo, Japan · Hasso Plattner Institute for Digital Engineering, Germany · Institute of Statistical Mathematics, Japan · The Graduate University for Advanced Studies, SOKENDAI, Japan · RIKEN Centre for Advanced Intelligence Project (AIP), Tokyo, Japan · Hasso Plattner Institute for Digital Health at the Icahn School of Medicine at Mount Sinai, USA
Modern deep generative models are primarily studied for their ability to generate realistic samples, yet the generative dynamics they learn can also serve as objects of statistical inference. We develop this idea for two-sample testing, the problem of deciding whether the same distribution generated two finite datasets. Using stochastic interpolants, we connect both distributions to a shared Gaussian bottleneck, so that each half of the resulting path is a Gaussian channel acting on a single population. We prove that the null hypothesis holds if and only if the population denoiser, or equivalently, the velocity fields of the two halves, coincide at any single noise level, which amounts to a reflection symmetry of the path about the bottleneck. Deviations from this symmetry yield a continuum of two-sample witnesses, which we estimate via held-out regression risks on learned denoisers and velocities and aggregate along the path; under an information-theoretic weighting, the aggregated discrepancy equals the Jeffreys divergence between the noise-smoothed distributions. Calibrating the resulting statistics by permutation yields tests that are valid in finite samples for any trained networks and consistent when the fields are learned accurately. On a synthetic benchmark and three image benchmarks, the proposed tests improve power over the strongest baseline by up to 33 percentage points at an equal total sample budget, with the best choice of regression representation and path weighting depending on the data modality. These results show that generative paths provide a principled representation for statistical testing, extending stochastic-interpolant models beyond generation.
Figures & tables
Figure 1: Path-based two-sample testing with a Gaussian-bottleneck interpolant. Samples from ρ0 are transported by the probability flow (solid curves) to the shared bottleneck N(0,21Id) at t=21 (open circles) and onwards to ρ1 , while dashed curves show the first half of each path mirrored under t↦1−t . Left: under H0 , the path is reflection symmetric about the bottleneck, so matched slices such as ρ1/4 and ρ3/4 coincide, each trajectory retraces its first half, and the mirrored curves are hidden beneath the flow. Right: under H1 , the second half of each path departs from the mirror image of the first, and this asymmetry, measured through the denoiser and velocity fields at matched noise levels, is the two-sample signal.
Variant
Fitted fields
Regression target y
Path weight w(s)
PBI-Denoiser_U
ηs0,ηs1
ω
1
PBI-Denoiser_J
ηs0,ηs1
ω
λsas/σs2
PBI-Velocity_U
us0,us1
as′ω+σs′z
1
PBI-Velocity_J
us0,us1
as′ω+σs′z
as/(λsσs2)
Table 1: The four PBI variants, each combining a regression representation with a path weight and run with Algorithms 1 , 2 and 3 . The two Jeffreys weights give the same population target, the Jeffreys divergence of Proposition 7 up to path truncation. All weights are positive and bounded on Iτ .
Figure 2: Test power at level κ=0.05 as a function of the per-population sample size n ( ntr=nte=n ; every method receives 4n observations in total). MNIST, OrganAMNIST and Fashion-MNIST use class contamination at a fraction 0.1 (3 → 8, left → right lung, sneaker → ankle boot). Lines show the mean over five training seeds (50 test sets each); shaded bands show ±1 standard deviation across seeds for the PBI variants. Highlighted baselines are those that are strongest at some sample size; the remaining baselines are shown in light grey. The HDGM axis is logarithmic.
Figure 3: Empirical Type-I error at level κ=0.05 , with both samples drawn from ρ0 , as a function of the per-population sample size n ( ntr=nte=n ). Lines show the mean rejection rate over five training seeds (50 test sets each); shaded bands show ±1 standard deviation across seeds for the PBI variants, and the dotted line marks κ . Baselines are styled as in Figure 2 , and the HDGM axis is logarithmic. The PBI variants remain near the nominal level on all four benchmarks, as guaranteed by Proposition 9 .
Figure 4: Samples from the fitted MNIST models at ntr=nte=50 per population, for (a) the denoiser and (b) the velocity representation. Rows show, from top to bottom, held-out observations from ρ0 (digit 3), samples from the ρ0 model, held-out observations from ρ1 (90% digit 3, 10% digit 8) and samples from the ρ1 model; columns are unpaired examples, so a row need not contain a target-class image. Samples are generated with the probability flow ODE with 150 steps. Although the samples are visibly noisy, models trained at this sample size already achieve the power in Table 4 .
Figure 5: Samples from the fitted OrganAMNIST models at ntr=nte=100 per population, for (a) the denoiser and (b) the velocity representation. Rows show, from top to bottom, held-out observations from ρ0 (Left Lung), samples from the ρ0 model, held-out observations from ρ1 (90% Left Lung, 10% Right Lung) and samples from the ρ1 model; columns are unpaired examples, so a row need not contain a target-class image. Samples are generated with the probability flow ODE with 150 steps. Although the samples are visibly noisy, models trained at this sample size already achieve the power in Table 6 .
Figure 6: Samples from the fitted Fashion-MNIST models at ntr=nte=50 per population, for (a) the denoiser and (b) the velocity representation. Rows show, from top to bottom, held-out observations from ρ0 (Sneaker), samples from the ρ0 model, held-out observations from ρ1 (90% Sneaker, 10% Ankle Boot) and samples from the ρ1 model; columns are unpaired examples, so a row need not contain a target-class image. Samples are generated with the probability flow ODE with 150 steps. Although the samples are visibly noisy, models trained at this sample size already achieve the power in Table 8 .
Aˉ
AUROC
Fields
Corruption C
ρ0
ρ1 matched
ρ1 corrupted
U
J
P@ k
Mirror
Rej.
Exact
new cluster
−0.005
−0.005
0.134
0.999
0.998
0.99
1.000
1.0
shifted mode
−0.012
−0.008
0.207
0.979
0.981
0.78
0.942
1.0
deformed mode
−0.005
−0.006
0.138
0.904
0.922
0.44
0.796
0.7
Fitted
new cluster
−0.006
−0.005
0.139
0.982
1.000
0.89
0.990
0.9
shifted mode
−0.010
−0.010
0.204
0.982
0.981
0.79
0.925
0.9
Table 12: Observation-level localisation under structural corruption, ρ1=0.9ρ0+0.1C , with 500 held-out observations per group, averaged over 10 seeds. Aˉ is the mean asymmetry (denoiser, uniform weight) of observations from ρ0 and of uncorrupted and corrupted observations from ρ1 . AUROC (uniform and Jeffreys weights) and P@ k (uniform weight) measure how well A ranks the corrupted observations within the ρ1 sample; random ranking gives AUROC 0.5 and P@ k=0.10 . Mirror is the AUROC of the geometric mismatch ∣X0−ω∣ , computed on 30 corrupted and 90 uncorrupted observations per seed. Rej. is the rejection rate of the test at κ=0.05 with R=999 permutations.
Figure 7: Observation-level localisation ( ρ1=0.9ρ0+0.1C , 500 held-out observations per group.) (a)–(c) Held-out ρ1 observations coloured by the asymmetry A(ω)=−ΦUden(ω) , computed with the exact denoisers for one seed; positive values mean the ρ1 denoiser reconstructs ω better than the ρ0 denoiser. Black edges mark corrupted observations, and dashed ellipses show the 2σ region of C . For the deformed mode, uncorrupted points near the shared centre also score highly, so only the tails are localised. (d) AUROC for ranking the corrupted observations within the ρ1 sample, averaged over 10 seeds, for A and for the geometric mirror mismatch ∣X0−ω∣ , with exact (filled) and fitted (open) fields. Setup and full results: Sections 5 and 12 .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
ρ0,ρ1
Unknown data densities; H0:ρ0=ρ1
Dℓ={xℓ(i)}i=1nℓ
Dataset drawn i.i.d. from ρℓ , ℓ∈{0,1}
κ
Significance level (often denoted α in the testing literature)
xt,ρt
Stochastic interpolant and its density at time t∈[0,1]
z∼N(0,Id)
Gaussian latent variable
s=∣t−21∣
Radial coordinate, the distance from the Gaussian bottleneck
Appendix
Table 13: Summary of notation.
Figure 8: Test power on HDGM at level κ=0.05 against the per-population sample size n ( ntr=nte=n ), with all ten methods shown individually; a condensed version appears in Figure 2 . Lines show the mean over five training seeds (50 test sets each); shaded bands show ±1 standard deviation across seeds; numerical values are in Table 2 .
Figure 9: Empirical Type-I error on HDGM at level κ=0.05 , with both samples drawn from ρ0 , against the per-population sample size n ( ntr=nte=n ). Lines show the mean over five training seeds (50 test sets each); shaded bands show ±1 standard deviation across seeds; the dotted line marks κ . The PBI variants stay close to the nominal level, as guaranteed by Proposition 9 .
Figure 10: Test power on MNIST at level κ=0.05 against the per-population sample size n ( ntr=nte=n ), with all ten methods shown individually; a condensed version appears in Figure 2 . Lines show the mean over five training seeds (50 test sets each); shaded bands show ±1 standard deviation across seeds; numerical values are in Table 4 .
Figure 11: Empirical Type-I error on MNIST at level κ=0.05 , with both samples drawn from ρ0 , against the per-population sample size n ( ntr=nte=n ). Lines show the mean over five training seeds (50 test sets each); shaded bands show ±1 standard deviation across seeds; the dotted line marks κ . The PBI variants stay close to the nominal level, as guaranteed by Proposition 9 .
Figure 12: Test power on OrganAMNIST at level κ=0.05 against the per-population sample size n ( ntr=nte=n ), with all ten methods shown individually; a condensed version appears in Figure 2 . Lines show the mean over five training seeds (50 test sets each); shaded bands show ±1 standard deviation across seeds; numerical values are in Table 6 .
Figure 13: Empirical Type-I error on OrganAMNIST at level κ=0.05 , with both samples drawn from ρ0 , against the per-population sample size n ( ntr=nte=n ). Lines show the mean over five training seeds (50 test sets each); shaded bands show ±1 standard deviation across seeds; the dotted line marks κ . The PBI variants stay close to the nominal level, as guaranteed by Proposition 9 .
Figure 14: Test power on Fashion-MNIST at level κ=0.05 against the per-population sample size n ( ntr=nte=n ), with all ten methods shown individually; a condensed version appears in Figure 2 . Lines show the mean over five training seeds (50 test sets each) and shaded bands ±1 standard deviation across seeds; numerical values are in Table 8 .
Figure 15: Empirical Type-I error on Fashion-MNIST at level κ=0.05 , with both samples drawn from ρ0 , against the per-population sample size n ( ntr=nte=n ). Lines show the mean over five training seeds (50 test sets each) and shaded bands ±1 standard deviation across seeds; the dotted line marks κ . The PBI variants stay close to the nominal level, as guaranteed by Proposition 9 .
Motivated by the success of modern flow-based generative models in modeling complex data, we study two-sample testing through the lens of flow-based methods. We propose the Zero-Flow Two-Sample Test (ZF2ST), built on the zero-flow criterion, which characterizes distributional equality through a time-reversal antisymmetry of a learnable velocity field. We extend this criterion and further develop the Zero-Flow Discrepancy, an identifying discrepancy that controls the Wasserstein distance, and derive a variational representation in terms of a witness function. This representation naturally leads to a witness-based test whose power is governed by the signal-to-noise ratio (SNR), allowing direct power maximization for witness learning. ZF2ST learns the witness on one data split and performs testing on held-out samples, thereby maintaining Type-I error control and admitting a simple asymptotic null distribution. Experimentally, ZF2ST performs competitively across various synthetic and real-world benchmarks, while showing particularly strong performance in distinguishing image distributions from different sources.
Yakun Wang, Leyang Wang, Song Liu +1
University of Bristol · RIKEN AIP · University of Tokyo & RIKEN AIP
Two-sample testing is a fundamental tool for detecting distributional differences across scientific domains, but classical tests (including kernel-based tests) can be ineffective on high-dimensional structured data such as images. Recent deep two-sample tests improve sensitivity in these settings by learning informative representations, yet they provide limited insight into which data features drive rejection of the null hypothesis H0. To address this issue, we propose a counterfactual explanation framework for deep two-sample testing that generates sample-level edits moving observations from a source group toward a target group while explicitly reducing the discrepancy measured by the test. Our method combines a diffusion autoencoder with a pretrained deep two-sample test model and optimizes a maximum mean discrepancy (MMD) objective in the test model's representation space to produce plausible counterfactuals. We quantify distribution-level effects through changes in the test statistic and the resulting two-sample p-values. We evaluate the method on synthetic 2D shape datasets and two MRI cohorts. Across both settings, the counterfactual transformations consistently increase p-values relative to the original samples, indicating that the edited source set becomes statistically closer to the target distribution under the test. We measure minimality using LPIPS to ensure the counterfactuals remain close to the original samples. The resulting edits provide interpretable evidence of the features associated with the detected group differences. On MRI, the localized changes are consistent with known anatomical differences between cohorts.
Wei-Cheng Lai, Marco Simnacher, Christoph Lippert
Digital Health and Machine Learning Hasso-Plattner-Institute, University of Potsdam Potsdam, Germany · Chair of Statistics Humboldt-Universität zu Berlin Berlin, Germany · Hasso Plattner Institute for Digital Health at Mount Sinai Icahn School of Medicine at Mount Sinai New York, United States of America
Deep learning methods have proved highly effective for classification and image recognition problems. In this paper, we ask whether this success can be transferred to hypothesis testing: if a neural network can distinguish, for example, an image of a handwritten digit from another, can it also distinguish an "image of a sample" (such as a scatter plot) generated under a given statistical model from one generated outside that model? Motivated by this idea, we propose a novel procedure called deep-testing, which approaches the classical inferential problem of hypothesis testing through deep learning. More specifically, the test statistic is a classification map learned by a deep neural network from simulated data satisfying the null and alternative hypotheses, leveraging its strong discriminating power to construct a highly powerful test. As a proof of concept, we apply deep-testing to the problem of independence testing, arguably one of the most important problems in statistics. In a large-scale simulation study, deep-testing achieves the highest overall power against nineteen competing methods across a broad range of complex dependence structures, confirming the viability of the proposed approach.
Gery Geenens, Pierre Lafaye de Micheaux, Ivan Muyun Zou
School of Mathematics and statistics, UNSW Sydney, Australia