Two-Sample Testing via Path-based Inference
Organizations: The University of Tokyo, Tokyo, Japan · Hasso Plattner Institute for Digital Engineering, Germany · Institute of Statistical Mathematics, Japan · The Graduate University for Advanced Studies, SOKENDAI, Japan · RIKEN Centre for Advanced Intelligence Project (AIP), Tokyo, Japan · Hasso Plattner Institute for Digital Health at the Icahn School of Medicine at Mount Sinai, USA
Abstract
Modern deep generative models are primarily studied for their ability to generate realistic samples, yet the generative dynamics they learn can also serve as objects of statistical inference. We develop this idea for two-sample testing, the problem of deciding whether the same distribution generated two finite datasets. Using stochastic interpolants, we connect both distributions to a shared Gaussian bottleneck, so that each half of the resulting path is a Gaussian channel acting on a single population. We prove that the null hypothesis holds if and only if the population denoiser, or equivalently, the velocity fields of the two halves, coincide at any single noise level, which amounts to a reflection symmetry of the path about the bottleneck. Deviations from this symmetry yield a continuum of two-sample witnesses, which we estimate via held-out regression risks on learned denoisers and velocities and aggregate along the path; under an information-theoretic weighting, the aggregated discrepancy equals the Jeffreys divergence between the noise-smoothed distributions. Calibrating the resulting statistics by permutation yields tests that are valid in finite samples for any trained networks and consistent when the fields are learned accurately. On a synthetic benchmark and three image benchmarks, the proposed tests improve power over the strongest baseline by up to 33 percentage points at an equal total sample budget, with the best choice of regression representation and path weighting depending on the data modality. These results show that generative paths provide a principled representation for statistical testing, extending stochastic-interpolant models beyond generation.
Figures & tables
| Variant | Fitted fields | Regression target | Path weight |
|---|---|---|---|
| PBI-Denoiser_U | |||
| PBI-Denoiser_J | |||
| PBI-Velocity_U | |||
| PBI-Velocity_J |
| AUROC | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Fields | Corruption | matched | corrupted | U | J | P@ | Mirror | Rej. | |
| Exact | new cluster | 0.999 | 0.998 | 0.99 | 1.000 | 1.0 | |||
| shifted mode | 0.979 | 0.981 | 0.78 | 0.942 | 1.0 | ||||
| deformed mode | 0.904 | 0.922 | 0.44 | 0.796 | 0.7 | ||||
| Fitted | new cluster | 0.982 | 1.000 | 0.89 | 0.990 | 0.9 | |||
| shifted mode | 0.982 | 0.981 | 0.79 | 0.925 | 0.9 | ||||
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| Unknown data densities; | |
| Dataset drawn i.i.d. from , | |
| Significance level (often denoted in the testing literature) | |
| Stochastic interpolant and its density at time | |
| Gaussian latent variable | |
| Radial coordinate, the distance from the Gaussian bottleneck |