Two-Sample Testing via Generative Processes
Organizations: The University of Tokyo, Tokyo, Japan · Hasso Plattner Institute for Digital Engineering, Germany · Institute of Statistical Mathematics, Japan · RIKEN Centre for Advanced Intelligence Project (AIP), Tokyo, Japan
Abstract
Deciding whether two samples come from the same distribution is a classical problem in statistics, and generative transport offers a new way to approach it. We build a stochastic interpolant directly between the two samples and observe that, for a symmetric schedule, its law is invariant under the time reflection whenever the two distributions coincide. We therefore test whether the marginals at times t and 1-t agree by computing their Jensen--Shannon divergence. Both marginals are explicit mixtures over all cross-pairs of observations, so nothing is learned, and permutation calibration gives an exact finite-sample level. For Gaussian noise, this divergence equals a time integral that pairs the reflection defects of the velocity field and of the score, so the test compares transport dynamics rather than endpoints alone. With a narrow-plus-broad noise design, the test attains the minimax separation rate n^{-2s/(4s+d)} over bounded, compactly supported densities whose difference has Sobolev smoothness s > 3d/4, with no lower bound on the densities. Fusing a dyadic grid of noise scales through their permutation ranks, without sample splitting, preserves exact level and adapts to unknown s at an iterated-logarithmic cost. Empirically, the test matches or outperforms state-of-the-art kernel two-sample tests.
Figures & tables
| Type I error (Blob-S) | Power (Blob-D) | |||||||
| Method | ||||||||
| Ours | 0.036 ∗ | 0.043 | 0.060 | 0.053 | 0.141 | 0.526 | 0.934 | 0.999 |
| MMDAgg | 0.047 | 0.041 | 0.053 | 0.057 | 0.123 | 0.405 | 0.767 | 0.935 |
| MMD-Fuse | 0.050 | 0.039 | 0.050 | 0.061 | 0.149 | 0.403 | 0.763 | 0.931 |
| C2ST-RF ‡ | 0.038 | 0.038 | 0.038 | 0.042 | 0.058 | 0.181 | 0.343 | 0.559 |
| AutoTST ‡ | 0.041 | 0.058 | 0.056 | 0.059 | 0.054 | 0.136 | 0.315 | 0.492 |
| Type I error | Power | |||||
| Method | Null | |||||
| Ours | 0.054 | 1.000 | 0.993 | 0.769 | 0.411 | 0.122 |
| MMDAgg | 0.050 | 1.000 | 0.956 | 0.585 | 0.252 | 0.087 |
| MMD-Fuse | 0.056 | 1.000 | 0.930 | 0.622 | 0.309 | 0.104 |
| C2ST-RF ‡ | 0.034 ∗ | 0.971 | 0.788 | 0.433 | 0.238 | 0.089 |
| AutoTST ‡ | 0.035 ∗ | 0.989 | 0.871 | 0.530 | 0.256 | 0.096 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Selection | Testing | Additional | Total |
|---|---|---|---|---|
| Our test, MMDAgg, MMD-FUSE | ||||
| MMD, median | ||||
| MMD, split | ||||
| C2ST, AutoTST | ||||
| MMD, oracle |
| Alternative | Retained digits |
|---|---|
| Schedule | ||
|---|---|---|
| Linear | ||
| Endpoint power, | ||
| Power ratio, | ||
| Quintic smoothstep | ||
| Normalized sigmoid, | ||
| normalized |