stat.MLOct 6, 2026

Anytime-valid simulation-based hypothesis testing

Authors: Patrick Forré, Lydia Brenner

Organizations: University of Amsterdam · Nikhef

Abstract

For a given data distribution (Xt)t∈N∼Q(X_t)_{t \in \mathbb{N}} \sim Q i.i.d., we investigate the hypothesis testing problem: H0:Q=P0H_0: Q = P_0 vs. H1:Q=P1H_1: Q = P_1, for two different model probability distributions P0P_0 and P1P_1. In contrast to the standard setting, where analytic densities p0p_0 and p1p_1 are given, here, we consider the density-free setting, where we only have access to i.i.d. simulations (Zt0)t∈N∼P0(Z^0_t)_{t \in \mathbb{N}} \sim P_0 and (Zt1)t∈N∼P1(Z^1_t)_{t \in \mathbb{N}} \sim P_1. For this simulation-based hypothesis testing setting, we construct an e-test martingale, resulting in a sequential test with anytime-valid type-I error guarantees, approximate growth optimality, geometrically decaying type-II error bounds, and asymptotic power one. Most ingredients used in our constructions are variants of well known concepts. The value of this paper lies in the compact presentation of an effective, anytime-valid solution for the density-free simulation-based sequential hypothesis testing case.

Explore similar work

May 27, 2026cs.CR

Optimal Rates for Differentially Private Hypothesis Testing with E-values

E-values have attracted considerable interest in recent years as flexible tools for enabling anytime-valid and adaptive data analysis. Hypothesis testing is at the core of many of these applications, which can often involve private or sensitive data. In this work, we answer a simple but important question: given two distributions P\mathbb{P} and Q\mathbb{Q}, what is the maximum achievable e-power when testing X∼PnX\sim \mathbb{P}^n against X∼QnX\sim\mathbb{Q}^n with e-values that satisfy ε\varepsilon-differential privacy? We characterize the optimal rate for this problem and provide an algorithm which matches it exactly. In the sequential setting, when observations arrive one-by-one and the analyst chooses when to halt, we give matching upper and lower bounds on the stopping times of any private e-process. Numerical experiments confirm the practicality of our algorithms, which require less data than the recently proposed DP-SPRT across a range of sequential testing problems and privacy levels.
May 22, 2026cs.DS

Entropy Equivalence Testing

We introduce the problem of \emph{entropy equivalence testing} for probability distributions, a relaxation of the well-studied closeness testing problem, where the distribution testing algorithm is now only required to distinguish, given samples from two unknown distributions p,qp,q and a parameter ε∈(0,1/2]\varepsilon \in(0,1/2], between p=qp=q and ∣H(p)−H(q)∣≥ε|H(p)-H(q)| \geq \varepsilon (where HH denotes the Shannon entropy). We provide a time- and sample-efficient algorithm for this task, showing that the optimal sample complexity for this task can be significantly lower than that of closeness testing. As an application, we leverage this result to provide the first non-trivial testing algorithm for (standard) closeness of low-degree \emph{Bayesian networks}, which significantly improves on either the sample or time complexity of a baseline based on full learning.
Oct 6, 2026stat.ML

Two-Sample Testing via Generative Processes

Deciding whether two samples come from the same distribution is a classical problem in statistics, and generative transport offers a new way to approach it. We build a stochastic interpolant directly between the two samples and observe that, for a symmetric schedule, its law is invariant under the time reflection t↦1−tt \mapsto 1-t whenever the two distributions coincide. We therefore test whether the marginals at times t and 1-t agree by computing their Jensen--Shannon divergence. Both marginals are explicit mixtures over all cross-pairs of observations, so nothing is learned, and permutation calibration gives an exact finite-sample level. For Gaussian noise, this divergence equals a time integral that pairs the reflection defects of the velocity field and of the score, so the test compares transport dynamics rather than endpoints alone. With a narrow-plus-broad noise design, the test attains the minimax separation rate n^{-2s/(4s+d)} over bounded, compactly supported densities whose difference has Sobolev smoothness s > 3d/4, with no lower bound on the densities. Fusing a dyadic grid of noise scales through their permutation ranks, without sample splitting, preserves exact level and adapts to unknown s at an iterated-logarithmic cost. Empirically, the test matches or outperforms state-of-the-art kernel two-sample tests.