Statistical Power Analysis

Latest papers 9

Oct 5, 2026stat.ML

Two-Sample Testing for Random Graphs without Vertex Correspondence

Two populations of graphs often have to be compared without any correspondence between their vertices, for instance when networks come from different communities, or when a graph generative model is evaluated against held-out graphs. We study how many graphs such an unaligned two-sample test needs, and which graph statistics can detect which differences. For an Erdős--Rényi null and a planted two-block difference that leaves every expected degree unchanged, we show that m≍t−3m\asymp t^{-3} graphs per group are necessary and sufficient when the per-graph signal-to-noise ratio is t<1t<1. Signed triangle counts attain this rate, and the lower bound holds for every graph size. With aligned vertices m≍t−1m\asymp t^{-1} graphs suffice, so misalignment costs a factor of order t−2t^{-2}. When the triangle signal cancels, the rate becomes t−4t^{-4} and 44-cycles are needed. Statistics built from trees have exactly the same expectation under both hypotheses, and tests based on finitely many of them have asymptotically no power. In the graphon limit, this class includes degree distributions and message-passing graph neural network features. For a non-constant null, a generic difference is visible at first order, and a simple motif test attains the aligned order of sample size, suggesting that misalignment is costly mainly for differences that are invisible at low orders. We also give an exactly valid test for one or two graphs per group, at a cost in power. In our simulations, the fitted exponents are close to the predicted ones, and degree-based and random-GNN evaluation metrics stay at their level in a setting where signed triangles need about 6565 graphs.
Sep 23, 2026stat.ME

Beyond the Illusion of Power: Calibrating Quasi-Experiments in Observational IS

Information systems (IS) researchers increasingly use quasi-experimental methods such as difference-in-differences (DiD) and instrumental variables (IV) to recover causal effects from observational panel data. Power calculations that justify these designs assume i.i.d. errors, but the deeper problem is what even a cluster-robust calculator cannot see. We report a Monte Carlo study over 9837 parameter conditions (approx 9.8 million datasets) and decompose the planned-versus-achieved power gap. The serial-correlation component is recoverable by an AR(1)-aware calculator when rho is known, and partially when rho must be estimated from short pre-periods, but panel attrition, staggered-adoption bias, and parallel-trends pretesting are captured by no closed-form formula; exogenous attrition alone costs approx 8 to 11 percentage points at the few-hundred-to-thousand sample sizes IS studies use. Treatment-correlated, outcome-dependent attrition instead induces bias, not just power loss. For IV, holding first-stage F fixed, larger N neither raises power nor curbs exclusion bias, though with a fixed instrument more data does sharpen the first stage, so identification rests on instrument strength, not sample size.
Sep 10, 2026cs.LG

A Statistical Approach to Estimating Sample Size of Machine Learning Models

Sample size determination for machine learning (ML) prediction models is challenging because conventional power analysis typically requires the predictor-outcome relationship and effect structure to be specified a priori. Nonlinear ML models learn complex prediction surfaces that do not admit straightforward analytical power calculations. We propose a framework that approximates nonlinear ML models with localized linear representations and estimates sample size requirements by evaluating statistical power across these local regions.
Aug 8, 2026cs.AI

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark in which an unknown fraction alpha of items was seen during training. With matched clean and seen controls, the behavioral channel is the sparse mixture Q_alpha = (1 - alpha) P_0 + alpha P_1, and an exact second-moment argument shows that detectability is governed by alpha * rho * sqrt(m), where rho^2 = chi^2(P_1 || P_0) measures behavioral separability. Any scalar detector reduces to its efficacy, ef = |E_1 f - E_0 f| / sqrt(Var_0(f)) <= rho, which can be estimated from controls before the audit is run. A separate sample-split certificate lower-bounds alpha distribution-free, without requiring an orientation assumption. Our empirical finding is two-sided. Frozen calibration efficacy predicts held-out power curves, with R^2 = 0.83-0.98 across six exact-permutation channels, but the efficacy-only Gaussian budget is miscalibrated at the small sample sizes it prescribes, failing in 9/9 gate-passing channels even though efficacy itself transports. The failure is in the inversion, not the calibration. A predeclared two-stage planner that simulates the deployed test repairs the budgets, is uniformly conservative, and abstains when its probe does not transport. The certificate is valid but vacuous at audit scale, and a five-seed paired injection study recovers the mechanism ordering verbatim > paraphrase > surface, in which the apparent answer-only signal is explained by baseline drift. We report the audit contract and its failures together: a non-rejection is interpretable only alongside the efficacy, budget, and validity gates that produced it.
Jul 6, 2026stat.ML

Beyond Heuristic Tuning: Power-Calibrated LLM Watermarking

Logit-based watermarking is a widely used mechanism for identifying LLM generated content, yet its effectiveness is governed by a fundamental trade-off between detectability and semantic distortion. Existing analyses provide limited guidance for principled hyperparameter selection, leaving practical deployments reliant on heuristic tuning. In this work, we develop a power-calibrated statistical framework that establishes explicit quantitative relationships between watermark hyperparameters, detection power, and distortion. This characterization transforms watermark design into a guided optimization problem. Building on these results, we derive practical parameter selection procedures that achieve optimal tradeoffs under constraints. Extensive experiments across multiple language models and datasets validate the theory and demonstrate that the proposed framework consistently identifies Pareto-optimal points.
Jul 5, 2026cs.CL

evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations

The dominant practice in language model evaluation is to report a single accuracy number per model and declare the higher one better, without testing whether the gap could plausibly be sampling noise. On benchmarks of a few thousand items, and under temperature sampling where a model can differ from itself run to run by more than the reported gap between models, this practice routinely overstates confidence in headline claims. The statistical machinery to fix this -- confidence intervals, paired significance tests, power analysis, clustered standard errors, multiple-comparison correction -- is well established, but no standard, pip-installable tool packages it in the shape an evaluation actually takes: a per-item results table. We present evalci, a pure-Python library (numpy/scipy/pandas only) that turns a per-item results table into a publication-ready claim -- e.g., "Model A beats Model B, Δ=3.1Δ=3.1 pts, 95% CI [1.2, 5.0], paired permutation p=0.002p=0.002, n=1,319n=1{,}319" -- in one function call, with adapters for lm-evaluation-harness and HELM output. Every routine is validated against an independent reference (statsmodels, or brute-force exact enumeration) rather than only against itself. As a case study, we re-analyze a public comparison of nine language models' MMLU accuracy and find that 3 of the 8 adjacent leaderboard-rank gaps are not statistically significant after correcting for the 36 pairwise comparisons the ranking implies. evalci is available at https://pypi.org/project/evalci/ (source: https://github.com/Shreyaskc/evalci, DOI: https://doi.org/10.5281/zenodo.21201815)
May 28, 2026cs.CL

Resolution Diagnostics for Paired LLM Evaluation

Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired evaluation design: 11 of 40 Open LLM Leaderboard v1 pairwise comparisons and 4 of 9 MMLU-Pro top-10 adjacent-rank pairs are unresolved at (alpha, 1-beta) = (0.05, 0.8). The MMLU-Pro count rises to 6/9 under real subject-level clustering and stays at 5-6 out of 9 in 99.9% of category-bootstrap resamples. We frame paired LLM evaluation as a hypothesis-testing problem, invert level-alpha, power-(1-beta) tests, and report a per-pair resolution ratio q = N/N* as the primary diagnostic. A sharp small-effect expansion with an explicit second-order constant shows that the widely-used unpaired Cohen-h-plus-(1-rho) shortcut deviates from the correct N* by approximately a factor of two in the close-comparison regime, a deficit that three of five off-the-shelf calculators(Cohen 1988, G*Power, R pwr) silently inherit when the user post-multiplies their per-arm output by (1-rho). The unresolved-pair pattern remains under multiplicity correction and anytime-valid sequential testing.
May 25, 2026cs.LG

Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

This is a planning-method note with an unpaired pilot audit. We adapt the classical paired-binary sample-size calculation (Miettinen, 1968) to quantization benchmarks, giving a conservative minimum detectable effect (MDE) bound δ∗≤(z1−α/2+z1−β)ρd/mδ^{*} \le (z_{1-α/2}+z_{1-β})\sqrt{ρ_d/m} in the paired item count mm and the FP16-NF4 disagreement rate ρdρ_d. The bound turns "how reliable is my quantization claim?" into a one-line budget a benchmark designer can commit to before running. We illustrate the bound on four models and four benchmarks (k=5k=5 splits of n=100n=100), and add a parallel MMLU prompt-template study to put the bound's quantization-noise scale alongside the prompt-noise scale. Assuming ρd=0.10ρ_d=0.10 (an unmeasured planning value), all observed NF4-FP16 deltas fall below the implied MDE, and most cross-split SDs lie within ±1.5\pm 1.5 pp of the binomial reference p(1−p)/n\sqrt{p(1-p)/n}, so much of the variance reported as "benchmark unreliability" on n=100n=100 subsamples is binomial sampling noise. The single borderline cell (OPT-WinoGrande, ∣Δ∣=3.2|Δ|=3.2 pp) is below the implied MDE at ρd=0.10ρ_d=0.10 but above it at ρd=0.05ρ_d=0.05, illustrating the planning trade-off the bound makes explicit. On MMLU, prompt-template ranges of 2-10 pp meet or exceed the largest observed quantization delta (3.2 pp), so a quantization audit that does not first fix the prompt template absorbs template variance into its noise floor. We complement the bound with a five-line pre-registration template.
Apr 25, 2026cs.LG

Unstable Rankings in Bayesian Deep Learning Evaluation

Standard evaluations of Bayesian deep learning methods assume that metric estimates are reliable, but we show this assumption fails under data scarcity. Method rankings are not only unreliable at small nn, but also dataset-dependent in ways that point estimates cannot reveal: the same method comparison yields P(MCD≺Ensemble)=1.000P(\mathrm{MCD} \prec \mathrm{Ensemble}) = 1.000 at n=50n = 50 on one dataset and remains below 0.950.95 even at n=500n = 500 on another. Across the datasets we consider, no universal sample size threshold exists, which is precisely why dataset-specific posterior inference is necessary. To address this, we use a Bayesian hierarchical model with method-specific variances to treat evaluation metrics as random variables across data realizations, and we use a predictive Minimum Detectable Difference curve to assess whether an observed gap would be detectable at a given training size. Across six Bayesian deep learning methods and five regression datasets, our results show that uncertainty-aware evaluation is necessary in low-data settings, because current evidence for method superiority and predictive detectability at the same training size can diverge substantially. Our framework provides practitioners with principled tools to determine whether their evaluation data is sufficient before drawing conclusions about method superiority.