stat.MLAug 21, 2026

The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

Authors: Adam Noonan

Organizations: Independent Researcher

Abstract

Suppose we want a cutoff that 90% of a population falls below. We estimate it from a sample, and another sample would give a different cutoff and a different fraction below it. We ask how much that fraction varies when observations come in independent groups, such as pupils in classrooms or sentences in news articles. We prove that grouping multiplies its large-sample variance by 1+(m−1)ρI(p)1+(m-1)ρ_I(p), where mm is the group size, pp is the target fraction, and ρI(p)ρ_I(p) measures whether two members of a group fall on the same side of the cutoff. That correlation can differ from the correlation between the scores themselves, and it changes with the target. We give a direct proof, a counterexample to using score correlation, and an extension to unequal group sizes. A dataset therefore does not have one effective sample size. How much information it contains depends on the question you ask. In our document experiment, the same 1,000 rows carried about 217 independent observations' worth of information at the median. At the 95th percentile, they carried about 621. Nothing about the dataset changed. We asked it a different question. The number of rows is a property of the dataset. The effective sample size belongs to the analysis.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Intrinsic effective sample size for manifold-valued Markov chain Monte Carlo via kernel discrepancy

    May 5, 2026Kisung YouMarkov Chain Monte CarloSample Size

  2. ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations

    Sep 29, 2026Anthony Lavertu, Jacob Cote, Sophie Gobeil +3Relational Deep LearningTraining-Inference Mismatch