The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering
Organizations: Independent Researcher
Abstract
Suppose we want a cutoff that 90% of a population falls below. We estimate it from a sample, and another sample would give a different cutoff and a different fraction below it. We ask how much that fraction varies when observations come in independent groups, such as pupils in classrooms or sentences in news articles. We prove that grouping multiplies its large-sample variance by , where is the group size, is the target fraction, and measures whether two members of a group fall on the same side of the cutoff. That correlation can differ from the correlation between the scores themselves, and it changes with the target. We give a direct proof, a counterexample to using score correlation, and an extension to unequal group sizes. A dataset therefore does not have one effective sample size. How much information it contains depends on the question you ask. In our document experiment, the same 1,000 rows carried about 217 independent observations' worth of information at the median. At the 95th percentile, they carried about 621. Nothing about the dataset changed. We asked it a different question. The number of rows is a property of the dataset. The effective sample size belongs to the analysis.
Figures & tables
| Mixture weight | Score correlation | Indicator correlation | Indicator design effect | Score-based substitute |
|---|---|---|---|---|
| 0.75 | 0.50 | 0.7222 | 1.7222 | 1.50 |
| 0.50 | 0.00 | 0.4444 | 1.4444 | 1.00 |
| 0.30 | 0.2222 | 1.2222 | 0.60 |
| Score correlation | Indicator correlation at | Effective size | Simulated sd | First-order sd | Score-based sd |
|---|---|---|---|---|---|
| 0.00 | 0.0000 | 200.0 | 0.0212 | 0.0212 | 0.0212 |
| 0.20 | 0.0798 | 161.4 | 0.0234 | 0.0236 | 0.0268 |
| 0.40 | 0.1847 | 128.7 | 0.0265 | 0.0264 | 0.0314 |
| 0.60 | 0.3221 | 101.7 | 0.0299 | 0.0297 | 0.0354 |
| 0.80 | 0.5135 | 78.7 | 0.0341 | 0.0337 | 0.0390 |
| 0.95 | 0.7545 | 61.3 | 0.0388 | 0.0382 | 0.0415 |
| Quantity | Measurement |
|---|---|
| Rows / question clusters | 25,028 / 500 |
| Cluster size: minimum / median / maximum | 8 / 44 / 135 |
| Mean / size-biased mean cluster size | 50.06 / 61.29 |
| Raw-score ICC | 0.599 |
| Indicator ICC at the reported proxy level | 0.495 |
| Plug-in design effect | 30.8 |
| Level | Indicator correlation | Design effect | Sentence draws | Document draws [95% band] | Ratio |
|---|---|---|---|---|---|
| 0.50 | 0.135 | 4.42 | 1,002 | 217 [208, 227] | 4.62× |
| 0.80 | 0.083 | 3.09 | 1,014 | 326 [312, 341] | 3.11× |
| 0.90 | 0.037 | 1.95 | 1,019 | 485 [465, 507] | 2.10× |
| 0.95 | 0.027 | 1.69 | 1,000 | 621 [595, 649] | 1.61× |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Analyte | Score ICC | ||||
|---|---|---|---|---|---|
| Sodium | 0.236 | 0.1133 | 0.0732 | 0.0732 | 0.0364 |
| Potassium | 0.104 | 0.0449 | 0.0215 | 0.0149 | 0.0074 |
| Albumin | 0.119 | 0.0307 | 0.0241 | 0.0128 | 0.0049 |
| Creatinine | 0.036 | 0.0136 | 0.0050 | 0.0036 | 0.0032 |
| Bilirubin | 0.114 | 0.0133 | 0.0042 | 0.0046 | 0.0023 |
| Mean drift | times drift | Candidate smooth coefficient, multiplied by | |
|---|---|---|---|
| 25 | |||
| 50 | |||
| 100 | |||
| 200 | |||
| 400 |
| Size profile | Mean size | Size-biased mean | Simulated sd | Predicted sd: size-biased mean | Predicted sd: plain mean |
|---|---|---|---|---|---|
| Equal, size 4 | 4.00 | 4.00 | 0.0339 | 0.0335 | 0.0335 |
| Sizes 3, 4, 5 | 3.98 | 4.15 | 0.0344 | 0.0341 | 0.0336 |
| Sizes 1, 2, 4, 9 | 4.00 | 6.38 | 0.0397 | 0.0391 | 0.0322 |
| Sizes 1, 2, 4, 8, 16 | 3.76 | 8.53 | 0.0467 | 0.0466 | 0.0330 |
| Choice of | Effective size | KS distance |
|---|---|---|
| Nominal row count | 25,028 | 0.3128 |
| Indicator-ICC plug-in | 812 | 0.0710 |
| Matched to the bootstrap variance | 1,290 | 0.0170 |
| Result | Source scripts |
|---|---|
| Gaussian dispersion comparison | verify_indicator_icc.py , sim_validation.py |
| Score-correlation counterexample and dependence checks | assumption_stress.py |
| Tail-limit calculations | verify_tail_limit.py , evt_tail_rate.py |
| Unequal-size and ICC estimation examples | ragged_and_estimation.py , icc_estimators.py |
| Released proxy measurements and resampling | prm_measurement.py , prm_dispersion.py , test_marginal_scope.py |
| SQuAD score–indicator comparison (Section 3.3) | score_squad2.py , measure.py , d01_measure.json |