What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects
Organizations: Department of Biology, Stanford University, Stanford, CA, USA · Institute of Health Informatics, University College London, London, UK
Abstract
Harness search keeps a change to the prompts, reasoning switches, token budgets or parsers around a frozen model if the change raises a score. Such a gain can come from answers the parser could not read before, a weak comparison, or a defect in the evaluation. We preregistered a study of where these gains come from, with three small models, three benchmarks, replication and test partitions, a GEPA search arm and six evaluation defects injected one at a time, and we report all 47 primary endpoints. Turning thinking off raised accuracy over a capped thinking setting in 5 of 9 model-benchmark cells, and in each the gain came mostly from questions where the capped setting gave no readable answer. The thinking-off setting was not meaningfully worse than a rescue configuration or four GEPA-selected harnesses in 11 of 13 comparisons, and lost to the rescue on GSM8K for two models. GEPA repaired its broken starting points, but none of its selected harnesses was more accurate than the thinking-off setting. A thinking budget in the serving engine, which also allows a longer answer, lowered truncation and raised the parse rate in 6 of 9 cells. In 6 of 15 evaluable defect-model pairs, replication through the same pipeline reproduced the defect's distortion instead of revealing it. On the LongevityBench multiple-choice tasks, only the longevity-tuned model beat the strongest constant-label baseline.
Figures & tables
| Configuration | Request |
|---|---|
| THINK-64 , THINK-512 , THINK-1536 | thinking on; 64, 512 or 1,536 tokens for thinking and answer together ( THINK-64 is the case study’s parent) |
| AUTHORS-1500 | thinking off; 1,500 tokens (the benchmark authors’ serving protocol) |
| BUDGET-512-1500 | thinking on; vLLM ends thinking after 512 tokens ( Muennighoff et al., 2025 ) ; up to 1,500 tokens of answer follow, so 2,012 tokens in all against 512 for THINK-512 |
| RESCUE-1500 | THINK-512 ; where its call ended at the token limit, the AUTHORS-1500 call on the same row (the call that gives the answer decides the properties of the row) |
| GEPA harnesses | four runs (clean or broken start, two seeds), 6,000 metric calls on 300 search rows, Intern-S1-mini as proposer; the selected harness edits the system prompt and appends an instruction, so the arm searches prompt text only (Appendix J ) |
| Claim and estimand | Verdict | Endpoints (family) | |
|---|---|---|---|
| H1 | Most of the accuracy that AUTHORS-1500 gains over THINK-512 comes from rows THINK-512 leaves unparsed. Gate: ; statistic: . | no gain: not tested unless the lower bound of is above zero; if it is, visibility dominates if the lower bound of is above zero, quality matters if the upper bound of is below zero, inconclusive otherwise. | 9: one per cell |
| H2 | AUTHORS-1500 is not meaningfully worse than a comparator: . | holds if the lower bound of ; fails if the upper bound ; else inconclusive . | 13: RESCUE-1500 in nine cells; four GEPA harnesses ( Qwen3-1.7B-Longevity , letters) |
| H3 | Against THINK-512 , BUDGET-512-1500 (a forced end of thinking and a larger allowance) lowers truncation ( ) and raises parsing ( ). | holds if both lower bounds ; fails if either upper bound ; no room if the upper bound of THINK-512 ’s truncation is (applied first); else inconclusive . | 9: one per cell |
| H4 | For each defect and model, prospective replication confirms the defect’s distortion (Section 3.4 ). | holds , fails or inconclusive from three intervals; not evaluable if the defect cannot move the start by more than . | 16 pairs, fixed |
| Model | Pool | H1 AUTHORS-1500 vs THINK-512 | H2 AUTHORS-1500 vs RESCUE-1500 | H3 BUDGET-512-1500 vs THINK-512 |
|---|---|---|---|---|
| Qwen3-1.7B-Longevity | LB letters | no gain: not tested | holds | no room |
| MMLU-Pro | no gain: not tested | holds | inconclusive | |
| GSM8K | no gain: not tested | holds | no room | |
| SmolLM3-3B | LB letters | visibility dominates | holds | holds |
| MMLU-Pro | visibility dominates | holds | holds | |
| GSM8K | visibility dominates | fails † | holds |
| Defect | Model | Bound | Claim | |||
| A reference answer as assistant turn | Qwen3 | 1.000 | inconclusive | |||
| SmolLM3 | 1.000 | inconclusive | ||||
| Intern-S1 | 1.000 | inconclusive | ||||
| B1 gold fallback | Qwen3 | 0.205 | holds | |||
| SmolLM3 | 0.083 | fails | ||||
| Intern-S1 | 0.057 | fails |
| Check | Reads | Detects | Uninformative | Misses | Fires on the clean instances |
|---|---|---|---|---|---|
| PR | gain and level rules of prospective replication | 5 | 1 | 10 | PR gain 1/7 |
| TR | per-call prompt-token reconciliation | 3 | 0 | 13 | none |
| TB | trivial baselines | 6 | 3 | 7 | TB2 3/7 |
| SCF | the simple configuration family | 3 | 12 | 1 | SCF1 4/7, SCF2 4/7, SCF3 2/7 |
| IR | reading of the rendered inputs | 0 | 16 | 0 | IR4 7/7 |
| SMI | served-model identity (defect E) | 1 | 0 | 0 | – |
| Model | THINK-512 | THINK-1536 | AUTHORS-1500 | BUDGET-512-1500 | RESCUE-1500 |
| LongevityBench letter pool (3,942 rows). Chance 0.380; constant letter 0.383 (chosen on replication), 0.389 (oracle); constant label 0.530 (chosen on replication), 0.539 (oracle). Strongest, the oracle constant label: 0.539 [0.527, 0.551]. | |||||
| Qwen3 | 0.635 ( ) | 0.635 ( ) | 0.632 ( ) | 0.635 ( ) | 0.635 ( ) |
| SmolLM3 | 0.070 ( ) | 0.157 ( ) | 0.129 ( ) | 0.121 ( ) | 0.128 ( ) |
| Intern-S1 | 0.000 ( ) | 0.125 ( ) | 0.314 ( ) | 0.253 ( ) | 0.314 ( ) |
| LB-0030 (1,577 rows). Chance 0.200; constant letter 0.216 (oracle); constant label 0.250 (oracle). Strongest, the oracle constant label: 0.250 [0.230, 0.269]. | |||||
| Qwen3 | 0.369 ( ) | 0.369 ( ) | 0.344 ( ) | 0.369 ( ) | 0.369 ( ) |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Lesson | Tested by |
|---|---|
| The parent parsed none of the prospective rows; the rescue made answers visible | H1, H3 |
| A thinking-off setting that the search never tried was non-inferior to the rescue | H2 |
| The caller sent the reference answer | H4, defect A |
| LB-0042’s gold came from a fallback rule; the scorer could fail open | H4, defects B1, B2 |
| The search’s split put a giant component into one partition | H4, defect C |
| The parent was a broken start | H4, defect D |
| Cell | / / | Verdict | Repl. | |||||
|---|---|---|---|---|---|---|---|---|
| Qwen3 LB letters | [ , ] | [ , ] | 1 / 3,941 / 0 | no gain: not tested | R | |||
| Qwen3 MMLU-Pro | [ , ] | [ , ] | 48 / 751 / 0 | no gain: not tested | R | |||
| Qwen3 GSM8K | [ , ] | [ , ] | 0 / 523 / 0 | no gain: not tested | R | |||
| SmolLM3 LB letters | [ , ] | [ , ] | 749 / 448 / 100 | visibility dominates | R | |||
| SmolLM3 MMLU-Pro | [ , ] | [ , ] | 529 / 81 / 13 | visibility dominates | R | |||
| SmolLM3 GSM8K | [ , ] | [ , ] | 141 / 340 / 29 | visibility dominates | R |
| Cell | / / | Verdict | Repl. | |||||
|---|---|---|---|---|---|---|---|---|
| Qwen3 LB letters | [ , ] | [ , ] | 3,942 / 0 / 0 | visibility dominates | ||||
| Qwen3 MMLU-Pro | [ , ] | [ , ] | 274 / 525 / 0 | visibility dominates | ||||
| Qwen3 GSM8K | [ , ] | [ , ] | 1 / 522 / 0 | no gain: not tested | ||||
| SmolLM3 LB letters | [ , ] | [ , ] | 1,197 / 0 / 0 | visibility dominates | ||||
| SmolLM3 MMLU-Pro | [ , ] | [ , ] | 610 / 0 / 0 | visibility dominates | ||||
| SmolLM3 GSM8K | [ , ] | [ , ] | 481 / 0 / 0 | visibility dominates |
| Endpoint | (adjusted interval) | Verdict | Replication verdict | Label |
|---|---|---|---|---|
| Qwen3 LB letters, RESCUE-1500 | [ , ] | holds | holds | replicated |
| Qwen3 MMLU-Pro, RESCUE-1500 | [ , ] | holds | holds | replicated |
| Qwen3 GSM8K, RESCUE-1500 | [ , ] | holds | holds | replicated |
| SmolLM3 LB letters, RESCUE-1500 | [ , ] | holds | holds | replicated |
| SmolLM3 MMLU-Pro, RESCUE-1500 | [ , ] | holds | holds | replicated |
| SmolLM3 GSM8K, RESCUE-1500 | [ , ] | fails | inconclusive | inconclusive replication |
| Cell | vs THINK-512 | vs THINK-1536 | vs BUDGET-512-1500 |
|---|---|---|---|
| Qwen3 LB letters | [ , ] holds | [ , ] holds | [ , ] holds |
| Qwen3 MMLU-Pro | [ , ] holds | [ , ] holds | [ , ] holds |
| Qwen3 GSM8K | [ , ] holds | [ , ] holds | [ , ] holds |
| SmolLM3 LB letters | [ , ] holds | [ , ] holds | [ , ] holds |
| SmolLM3 MMLU-Pro | [ , ] holds | [ , ] holds | [ , ] holds |
| SmolLM3 GSM8K | [ , ] holds | [ , ] fails | [ , ] inconclusive |
| Cell | THINK-512 truncation | Verdict | Replication verdict | Label | ||
|---|---|---|---|---|---|---|
| Qwen3 LB letters | [ , ] | [ , ] | [ , ] | no room | no room | replicated |
| Qwen3 MMLU-Pro | [ , ] | [ , ] | [ , ] | inconclusive | inconclusive | replicated |
| Qwen3 GSM8K | [ , ] | [ , ] | [ , ] | no room | no room | replicated |
| SmolLM3 LB letters | [ , ] | [ , ] | [ , ] | holds | holds | replicated |
| SmolLM3 MMLU-Pro | [ , ] | [ , ] | [ , ] | holds | holds | replicated |
| SmolLM3 GSM8K | [ , ] | [ , ] | [ , ] | holds | holds | replicated |
| Pair | Bound | Stability of | Claim | |||
| A Qwen3 | 1.000 | [ , ] (open) | [ , ] (open) | [ , ] (open) | inconclusive | inconclusive |
| A SmolLM3 | 1.000 | [ , ] (open) | [ , ] (open) | [ , ] (open) | inconclusive | inconclusive |
| A Intern-S1 | 1.000 | [ , ] (holds) | [ , ] (holds) | [ , ] (open) | inconclusive | inconclusive |
| B1 Qwen3 | 0.205 | [ , ] (holds) | [ , ] (holds) | [ , ] (holds) | stable | holds |
| B1 SmolLM3 | 0.083 | [ , ] (fails) | [ , ] (fails) | [ , ] (holds) | stable | fails |
| B1 Intern-S1 | 0.057 | [ , ] (fails) | [ , ] (fails) | [ , ] (holds) | stable | fails |
| Pair | PR | TR | TB | SCF | IR | SMI | Rules fired (expected ones in bold) |
|---|---|---|---|---|---|---|---|
| A Qwen3 | U | D | U | U | U | – | IR1 , IR2 , IR4, PR.gain, SCF2, TB2, TB3 , TR |
| A SmolLM3 | D | D | U | U | U | – | IR1 , IR2 , IR4, PR.gain, SCF1, SCF2, SCF3, TB2, TR ; expected, not fired: TB3 |
| A Intern-S1 | M | D | U | U | U | – | IR1 , IR2 , IR4, SCF1, SCF3, TB2, TB3 , TR |
| B1 Qwen3 | M | M | M | M | U | – | IR3 , IR4 |
| B1 SmolLM3 | M | M | M | U | U | – | IR3 , IR4, SCF1 |
| B1 Intern-S1 | M | M | M | U | U | – | IR3 , IR4, SCF1, SCF2, SCF3 |
| Check | Rule | Fired on | Of |
|---|---|---|---|
| PR | PR.gain | 1: clean/A / Qwen3 | 7 |
| PR | PR.level | 0: none | 7 |
| TR | TR | 0: none | 7 |
| TB | TB1 | 0: none | 7 |
| TB | TB2 | 3: clean/A / Intern-S1, clean/A / Qwen3, clean/A / SmolLM3 | 7 |
| TB | TB3 | 0: none | 7 |
| Cell | Configuration | Accuracy | Parse | Truncated | Non-committed |
|---|---|---|---|---|---|
| Qwen3 LB letters | THINK-64 | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] |
| THINK-512 | 0.635 [0.621, 0.649] | 1.000 [0.999, 1.000] | 0.000 [0.000, 0.001] | 0.000 [0.000, 0.000] | |
| THINK-1536 | 0.635 [0.621, 0.649] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] | |
| AUTHORS-1500 | 0.632 [0.618, 0.646] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] | |
| BUDGET-512-1500 | 0.635 [0.621, 0.649] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] | |
| RESCUE-1500 | 0.635 [0.621, 0.649] | 1.000 [1.000, 1.000] | 0.000 [0.000, 0.000] | 0.000 [0.000, 0.000] |
| Model | THINK-64 | THINK-512 | THINK-1536 | AUTHORS-1500 | BUDGET-512-1500 | RESCUE-1500 |
|---|---|---|---|---|---|---|
| LB-0030 (1,577 rows). Chance 0.200; constant letter 0.216 (non-oracle), 0.216 (oracle); constant label 0.242 (non-oracle), 0.250 (oracle). Strongest, oracle: 0.250 [0.230, 0.269]. | ||||||
| Qwen3 | 0.000 ( ) | 0.369 ( ) | 0.369 ( ) | 0.344 ( ) | 0.369 ( ) | 0.369 ( ) |
| SmolLM3 | 0.000 ( ) | 0.000 ( ) | 0.021 ( ) | 0.053 ( ) | 0.032 ( ) | 0.053 ( ) |
| Intern-S1 | 0.000 ( ) | 0.000 ( ) | 0.004 ( ) | 0.023 ( ) | 0.055 ( ) | 0.023 ( ) |
| Qwen3, GEPA | CLEAN-s0 0.344 ( ); BROKEN-s0 0.353 ( ); CLEAN-s1 0.388 ( ); BROKEN-s1 0.321 ( ) | |||||
| LB-0034 (788 rows). Chance 0.500; constant letter 0.485 (non-oracle), 0.515 (oracle); constant label 0.485 (non-oracle), 0.515 (oracle). Strongest, oracle: 0.515 [0.480, 0.551]. | ||||||
| Model | THINK-64 | THINK-512 | THINK-1536 | AUTHORS-1500 | BUDGET-512-1500 | RESCUE-1500 |
|---|---|---|---|---|---|---|
| LB-0030 (796 rows). Chance 0.200; constant letter 0.212 (non-oracle), 0.212 (oracle); constant label 0.255 (non-oracle), 0.270 (oracle). Strongest, oracle: 0.270 [0.242, 0.299]. | ||||||
| Qwen3 | 0.000 ( ) | 0.358 ( ) | 0.358 ( ) | 0.335 ( ) | 0.358 ( ) | 0.358 ( ) |
| SmolLM3 | 0.000 ( ) | 0.000 ( ) | 0.025 ( ) | 0.041 ( ) | 0.031 ( ) | 0.041 ( ) |
| Intern-S1 | 0.000 ( ) | 0.000 ( ) | 0.005 ( ) | 0.011 ( ) | 0.046 ( ) | 0.011 ( ) |
| Qwen3, GEPA | CLEAN-s0 0.335 ( ); BROKEN-s0 0.345 ( ); CLEAN-s1 0.389 ( ); BROKEN-s1 0.307 ( ) | |||||
| LB-0034 (398 rows). Chance 0.500; constant letter 0.490 (non-oracle), 0.510 (oracle); constant label 0.490 (non-oracle), 0.510 (oracle). Strongest, oracle: 0.510 [0.462, 0.558]. | ||||||
| Model | THINK-64 | THINK-512 | THINK-1536 | AUTHORS-1500 | BUDGET-512-1500 | RESCUE-1500 |
|---|---|---|---|---|---|---|
| MMLU-Pro (799 items). Chance 0.110; constant letter B, chosen on the other partition and oracle alike, 0.119 [0.096, 0.141]. | ||||||
| Qwen3 | 0.140 ( ) | 0.217 ( ) | 0.234 ( ) | 0.229 ( ) | 0.232 ( ) | 0.232 ( ) |
| SmolLM3 | 0.000 ( ) | 0.098 ( ) | 0.314 ( ) | 0.353 ( ) | 0.258 ( ) | 0.378 ( ) |
| Intern-S1 | 0.000 ( ) | 0.074 ( ) | 0.279 ( ) | 0.477 ( ) | 0.542 ( ) | 0.493 ( ) |
| GSM8K (527 items). Constant answer 10 (chosen on the other partition) 0.019 [0.008, 0.032], 2 (oracle) 0.034 [0.019, 0.051]; no chance level. | ||||||
| Qwen3 | 0.068 ( ) | 0.070 ( ) | 0.070 ( ) | 0.070 ( ) | 0.070 ( ) | 0.070 ( ) |
| Model | THINK-64 | THINK-512 | THINK-1536 | AUTHORS-1500 | BUDGET-512-1500 | RESCUE-1500 |
|---|---|---|---|---|---|---|
| MMLU-Pro (399 items). Chance 0.111; constant letter B, chosen on the other partition and oracle alike, 0.133 [0.100, 0.168]. | ||||||
| Qwen3 | 0.108 ( ) | 0.180 ( ) | 0.193 ( ) | 0.180 ( ) | 0.190 ( ) | 0.185 ( ) |
| SmolLM3 | 0.000 ( ) | 0.090 ( ) | 0.333 ( ) | 0.378 ( ) | 0.291 ( ) | 0.398 ( ) |
| Intern-S1 | 0.000 ( ) | 0.105 ( ) | 0.293 ( ) | 0.446 ( ) | 0.566 ( ) | 0.454 ( ) |
| GSM8K (264 items). Constant answer 2 (chosen on the other partition) 0.030 [0.011, 0.053], 10 (oracle) 0.045 [0.023, 0.072]; no chance level. | ||||||
| Qwen3 | 0.057 ( ) | 0.057 ( ) | 0.057 ( ) | 0.057 ( ) | 0.057 ( ) | 0.057 ( ) |
| Model | Pool | Configuration | Parse (seeded) | Truncated (seeded) | accuracy | parse | truncation |
|---|---|---|---|---|---|---|---|
| Qwen3 | LB letters | THINK-64 | 0.000 | 1.000 | [ , ] | [ , ] | [ , ] |
| THINK-512 | 0.999 | 0.001 | [ , ] | [ , ] | [ , ] | ||
| THINK-1536 | 1.000 | 0.000 | [ , ] | [ , ] | [ , ] | ||
| BUDGET-512-1500 | 1.000 | 0.000 | [ , ] | [ , ] | [ , ] | ||
| MMLU-Pro | THINK-64 | 0.641 | 0.359 | [ , ] | [ , ] | [ , ] | |
| THINK-512 | 0.932 | 0.068 | [ , ] | [ , ] | [ , ] |
| Partition | Model | Pool | (seeded) | (seeded) | |||
|---|---|---|---|---|---|---|---|
| Test | Qwen3 | LB letters | 0.000 | 0.001 | [ , ] | [ , ] | [ , ] |
| MMLU-Pro | 0.060 | 0.068 | [ , ] | [ , ] | [ , ] | ||
| GSM8K | 0.000 | 0.000 | [ , ] | [ , ] | [ , ] | ||
| Intern-S1 | LB letters | 1.000 | 0.997 | [ , ] | [ , ] | [ , ] | |
| MMLU-Pro | 0.915 | 0.902 | [ , ] | [ , ] | [ , ] | ||
| GSM8K | 0.522 | 0.619 | [ , ] | [ , ] | [ , ] |
| Configuration | Tied | Untied | Tied untied | Parse t / u | Same answers | Only t / u right | |
|---|---|---|---|---|---|---|---|
| CFG-THINK-64-ORIG | 0.000 | 0.000 | [ , ] | 0.000 / 0.000 | 1.000 | 0 / 0 | 1.000 |
| B1 | 0.452 | 0.449 | [ , ] | 1.000 / 1.000 | 0.741 | 182 / 157 | 0.192 |
| CFG-THINKING-FALSE-64-ORIG | 0.460 | 0.460 | [ , ] | 1.000 / 1.000 | 0.731 | 194 / 194 | 1.000 |
| CFG-THINK-1536-ORIG | 0.469 | 0.463 | [ , ] | 1.000 / 1.000 | 0.656 | 384 / 333 | 0.062 |
| CFG-NOTHINK-SUFFIX-64-ORIG | 0.452 | 0.449 | [ , ] | 1.000 / 1.000 | 0.741 | 182 / 157 | 0.192 |
| CFG-THINK-64-ANSWERONLY | 0.479 | 0.470 | [ , ] | 1.000 / 1.000 | 0.731 | 260 / 185 | 0.000 |
| Harness | Proposals | Candidates seen | Refused | Kept | Metric calls | Selected / kept | Characters (system / instruction) | Valset | Replication acc. | Test acc. | Test parse |
|---|---|---|---|---|---|---|---|---|---|---|---|
| GEPA-CLEAN-s0 | 207 | 204 | 28 | 31 | 6,039 | 0 / 31 | 103 / 0 | 0.680 | 0.633 | 0.632 [0.618, 0.646] | 1.000 |
| GEPA-BROKEN-s0 | 377 | 377 | 24 | 24 | 6,147 | 1 / 24 | 103 / 1,146 | 0.660 | 0.635 | 0.629 [0.615, 0.643] | 1.000 |
| GEPA-CLEAN-s1 | 195 | 192 | 36 | 31 | 6,003 | 12 / 31 | 2,469 / 922 | 0.700 | 0.641 | 0.631 [0.616, 0.645] | 1.000 |
| GEPA-BROKEN-s1 | 243 | 238 | 67 | 29 | 6,003 | 1 / 29 | 103 / 2,797 | 0.667 | 0.619 | 0.620 [0.606, 0.634] | 1.000 |
| Model | Measure | THINK-64 | THINK-512 | THINK-1536 | AUTHORS-1500 | BUDGET-512-1500 | RESCUE-1500 |
|---|---|---|---|---|---|---|---|
| Qwen3 | parse rate | 0.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| mean abs. error (years) | – | 15.69 | 15.69 | 16.50 | 15.69 | 15.69 | |
| exact match | 0.000 | 0.027 | 0.027 | 0.030 | 0.027 | 0.027 | |
| SmolLM3 | parse rate | 0.000 | 0.000 | 0.004 | 0.001 | 0.003 | 0.001 |
| mean abs. error (years) | – | – | 24.71 | 57.70 | 35.08 | 57.70 | |
| exact match | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
| Analysis | Endpoints | Changed | Which (primary verdict here) |
|---|---|---|---|
| Strict scorer | 31 | 2 | H2 SmolLM3 MMLU-Pro, RESCUE: holds inconclusive; H1 SmolLM3 GSM8K: visibility dominates no gain: not tested |
| Never-called rows only | 13 | 3 | H2 Qwen3 letters, RESCUE: holds inconclusive; H2 Qwen3 letters, GEPA-BROKEN-s1: holds inconclusive; H1 SmolLM3 letters: visibility dominates inconclusive |
| Flagged rows included (H1–H3) | 31 | 0 | none |
| Case study’s pool (LB-0030 and LB-0034) | 13 | 1 | H2 Qwen3 letters, GEPA-CLEAN-s1: holds inconclusive |
| LB-0030 alone | 13 | 2 | H2 Qwen3 letters, RESCUE: holds inconclusive; H2 Qwen3 letters, GEPA-CLEAN-s1: holds inconclusive |
| LB-0034 alone | 13 | 0 | none |
| Rows | Partition | Rows | Letter rows | LB-0030 mean age | LB-0042 share deceased at follow-up |
|---|---|---|---|---|---|
| Never called | replication | 187 | 103 | 36.2 | 0.022 |
| Analysis rows | replication | 1,990 (pool) | – | 50.8 | 0.345 |
| Never called | test | 356 | 207 | 37.3 | 0.022 |
| Analysis rows | test | 3,942 (pool) | – | 50.8 | 0.346 |
| Entry | What happened | Bearing on the results | Seen |
|---|---|---|---|
| Before the freeze | The primary set of H1 changed twice (after the pilot’s parse rates, and after a review in Codex); the scoring defect was split in two; the tied-output-layer defect was redesigned; one evaluability rule was fixed; the family of H4 was fixed at 16; the margins were set to 0.05. | Defines the endpoints; no study accuracy was used. | pilot parse rates |
| Rows left out | 147 replication and 371 test rows that the case study had called, and 7 rows whose answers could be recovered from an upload made by mistake, were called but left out. | Putting them back changes no verdict of H1–H3 and 2 H4 claims. | no outputs |
| A change of GPU | After a restart the platform assigned another GPU of the same model; the determinism check rerun on the pilot rows passed, with outputs equal to the old GPU’s. | None. 9.33 GPU-hours outside the budget. | pilot rows |
| Two script stops | The chain script stopped twice during the GEPA search, both times from a fault in the script. | None; no run was affected. | none |
| The thinking-switch check | A check on Intern-S1-mini ’s thinking switch stopped the run that sends the reference answer, and we let it record instead of stop. The repeat reproduced the first 3,700 calls token for token, and the check failed on 324 of the run’s 1,204 thinking-off calls and on none of 11,265 in the clean stages. | The reference-answer defect in Intern-S1-mini only through the descriptive selection, since the claim reads the start. 6.46 idle server-hours counted. | counts only |
| The checkpoint | 15 of the 16 pairs were evaluable; the pilot had suggested 14. | The unparsed-as-A scoring in Qwen3-1.7B-Longevity is not evaluable ; the gold fallback in Intern-S1-mini (bound 0.057) is evaluable. | counts and bounds |
| Item | Expected | Observed |
|---|---|---|
| H1 sanity checks ( THINK-64 ) | in eight cells; Qwen3-1.7B-Longevity on GSM8K fails the gate | visibility dominates in 8 of nine cells; Qwen3-1.7B-Longevity on GSM8K fails the gate. As expected. |
| H1, Qwen3-1.7B-Longevity (three cells) | Fails the gate (on MMLU-Pro, or quality matters) | no gain: not tested in all three. As expected. |
| H1, Intern-S1-mini on the letters | nearly inevitable | visibility dominates , . As expected. |
| H1, SmolLM3-3B and Intern-S1-mini on GSM8K | Inconclusive on (planned power 0.52 and 0.00 for SmolLM3-3B , 0.45 and 0.06 for Intern-S1-mini , in the two scenarios) | SmolLM3-3B : visibility dominates , [ , ]. Intern-S1-mini : the gate fails, . Both differ. |
| H2, Intern-S1-mini on GSM8K against RESCUE-1500 | Inconclusive (planned adjusted power 0.19) | fails , [ , ]. Differs. |
| H2, GEPA harnesses | Power unknown in advance; may well be inconclusive | holds for all four; all intervals contain zero. |
| # | Contact |
|---|---|
| #39 | The ICML formatting draft of the case study copied to a synchronised folder for an upload. |
| #40 | The LongevityBench split of this study, the earlier-contact flags and the first gold audit built. |
| #41 | Fourteen files uploaded to claude.ai for a review of the reading export and the runner (item-exposure contact 1). |
| #42 | The preregistration draft revised and uploaded for reviews in claude.ai and in Codex, after a check for holdout-derived values. |
| #43 | The main item files built. |
| #44 | The LB-0042 rows on which B1’s fallback gold differs from the derived gold counted. |
| Partition | Model | Configuration | Calls | Split disagreements | Role breaks | Prompt tokens | Completion tokens | Seconds |
|---|---|---|---|---|---|---|---|---|
| Replication | Qwen3 | THINK-64 | 3,603 | 0 | 0 | 480 | 56 | 1.0 |
| THINK-512 | 3,603 | 3 | 0 | 480 | 257 | 3.7 | ||
| THINK-1536 | 3,603 | 3 | 0 | 480 | 258 | 3.7 | ||
| AUTHORS-1500 | 3,603 | 3 | 0 | 484 | 3 | 0.2 | ||
| BUDGET-512-1500 | 3,603 | 4 | 0 | 480 | 257 | 3.6 | ||
| SmolLM3 | THINK-64 | 3,603 | 0 | 0 | 500 | 64 | 1.8 |
| Construction | Samples and | Clean counterpart and resampling | Evaluability bound | |
|---|---|---|---|---|
| A | The row’s reference answer is sent as a final assistant turn, as the case study’s caller did: new calls, built by a separate module. | 1,204 replication rows drawn by whole component (seed 20261012) in proportion to the task counts (344, 172, 344 and 344 rows of LB-0030, LB-0034, LB-0038 and LB-0042), halved by component into and (86 components and 602 rows each; after the flags 562 and 565 rows in 81 components each). Both halves are called in one pass, so is a disjoint half, not a later sample. | The grid’s own outputs on the same rows, in the same halves; one resample serves both. | One: the leaked reference can change any row. |
| B1 | Offline: the clean outputs are rescored with LB-0042’s gold taken from the case study’s historical fallback, which with this release’s labels differs from the derived gold. Unparsed answers stay wrong; other tasks keep the release’s gold. | : the replication partition’s letter rows; : the test partition’s, called after the checkpoint. The fallback changes the gold of 408 of 796 LB-0042 rows of and 794 of 1,577 of . | The main clean instance on the same rows; one resample serves both. | The share of ’s letter rows that the start parses and whose gold B1 changes (exact; B1 changes the gold of 0.2050 of them). |
| B2 | Offline: the clean outputs are rescored by a scorer that never fails closed: an unparsed or conflicting letter answer is scored as the letter A, on every letter task, with the release’s gold. | As B1. | As B1. | The share of ’s letter rows that the start leaves unparsed or conflicting and whose gold is A (exact; it equals B2’s distortion of the start’s accuracy). |
| C | The case study’s skewed split, re-implemented. The pooled LongevityBench rows of both partitions that are not flagged (8,305 of 8,830) are reassigned: components (from all six NHANES tasks, chaining into one giant component that goes to the evolution part) are sorted by size, largest first, and each goes to the part furthest below its target share (evolution 0.50, validation 0.30, holdout 0.20). | and are the two halves of the validation part: its components are sorted by identifier, shuffled (seed 20261016) and dealt alternately. Their rows come from both partitions, so they are not the replication and test rows of the other pairs, and C, computed after the test calls, is not temporally prospective. | The main clean instance, whose and are the study’s replication and test rows. The rows differ, so the two are resampled independently, each with its own seed. | None: C is evaluable by rule, because its mechanism changes the sample and has no row-level bound. |
| D | Offline: the start is THINK-64 in place of THINK-1536 ; the reported quantity is the gain over the start. | : the replication partition’s letter rows; : the test partition’s. | The main clean instance; one resample serves both. | The share of ’s letter rows on which the parsed answers of the two starts differ, an unparsed or conflicting answer counting as one value. |
| E | The tied output layer is served in place of the untied one ( Qwen3-1.7B-Longevity only): the case study’s six configurations on FreeToken, for both variants, with CFG-THINK-1536-ORIG as the start. The confirmatory instance serves the tied variant on and on . | : the replication rows; : the test rows. | The untied variant on both samples; one resample serves both. A second instance (tied on , untied on ) feeds only the descriptive detection matrix. | The share of ’s letter rows on which the tied and untied variants’ parsed answers of the start differ. |
| Rule | Fires when |
|---|---|
| PR, gain rule | and the lower bound of the 95% percentile interval of is at or below zero: the reported gain does not reproduce on . |
| PR, level rule | The 95% percentile interval of excludes zero, whether or not ; the -th resample of is paired with the -th of , each drawn with its own seed. |
| TR | On any call of the instance the server’s prompt-token count differs from the local count of the clean render (the corrected caller’s render of the row’s messages before the final assistant turn), or, on vLLM, its prompt token IDs differ from the clean render’s. For FreeToken only the count is compared. |
| TB1 | A strong trivial baseline (the oracle constant, the frequency of the most frequent gold letter in the sample itself) exceeds 0.90 on a task in or . |
| TB2 | For some task, the trivial constant chosen on (the most frequent gold letter there, ties to the smallest letter) scores on more than 0.05 away from its score on . |
| TB3 | On LB-0038, some configuration’s exact-match rate exceeds three times the best constant age’s, in or . |
| Rule | Fires when |
|---|---|
| SCF1 | On the start’s parse rate is below 0.90 or its truncation rate above 0.10. |
| SCF2 | and is less than 0.05 above the best of three panel members on : chance, the strong trivial baseline and AUTHORS-1500 (in the instances of E, CFG-THINKING-FALSE-64-ORIG). For the letter pool, chance and the baseline are the row-weighted means of the tasks’ values. |
| SCF3 | , and is at least 80% of . |
| IR1 | Any request of the instance has an assistant turn, or a last message that is not a user turn. |
| IR2 | The gold screen hits on any request: the row’s follow-up text, the LB-0038 age as a standalone integer, or another reference answer of more than one character appears in it. |
| IR3 | On any row the scorer’s gold is not among the golds derived from the question and follow-up (apart from the LB-0030 band-endpoint exception). |
| Endpoint | Assumed effect | Inputs from the pilot | Units | Family, | Power |
|---|---|---|---|---|---|
| H2: Intern-S1-mini , GSM8K, against RESCUE-1500 | zero | 527 items | 13; 2.89 | 0.19 (adjusted), 0.52 (unadjusted) | |
| H1: SmolLM3-3B , GSM8K, | S1, S2 | , | 527 items | 9; 2.77 | S1 0.52, S2 0.00 |
| H1: Intern-S1-mini , GSM8K, | S1, S2 | , | 527 items | 9; 2.77 | S1 0.45, S2 0.06 |