Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability
Organizations: University of Toronto · Stevens Institute of Technology · University of California, Santa Cruz · University of Waterloo · Vector Institute · Northwestern University
Abstract
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for the surprising finding that even random rewards can improve the performance of large language models (LLMs): one attributes the gains to particular mechanisms within RL training; the other to data contamination. Our results motivate a different view: spurious-reward RL can probe a model's reachability, or what further training can attain from its current state under specified constraints, beyond what is reflected in its current performance. Two OLMo checkpoints with the same accuracy on synthetic arithmetic (3.5%), for example, reach 8.5% and 55% in their best runs under the same correctness-rewarded RL. Examining OLMo checkpoints across pre-training and mid-training reveals three distinct regimes of training response: early on, RL produces little improvement even when correct answers are rewarded; later in pre-training, rewarding correct answers becomes effective while random rewards remain weak; and, upon entering mid-training, even random rewards can produce large gains. A similar ordering appears in a number-masked supervised fine-tuning (SFT) analysis of these checkpoints, suggesting that the pattern is not specific to a particular RL mechanism. Moreover, RL with random rewards offers a distinctive perspective on what training can make an LLM do, since its reward signal supplies no information about which answers are correct. By asking what training can attain without correctness feedback, it addresses the label-leakage side of a central problem in decodability-based probing: whether a successful probe reveals the model's capabilities or learns the task itself.
Figures & tables
| Response | Learner | Initial | Terminal | Mean |
| source | score | max | gain | |
| P1 | P1 | 2.5 | 9.5 | +0.6 |
| P2 | 13.5 | 80.0 | +4.5 | |
| P2 | P1 | 2.5 | 14.0 | +2.9 |
| P2 | 13.5 | 57.0 | +13.7 |
| Synthetic GSM | Code | Invented-fact lookup | |
| P1 / P2 checkpoints | 25 / 12 | 8 / 4 | 8 / 4 |
| Training / evaluation problems | 800 / 200 | 764 / 200 | 800 / 200 |
| GT seeds, P1 / P2 | 4 / 4 | 1 / 4 | 2 / 2 |
| Random seeds, P1 / P2 | 8 / 8 | 2 / 8 | 8 / 8 |
| Response cap (tokens) | 96 | 256 | 96 |
| Training steps | 500 | 500 | 500 |
| Arm | Training content | Math template | Seeds |
| B | Original GSM question | Yes | 32 (8 reused) |
| C | Word-shuffled GSM question | Yes | 16 |
| A | Unrelated MMLU question stem | Yes | 16 |
| D | Random vocabulary tokens | Yes | 32 |
| E | Random-token complete prompt | No | 32 |
| ID | Training position | Training step |
| SmolLM2-1.7B | ||
| S1 | 0.26T | step-125000 |
| S2 | 5.77T | step-2750000 |
| S3 | 9.96T | step-4750000 |
| S4 | 10.75T | step-5125000 |
| OLMo-3-7B | ||
| configuration | reward | reference coefficient | |
| r_g4 / r_g8 | independent random | 0.01 | 4 / 8 |
| gt_g4 | ground truth | 0 | 4 |
| gt_K50 | ground truth for 50 steps, then random | 0.01 | 8 |
| anti_K50 | inverted ground truth for 50 steps, then random | 0.01 | 8 |
| Batch | Runs | Flagged steps / all steps | Runs with a flagged step |
| Completed checkpoint scan | 636 | 519 / 318,000 | 155 |
| Expanded input controls | 152 | 492 / 76,000 | 112 |
| SmolLM2 / OLMo-3 checkpoints | 80 | 46 / 40,000 | 21 |
| Earlier control study | 55 | 0 / 27,500 | 0 |
| Checkpoint | Seed | Lenient | Strict | Format rate |
| P2 +5B | 31003 | 67.5 | 0.0 | 100.0 |
| P2 +5B | 31004 | 59.0 | 0.0 | 0.0 |
| P2 +5B | 31006 | 77.5 | 0.0 | 0.0 |
| P2 +5B | 31007 | 77.0 | 0.0 | 50.5 |
| P2 +9B | 31004 | 78.0 | 43.0 | 54.0 |
| P2 +9B | 31008 | 55.0 | 4.5 | 9.5 |
| Checkpoint | Base | 1 | 2 | 3 | 4 | Max | |
| P1 5B | 0.0 | 0.0 | 0.0 | 0.0 | 0.5 | 0.5 | 4 |
| P1 34B | 0.5 | 2.0 | 0.5 | 0.0 | 1.0 | 2.0 | 4 |
| P1 462B | 3.5 | 4.5 | 3.5 | 4.5 | 8.5 | 8.5 | 4 |
| P1 839B | 3.5 | 55.0 | 33.5 | 16.5 | 19.5 | 55.0 | 4 |
| P1 1,259B | 2.0 | 66.5 | 72.5 | 1.5 | 57.0 | 72.5 | 4 |
| P1 1,469B | 2.5 | 18.5 | 18.0 | 12.0 | 19.5 | 19.5 | 4 |
| Checkpoint | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | Max |
| P1 5B | 0.0 | 0.0 | 0.0 | 0.0 | 0.5 | 0.5 | 0.5 | 0.5 | 0.5 |
| P1 34B | 0.5 | 2.0 | 1.0 | 1.5 | 1.0 | 0.0 | 0.5 | 1.5 | 2.0 |
| P1 462B | 4.0 | 2.5 | 4.0 | 3.5 | 4.5 | 0.0 | 2.5 | 3.0 | 4.5 |
| P1 839B | 1.5 | 2.5 | 1.5 | 3.0 | 4.0 | 4.0 | 4.0 | 3.5 | 4.0 |
| P1 1,259B | 3.0 | 2.5 | 2.5 | 1.0 | 2.0 | 2.5 | 2.5 | 3.5 | 3.5 |
| P1 1,469B | 2.0 | 1.0 | 1.0 | 1.5 | 2.0 | 2.5 | 1.0 | 0.0 | 2.5 |
| Checkpoint | Base | 1 | 2 | 3 | 4 | Max | |
| P1 5B | 0.0 | 0.0 | — | — | — | 0.0 | 1 |
| P1 462B | 0.5 | 0.5 | — | — | — | 0.5 | 1 |
| P1 839B | 0.0 | 0.0 | — | — | — | 0.0 | 1 |
| P1 1,259B | 0.5 | 0.5 | — | — | — | 0.5 | 1 |
| P1 2,098B | 1.0 | 11.0 | — | — | — | 11.0 | 1 |
| P1 2,937B | 0.0 | 0.5 | — | — | — | 0.5 | 1 |
| Checkpoint | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | Max |
| P1 5B | 0.0 | 0.0 | — | — | — | — | — | — | 0.0 |
| P1 462B | 0.0 | 1.0 | — | — | — | — | — | — | 1.0 |
| P1 839B | 0.0 | 0.0 | — | — | — | — | — | — | 0.0 |
| P1 1,259B | 0.0 | 0.0 | — | — | — | — | — | — | 0.0 |
| P1 2,098B | 1.5 | 0.0 | — | — | — | — | — | — | 1.5 |
| P1 2,937B | 0.5 | 0.5 | — | — | — | — | — | — | 0.5 |
| Checkpoint | Base | 1 | 2 | 3 | 4 | Max | |
| P1 5B | 0.0 | 0.0 | 0.0 | — | — | 0.0 | 2 |
| P1 462B | 3.0 | 97.5 | 97.0 | — | — | 97.5 | 2 |
| P1 839B | 22.5 | 98.5 | 99.5 | — | — | 99.5 | 2 |
| P1 1,259B | 27.0 | 99.0 | 98.0 | — | — | 99.0 | 2 |
| P1 2,098B | 20.0 | 99.5 | 99.0 | — | — | 99.5 | 2 |
| P1 2,937B | 5.0 | 99.0 | 98.5 | — | — | 99.0 | 2 |
| Checkpoint | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | Max |
| P1 5B | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| P1 462B | 26.0 | 2.0 | 25.0 | 1.0 | 3.0 | 5.5 | 5.5 | 4.0 | 26.0 |
| P1 839B | 20.5 | 20.5 | 22.0 | 0.0 | 27.0 | 23.5 | 22.0 | 1.5 | 27.0 |
| P1 1,259B | 17.0 | 24.5 | 3.5 | 33.0 | 17.5 | 33.0 | 24.0 | 0.0 | 33.0 |
| P1 2,098B | 12.0 | 11.5 | 16.5 | 19.0 | 15.0 | 15.5 | 10.0 | 12.5 | 19.0 |
| P1 2,937B | 0.0 | 2.0 | 3.0 | 3.0 | 23.0 | 16.0 | 7.0 | 0.0 | 23.0 |
| Checkpoint | Raw | GT max | Random max | Random mean | Gain | Random peak |
| S1 | 0.5 | 1.0 | 1.0 | 0.56 | 0/8 | 1.0 |
| S2 | 0.0 | 1.0 | 4.0 | 1.12 | 0/8 | 4.0 |
| S3 | 1.0 | 3.0 | 8.5 | 2.44 | 0/8 | 8.5 |
| S4 | 3.0 | 78.0 | 7.5 | 4.69 | 0/8 | 7.5 |
| O1 | 1.0 | 0.5 | 1.0 | 0.63 | 0/8 | 1.0 |
| O2 | 15.5 | 93.5 | 15.0 | 8.31 | 0/8 | 31.5 |
| GT | Random | |||||||||
| ID | 1 | 2 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
| S1 | 1.0 | 0.5 | 0.5 | 0.5 | 0.5 | 0.5 | 0.5 | 1.0 | 0.5 | 0.5 |
| S2 | 0.5 | 1.0 | 0.5 | 0.5 | 2.0 | 1.5 | 0.0 | 0.0 | 0.5 | 4.0 |
| S3 | 3.0 | 1.0 | 8.5 | 2.0 | 1.0 | 4.0 | 1.0 | 1.5 | 0.5 | 1.0 |
| S4 | 54.0 | 78.0 | 7.0 | 3.5 | 2.5 | 5.0 | 7.5 | 5.0 | 5.0 | 2.0 |
| O1 | 0.5 | 0.5 | 0.5 | 0.5 | 1.0 | 1.0 | 1.0 | 0.0 | 0.5 | 0.5 |
| Checkpoint | S1 | S2 | S3 | S4 | O1 | O2 | O3 | O4 |
| Full output | 0.5 | 0.0 | 1.0 | 3.0 | 1.0 | 15.5 | 18.0 | 30.0 |
| First block | 3.0 | 4.0 | 14.0 | 27.5 | 1.0 | 53.0 | 35.0 | 34.0 |
| Training input | Max | Mean | Median | Gain | Collapse | |
| B: GSM | 32 | 78.0 | 13.98 | 6.75 | 3/32 | 15/32 |
| C: shuffled GSM | 16 | 55.5 | 15.56 | 14.25 | 2/16 | 6/16 |
| A: MMLU | 16 | 86.5 | 23.28 | 18.50 | 2/16 | 1/16 |
| D: tokens + math template | 32 | 68.5 | 17.50 | 13.50 | 4/32 | 8/32 |
| E: raw tokens | 32 | 23.0 | 20.36 | 20.50 | 0/32 | 0/32 |
| Seed | B | C | A | D | E |
| 31001 | 23.5 | 8.5 | 23.0 | 16.5 | 19.5 |
| 31002 | 7.5 | 0.5 | 15.5 | 22.0 | 19.5 |
| 31003 | 15.0 | 21.0 | 22.5 | 8.5 | 23.0 |
| 31004 | 6.0 | 35.0 | 50.5 | 68.5 | 21.0 |
| 31005 | 0.0 | 20.5 | 19.0 | 23.0 | 20.5 |
| 31006 | 21.5 | 55.5 | 29.5 | 22.5 | 21.5 |
| Model | Input | Max | Mean | SD | Gain | Collapse |
| Qwen2.5-7B | D | 91.5 | 50.75 | 29.78 | 4/8 | 1/8 |
| E | 51.0 | 46.38 | 2.45 | 0/8 | 0/8 | |
| Llama-3.1-8B | D | 64.0 | 16.63 | 19.30 | 1/8 | 3/8 |
| E | 31.5 | 19.56 | 6.94 | 0/8 | 0/8 |
| Qwen2.5-7B | Llama-3.1-8B | |||
| Seed | D | E | D | E |
| 31001 | 57.0 | 51.0 | 1.0 | 31.5 |
| 31002 | 34.0 | 43.0 | 3.0 | 14.5 |
| 31003 | 0.0 | 45.0 | 1.0 | 9.0 |
| 31004 | 25.5 | 48.0 | 19.0 | 16.0 |
| 31005 | 91.5 | 44.0 | 13.5 | 18.0 |
| Source | Paired quantity (P2 minus P1) | Counts | Holm | |
| P1 | Gain | 6/9/1 | 0.6072 | 0.8405 |
| P1 | Improved | 3/0 | 0.2500 | 0.8405 |
| P1 | Absolute gain | 13/2/1 | 0.0074 | 0.0443 |
| P2 | Gain | 11/5/0 | 0.2101 | 0.8405 |
| P2 | Improved | 5/1 | 0.2188 | 0.8405 |
| P2 | Absolute gain | 12/4/0 | 0.0768 | 0.3841 |
| Source | Learner | Primary score | Prefix score | Prefix gain | Prefix gain |
| before after | before after | pp | pp | ||
| P1 | P1 | 2/16 | 10/16 | ||
| P1 | P2 | 8/16 | 0/16 | ||
| P2 | P1 | 1/16 | 11/16 | ||
| P2 | P2 | 8/16 | 0/16 |
| Replicate | P1 P1 | P1 P2 | P2 P1 | P2 P2 |
| 51001 | 3.0 | 4.0 | 1.5 | 48.5 |
| 51002 | 4.0 | 44.0 | 7.0 | 14.5 |
| 51003 | 4.0 | 23.0 | 1.0 | 13.5 |
| 51004 | 8.0 | 2.0 | 4.0 | 10.5 |
| 51005 | 1.0 | 43.5 | 3.0 | 47.5 |
| 51006 | 0.5 | 11.5 | 7.0 | 19.0 |
| seed | P1 (3,896B) | P2 +5B | P2 +50B |
| r_g4 ( = 4) | |||
| 777 | 3.0 / 7.5 (+4.5) | 13.0 / 26.5 (+13.5) ⋆ | 21.0 / 1.5 ( 19.5) |
| 888 | 1.0 / 27.0 (+26.0) ⋆ | 9.0 / 2.5 ( 6.5) | 17.5 / 0.5 ( 17.0) |
| 999 | 4.5 / 9.0 (+4.5) | 20.0 / 17.5 ( 2.5) | 27.0 / 76.5 (+49.5) ⋆ |
| 1001 | 4.5 / 0.0 ( 4.5) | 14.0 / 19.5 (+5.5) | 19.5 / 6.0 ( 13.5) |
| 1002 | 1.0 / 8.0 (+7.0) | 14.5 / 49.0 (+34.5) ⋆ | 16.0 / 85.5 (+69.5) ⋆ |
| Checkpoint | = 4 successes | = 4 mean | = 8 successes | = 8 mean |
| P1 3,896B | 1/5 | +7.5 | 0/5 | +0.8 |
| P2 +5B | 2/5 | +8.9 | 0/5 | 4.0 |
| P2 +50B | 2/5 | +13.8 | 0/5 | 6.1 |
| arm | after accuracy (777 / 888 / 999 / 1001 / 1002) | mean |
| gt_g4 P1 | 90.5 / 92.5 / 81.5 / 81.0 / 84.0 | +83.1 |
| gt_g4 P2 +5B | 99.0 / 97.0 / 98.0 / 99.5 / 100.0 | +84.6 |
| gt_g4 P2 +50B | 100.0 / 98.0 / 100.0 / 98.0 / 21.5 | +63.3 |
| gt_K50 | 90.5 / 1.0 / 62.0 / 3.0 / 54.0 | +28.0 |
| anti_K50 | 1.5 / 1.5 / 16.5 / 3.5 / 9.0 | 7.7 |
| Quantity | D | E | Difference | |||
| OLMo-2 | ||||||
| Mean signed gain (pp) | 3.00 | 0.14 | 2.86 | 0.3337 | 1 | 1 |
| Mean absolute change (pp) | 12.59 | 1.11 | +11.48 | |||
| Gain pp | 4/32 | 0/32 | +12.50 | 0.125 | 1 | 1 |
| Absolute change pp | 16/32 | 0/32 | +50.00 | 0.0088 | ||
| Terminal score at most 5% | 8/32 | 0/32 | +25.00 | 0.0078 | 0.1016 | 1 |
| Model | Input | Gain: count [95% CI] | Change: count [95% CI] |
| OLMo-2 | B | 3/32 [2.0, 25.0] | 21/32 [46.8, 81.4] |
| OLMo-2 | C | 2/16 [1.6, 38.3] | 10/16 [35.4, 84.8] |
| OLMo-2 | A | 2/16 [1.6, 38.3] | 5/16 [11.0, 58.7] |
| OLMo-2 | D | 4/32 [3.5, 29.0] | 16/32 [31.9, 68.1] |
| OLMo-2 | E | 0/32 [0.0, 10.9] | 0/32 [0.0, 10.9] |
| Qwen2.5 | D | 4/8 [15.7, 84.3] | 7/8 [47.3, 99.7] |
| Pair | Mean diff. | Rate diff. | |||||
| B C | 16 | 2.59 | 0.5857 | 1 | 6.25 | 1 | 1 |
| B A | 16 | 10.31 | 0.0023 | 0.0417 | 6.25 | 1 | 1 |
| B D | 32 | 3.64 | 0.4467 | 1 | 3.12 | 1 | 1 |
| B E | 32 | 6.50 | 0.0672 | 1 | +9.38 | 0.25 | 1 |
| C A | 16 | 7.72 | 0.1971 | 1 | 0.00 | 1 | 1 |
| C D | 16 | 5.16 | 0.3455 | 1 | 0.00 | 1 | 1 |
| Comparison | Seeds | Difference | ||
| GSM GT: P1 3,896B P1 34B | 4 | +81.12 | 0.125 | 0.875 |
| GSM GT: P2 +5B P1 3,896B | 4 | +17.38 | 0.125 | 0.875 |
| GSM GT: P1 839B P1 462B | 4 | +25.88 | 0.125 | 0.875 |
| GSM Random: P1 3,896B P1 34B | 8 | +5.81 | 0.0156 | 0.1719 |
| GSM Random: P2 +5B P1 3,896B | 8 | +33.44 | 0.0156 | 0.1719 |
| GSM Random: P1 839B P1 462B | 8 | 0.00 | 1 | 1 |
| Pair | GT diff. | Random diff. | ||||
| S2 S1 | 0.00 | 1 | 1 | +0.56 | 0.4375 | 1 |
| S3 S1 | +1.25 | 0.5 | 1 | +1.88 | 0.0156 | 0.4219 |
| S4 S1 | +65.25 | 0.5 | 1 | +4.12 | 0.0078 | 0.25 |
| S3 S2 | +1.25 | 1 | 1 | +1.31 | 0.3594 | 1 |
| S4 S2 | +65.25 | 0.5 | 1 | +3.56 | 0.0234 | 0.5859 |
| S4 S3 | +64.00 | 0.5 | 1 | +2.25 | 0.0469 | 1 |
| Comparison | Mean diff. | Rate diff. | ||||
| P1 3,896B | +6.70 | 0.3125 | 1 | +20.00 | 1 | 1 |
| P2 +5B | +12.90 | 0.1875 | 1 | +40.00 | 0.5 | 1 |
| P2 +50B | +19.90 | 0.4375 | 1 | +40.00 | 0.5 | 1 |
| Three-checkpoint mean | +13.17 | 0.3125 | 1 | +33.33 | 0.125 | 1 |
| GT anti steering | +35.70 | 0.25 | 1 | +60.00 | 0.25 | 1 |
| Checkpoint | GT | R | GT R | ||||||
| P1 5B | 1/0/3 | 1 | 1 | 4/0/4 | 0.125 | 1 | +0.12 | 1 | 1 |
| P1 34B | 2/1/1 | 1 | 1 | 5/1/2 | 0.2188 | 1 | 0.38 | 0.75 | 1 |
| P1 462B | 3/0/1 | 0.25 | 1 | 3/4/1 | 1 | 1 | +1.75 | 0.125 | 1 |
| P1 839B | 4/0/0 | 0.125 | 1 | 3/4/1 | 1 | 1 | +29.00 | 0.125 | 1 |
| P1 1,259B | 3/1/0 | 0.625 | 1 | 6/1/1 | 0.125 | 1 | +47.12 | 0.25 | 1 |
| P1 1,469B | 4/0/0 | 0.125 | 1 | 1/5/2 | 0.2188 | 1 | +15.62 | 0.125 | 1 |
| Checkpoint | GT | R | GT R | ||||||
| P1 5B | 0/0/1 | – | – | 0/0/2 | 1 | 1 | 0.00 | – | – |
| P1 462B | 0/0/1 | – | – | 1/1/0 | 1 | 1 | +0.50 | – | – |
| P1 839B | 0/0/1 | – | – | 0/0/2 | 1 | 1 | 0.00 | – | – |
| P1 1,259B | 0/0/1 | – | – | 0/2/0 | 0.5 | 1 | +0.50 | – | – |
| P1 2,098B | 1/0/0 | – | – | 1/1/0 | 1 | 1 | +9.50 | – | – |
| P1 2,937B | 1/0/0 | – | – | 2/0/0 | 0.5 | 1 | 0.00 | – | – |
| Checkpoint | GT | R | GT R | ||||||
| P1 5B | 0/0/2 | 1 | 1 | 0/0/8 | 1 | 1 | 0.00 | 1 | 1 |
| P1 462B | 2/0/0 | 0.5 | 1 | 5/3/0 | 0.7266 | 1 | +83.25 | 0.5 | 1 |
| P1 839B | 2/0/0 | 0.5 | 1 | 2/5/1 | 0.4531 | 1 | +78.50 | 0.5 | 1 |
| P1 1,259B | 2/0/0 | 0.5 | 1 | 2/6/0 | 0.2891 | 1 | +77.75 | 0.5 | 1 |
| P1 2,098B | 2/0/0 | 0.5 | 1 | 0/8/0 | 0.0078 | 1 | +87.50 | 0.5 | 1 |
| P1 2,937B | 2/0/0 | 0.5 | 1 | 3/5/0 | 0.7266 | 1 | +97.75 | 0.5 | 1 |
| Checkpoint | GT | R | GT R | ||||||
| O1 | 0/2/0 | 0.5 | 1 | 0/5/3 | 0.0625 | 1 | 0.00 | 1 | 1 |
| O2 | 2/0/0 | 0.5 | 1 | 0/8/0 | 0.0078 | 1 | +79.25 | 0.5 | 1 |
| O3 | 2/0/0 | 0.5 | 1 | 3/5/0 | 0.7266 | 1 | +49.25 | 0.5 | 1 |
| O4 | 2/0/0 | 0.5 | 1 | 4/4/0 | 1 | 1 | +44.50 | 0.5 | 1 |
| S1 | 1/0/1 | 1 | 1 | 1/0/7 | 1 | 1 | +0.25 | 1 | 1 |
| S2 | 2/0/0 | 0.5 | 1 | 6/0/2 | 0.0312 | 1 | +0.25 | 1 | 1 |
| Cell | Mean gain | Gain: count [95% CI] | |||
| OLMo-2 B | 6.64 | 6/25/1 | 0.1317 | 3/32 [2.0, 25.0] | |
| OLMo-2 C | 4.94 | 6/9/1 | 0.6072 | 1 | 2/16 [1.6, 38.3] |
| OLMo-2 A | +2.78 | 7/9/0 | 0.8036 | 1 | 2/16 [1.6, 38.3] |
| OLMo-2 D | 3.00 | 13/19/0 | 0.3771 | 1 | 4/32 [3.5, 29.0] |
| OLMo-2 E | 0.14 | 13/13/6 | 1 | 1 | 0/32 [0.0, 10.9] |
| Qwen2.5 D | +6.56 | 4/4/0 | 1 | 1 | 4/8 [15.7, 84.3] |
| Checkpoint / task | Gain: count [95% CI] | High score: count [95% CI] |
| GSM P1 5B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| GSM P1 34B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| GSM P1 462B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| GSM P1 839B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| GSM P1 1,259B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| GSM P1 1,469B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| Checkpoint / task | Gain: count [95% CI] | High score: count [95% CI] |
| GSM P1 3,896B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| GSM P2 +5B | 4/8 [15.7, 84.3] | 4/8 [15.7, 84.3] |
| GSM P2 +9B | 2/8 [3.2, 65.1] | 2/8 [3.2, 65.1] |
| GSM P2 +13B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| GSM P2 +17B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| GSM P2 +21B | 3/8 [8.5, 75.5] | 3/8 [8.5, 75.5] |
| Checkpoint / task | Gain: count [95% CI] | High score: count [95% CI] |
| CODE P2 +50B | 1/8 [0.3, 52.7] | 0/8 [0.0, 36.9] |
| Lookup P1 5B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| Lookup P1 462B | 2/8 [3.2, 65.1] | 0/8 [0.0, 36.9] |
| Lookup P1 839B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| Lookup P1 1,259B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |
| Lookup P1 2,098B | 0/8 [0.0, 36.9] | 0/8 [0.0, 36.9] |