What Pretraining and Midtraining Make Learnable from Rewards?
Organizations: City University of Hong Kong · University of New South Wales
Abstract
A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and reward adaptation in the same parameters, proving how prediction acquires execution or retrieval and rewards learn their task-specific use. Experiments with pretrained Qwen2.5 checkpoints test this division of labor. Across eight worlds, Sequential models trained with correct source and first-operation supervision reach 82.61% success, versus 44.15% for a private-random source control. Memory replay preserves retrieval during reward adaptation, and an independent eight-world confirmation achieves 75.32% task success versus 49.86% after matched alternative-retrieval training. GSM8K and HotpotQA separate accuracy at reward entry, subsequent gain and final performance. Together, these results connect information acquisition, executable computation and reward-guided task learning.
Figures & tables
Appendix figures & tables52 assets
Supplementary material from the paper’s appendix.
Appendix
| Module | Effect | 95% CI | Exact | Holm |
| Sequential | 67.57 | [37.04, 98.10] | 0.007812 | 0.015625 |
| Memory | 0.67 | [-1.22, 2.57] | 0.640625 | 0.640625 |
| Module | Comparison | Metric | Effect | 95% CI | Exact |
| Seq. | Correct: action–state | Select. | 43.55 | [1.47, 85.64] | 0.08594 |
| Seq. | Correct: action–state | Success | 54.62 | [23.75, 85.50] | 0.01562 |
| Seq. | Control: action–state | Select. | 0.96 | [-0.09, 2.02] | 0.04688 |
| Seq. | Control: action–state | Success | 31.74 | [13.15, 50.32] | 0.01562 |
| Seq. | Source interaction | Select. | 42.59 | [0.82, 84.36] | 0.08594 |
| Seq. | Source interaction | Success | 22.88 | [-1.44, 47.20] | 0.08594 |
| Module | Route | Entry | Gain | Final |
| Sequential | Correct/state | |||
| Sequential | Control/state | |||
| Sequential | Correct/action | |||
| Sequential | Control/action | |||
| Memory | Correct/pure | |||
| Memory | Control/pure |
| World | Selected A | Selected B | Other A | Other B | Accuracy | Selectivity |
| 508 | 73.34 | 75.10 | 52.44 | 50.68 | 22.66 | 45.31 |
| 509 | 83.69 | 91.31 | 49.32 | 49.32 | 38.18 | 76.37 |
| 510 | 80.66 | 81.84 | 49.80 | 49.41 | 31.64 | 63.28 |
| 511 | 70.90 | 73.05 | 47.56 | 48.63 | 23.88 | 47.75 |
| 512 | 73.54 | 71.09 | 50.10 | 50.59 | 21.97 | 43.95 |
| 513 | 73.83 | 72.27 | 48.83 | 50.00 | 23.63 | 47.27 |
| Module | Effect | 95% interval | ||
| Sequential | ||||
| Memory |
| Comparison | Independent seeds and scope |
| Direct RL, correct source, paired wrong source | 8 sequential and 5 contextual, 0.5B |
| Random initialization and 1.5B transfer | 3 sequential seeds per three-arm panel |
| Source dose | 5 sequential seeds, 0.5B, shared source |
| Alternative learning routes | 3 sequential LM seeds |
| Alternate downstream task choice | 3 sequential seeds, shared source states |
| Format-only source then RL | 3 sequential seeds |
| Task / model / initialization | Source rate | Reward rate |
| Contextual memory / 0.5B / pretrained | ||
| Sequential / 0.5B / pretrained | ||
| Sequential / 0.5B / random | ||
| Sequential / 1.5B / pretrained | ||
| GSM8K / 1.5B / pretrained | ||
| HotpotQA / 1.5B / pretrained |
| Measurement | Run A | Run B |
| Two-candidate completion accuracy | 46/96 | 49/96 |
| Original records correct | 14/32 | 17/32 |
| Prediction reverses on relevant-value flip | 1/32 | 0/32 |
| Original and relevant-flip pair both correct | 0/32 | 0/32 |
| Query-field CE, before | 0.2045 | 0.2114 |
| Query-field CE, after | 0.0430 | 0.0397 |
| Task | Method | Decoding | Metric | Mean and 95% CI |
|---|---|---|---|---|
| GSM8K | Direct RL | greedy | numeric correct | |
| GSM8K | Direct RL | greedy | Format error | |
| GSM8K | Direct RL | greedy | Length (tokens) | |
| GSM8K | Direct RL | greedy | truncated | |
| GSM8K | Correct source | greedy | numeric correct | |
| GSM8K | Correct source | greedy | Format error |
| Method | Decoding | Type | Metric | Mean and 95% CI |
|---|---|---|---|---|
| Direct RL | greedy | bridge | Answer F1 | |
| Direct RL | greedy | bridge | success | |
| Direct RL | greedy | comparison | Answer F1 | |
| Direct RL | greedy | comparison | success | |
| Correct source | greedy | bridge | Answer F1 | |
| Correct source | greedy | bridge | success |
| Training | Method | Evaluation | Checkpoint | Seeds | Success (%) |
|---|---|---|---|---|---|
| GSM8K | Base/Bo16 | GSM8K | base | 1 | |
| GSM8K | Correct source/Bo16 | GSM8K | source | 5 | |
| GSM8K | Direct RL | BBH logic | feedback | 5 | |
| GSM8K | Direct RL | GSM-Symbolic | feedback | 5 | |
| GSM8K | Direct RL | BBH tracking | feedback | 5 | |
| GSM8K | Correct source | BBH logic | feedback | 5 |
| Training route | Seeds | Before (%) | After (%) | Gain (pp) |
| Sequential | ||||
| Correct source | 8 | |||
| Direct RL | 8 | |||
| Wrong source | 8 | |||
| Format source | 3 | |||
| Memory | ||||
| Source prediction tokens | Worlds | Final success (%) |
| 5 | ||
| 5 | ||
| 5 |
| Setting | Route | Worlds | Success (%) |
| 0.5B, pretrained, source tokens | Balanced source | 3 | |
| Interleaved source updates | 3 | ||
| Classical table learner | 3 | ||
| 0.5B, random initialization, source tokens | Correct source | 3 | |
| Wrong source | 3 | ||
| Direct RL | 3 |
| Route | Seeds | Public path (%) | Correct public (%) | Private public (%) |
| Correct source | 8 | |||
| Wrong source | 8 | |||
| Direct RL | 8 | – | ||
| Format source | 3 | – |
| Route | Seed | ||||
| Correct source | 0 | 1024 | 1012 | 695 | 345 |
| Correct source | 1 | 1024 | 1024 | 1021 | 515 |
| Correct source | 2 | 1024 | 1024 | 824 | 381 |
| Correct source | 3 | 1024 | 1024 | 773 | 412 |
| Correct source | 4 | 1024 | 1024 | 1017 | 497 |
| Correct source | 5 | 1024 | 1024 | 771 | 405 |
| Route | Reward task | Test A (%) | Test B (%) | A A or B (%) |
| Correct source | A | |||
| Correct source | B | |||
| Direct RL | A | – | ||
| Direct RL | B | – |
| Route | Seed | ||||
| Correct source | 3 | 1024 | 773 | 412 | 361 |
| Correct source | 4 | 1024 | 1017 | 497 | 520 |
| Correct source | 5 | 1024 | 771 | 405 | 366 |
| Source / task B | 3 | 1024 | 824 | 439 | 385 |
| Source / task B | 4 | 1024 | 988 | 482 | 506 |
| Source / task B | 5 | 1024 | 685 | 364 | 321 |
| Route | Seeds | Valid (%) | Correct (%) | Correct valid (%) | Constant runs |
| Correct source | 5 | 5 | |||
| Wrong source | 5 | 5 | |||
| Direct RL | 5 | – | 0 |
| Route | Malformed | Path error | Public error | Private error | Success |
| Correct source | |||||
| Wrong source | |||||
| Direct RL | |||||
| Format source |
| Task | Evaluation | Comparator | Seeds | Difference (pp) |
| sequential | greedy | Direct RL | 8 | |
| sequential | greedy | Wrong source | 8 | |
| sequential | greedy | Format source | 3 | |
| contextual | greedy | Direct RL | 5 | |
| contextual | greedy | Wrong source | 5 | |
| GSM8K | greedy | Direct RL | 5 |
| Training route | Seeds | Before (%) | After (%) | Gain (pp) |
| Sequential | ||||
| Correct source | 8 | |||
| Direct RL | 8 | |||
| Wrong source | 8 | |||
| Format source | 3 | |||
| Memory | ||||
| Route | Seeds | Source tokens | Answer labels | Success (%) |
| Source only, all worlds | 8 | 1,048,576 | 0 | |
| Source only, paired worlds | 3 | 1,048,576 | 0 | |
| Continue source prediction | 3 | 2,097,152 | 0 | |
| Answer-only adaptation | 3 | 1,048,576 | 8,192 |
| Candidates | Seeds | Prompts | Test queries/seed | Success (%) |
| 1 | 3 | 1,024 | 1,024 | |
| 2 | 3 | 1,024 | 2,048 | |
| 4 | 3 | 1,024 | 4,096 | |
| 8 | 3 | 1,024 | 8,192 |
| Checkpoint | Independent units | Prompts | Success (%) |
| Logical deduction (250 prompts) | |||
| Base checkpoint | 1 fixed model | 250 | |
| After source prediction | 3 training seeds | 250 | |
| After source + reward | 3 training seeds | 250 | |
| Object tracking (250 prompts) | |||
| Base checkpoint | 1 fixed model | 250 | |
| Task | Seeds | Prompts/seed | Answer labels/seed | Success (%) |
| GSM8K | 3 | 1,319 | 8,192 | |
| HotpotQA | 3 | 1,024 | 8,192 |
| Source tokens | Before RL (%) | After RL (%) | Gain (pp) |
| GSM8K; 8,192 reward queries | |||
| None (direct) | 8.57\,\text{\footnotesize\pm 0.00} | 62.32\,\text{\footnotesize\pm 4.14} | 53.75\,\text{\footnotesize\pm 4.14} |
| 3.11\,\text{\footnotesize\pm 0.57} | 61.76\,\text{\footnotesize\pm 3.15} | 58.65\,\text{\footnotesize\pm 3.67} | |
| 2.32\,\text{\footnotesize\pm 0.39} | 62.05\,\text{\footnotesize\pm 2.93} | 59.73\,\text{\footnotesize\pm 3.16} | |
| HotpotQA; 8,192 reward queries | |||
| None (direct) | 23.93\,\text{\footnotesize\pm 0.00} | 47.60\,\text{\footnotesize\pm 1.02} | 23.67\,\text{\footnotesize\pm 1.02} |
| Contrast | Endpoint difference (pp) | Gain difference (pp) |
| GSM8K | ||
| minus direct | -0.56\,\text{\footnotesize\pm 4.87} | 4.90\,\text{\footnotesize\pm 5.28} |
| minus direct | -0.27\,\text{\footnotesize\pm 3.63} | 5.97\,\text{\footnotesize\pm 3.66} |
| minus | -0.29\,\text{\footnotesize\pm 5.64} | -1.08\,\text{\footnotesize\pm 6.38} |
| HotpotQA | ||
| minus direct | -1.37\,\text{\footnotesize\pm 0.37} | -12.99\,\text{\footnotesize\pm 1.67} |
| Representation | Objective | World | Content before | Content after | Pair before/after |
|---|---|---|---|---|---|
| Sequential; content /512 | |||||
| Original | State | 1 | 512 | 11 | – |
| Original | State | 2 | 512 | 12 | – |
| Original | Weighted | 1 | 512 | 6 | – |
| Original | Weighted | 2 | 512 | 14 | – |
| Curriculum | State | 1 | 512 | 6 | – |
| Route and remaining reward problem | Observed source tokens | Training verifier queries |
| Information resources: identify a mechanism and its task binding | ||
| Few source observations; reward queries recover the mechanism class | ||
| Logarithmic ; with fixed ; Theorem C.2 . | ||
| Learn a transition table first; rewards identify only the initial operation | At most | |
| Neural training: acquire execution, then adapt the same parameters | ||
| Source Adam and two sampled reward updates | At most | |