cs.LGSep 29, 2026

What Pretraining and Midtraining Make Learnable from Rewards?

Authors: Chiwun Yang, Xiaoyu Li

Organizations: City University of Hong Kong · University of New South Wales

Abstract

A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and reward adaptation in the same parameters, proving how prediction acquires execution or retrieval and rewards learn their task-specific use. Experiments with pretrained Qwen2.5 checkpoints test this division of labor. Across eight worlds, Sequential models trained with correct source and first-operation supervision reach 82.61% success, versus 44.15% for a private-random source control. Memory replay preserves retrieval during reward adaptation, and an independent eight-world confirmation achieves 75.32% task success versus 49.86% after matched alternative-retrieval training. GSM8K and HotpotQA separate accuracy at reward entry, subsequent gain and final performance. Together, these results connect information acquisition, executable computation and reward-guided task learning.

Figures & tables

Appendix figures & tables52 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Replay on Demand: An Emergent Curriculum for Balancing Adaptation and Forgetting in Continued Pretraining

    Sep 30, 2026Lukas Thede, Shengzhuang Chen, Stefan Winzeck +3ReplayPretraining

  2. What do Reward Models Memorize?

    Jul 27, 2026Ivo Verhoeven, Pushkar Mishra, Ekaterina ShutovaProgress Reward ModelingPreference Datasets

  3. Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

    Sep 28, 2026Bohao Wang, Xiaoyan Zhao, Yang Zhang +4Large Language Model PersonalizationPreference Alignment Learning