Language Models as AI Research World Models
Organizations: Amazon · UC Santa Cruz · Emory University · Pennsylvania State University · UC Santa Barbara
Abstract
AI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement. Yet their ability to propose experiments outpaces their capacity to execute them in real environments, making outcome prediction a key capability for sustained self-improvement under limited experimental budgets. We investigate language models as Research World Models (RWMs), which predict the outcomes of candidate interventions across research environments. Our evaluation draws on over 2,600 experimental records from nine research environments spanning pretraining, post-training, and inference, representing more than 171,000 H100 GPU-hours of experimentation. Research knowledge acquired from real experimental experience improves RWM predictions of unseen interventions within the same environment (Spearman +0.27), and can be reused across environments. For example, using only pretraining experience from OLMo3, Marin, and Nanochat, an RWM reduces selection regret in the Qwen3 environment by 78% compared with zero-experience setting. These benefits extend to multi-round Autoresearch under a fixed selection budget: RWMs with in-env and cross-env research knowledge increase the best gain achieved by 15.8% and 11.6%, respectively. Ablations across 13 language models used as RWMs show that adding research knowledge can improve intervention ranking more than changing models or increasing reasoning effort alone. These findings support language models as RWMs and motivate accumulating experimental data for future RWM training.
Figures & tables
| Research Env. | Model / System | Eval metric | Records | H100 GPU-hours | In-env | Cross-env |
| Pretraining | ||||||
| OLMo3-100M | OLMo3-100M | mean BPB | 480 | 8,572 | ||
| Marin | Marin-153M | mean BPB | 132 | 12,101 | — | |
| Nanochat | Nanochat-149M | mean BPB | 132 | 2,584 | — | |
| Qwen3 | Qwen3-153M | mean BPB | 132 | 15,447 | — | |
| OLMo3-190M | OLMo3-190M | mean BPB | 132 | 5,582 | — | |
| MAE | Spearman | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Environment (unit) | History | Test | kNN | ridge | kNN | ridge | ||||
| OLMo3-100M (BPB) | 320 | 160 | 2.93e-2 | 2.58e-2 | 3.13e-2 | 2.03e-2 | 0.62 | 0.78 | 0.67 | 0.90 |
| Diffusion (BPB) | 536 | 125 | 1.95e-2 | 1.58e-2 | 1.99e-2 | 8.77e-3 | 0.46 | 0.49 | 0.50 | 0.78 |
| Math distillation (pp) | 180 | 100 | 1.58 | 1.15 | 1.11 | 0.91 | 0.43 | 0.61 | 0.62 | 0.70 |
| Code RL (pp) | 34 | 115 | 5.98 | 5.45 | 5.02 | 4.17 | 0.48 | 0.29 | 0.48 | 0.62 |
| Inference optimization (tokens/s) | 457 | 98 | 13.19 | 13.18 | 12.73 | 1.34 | 0.54 | 0.62 | 0.72 | 0.88 |
| Spearman | MAE (BPB) | |||||||
|---|---|---|---|---|---|---|---|---|
| Target Environment | kNN | ridge | kNN | ridge | ||||
| OLMo3-100M | 0.69 | 0.56 | 0.26 | 0.83 | 1.11e-2 | 1.18e-2 | 1.65e-2 | 9.28e-3 |
| OLMo3-190M | 0.72 | 0.43 | 0.26 | 0.78 | 7.94e-3 | 1.08e-2 | 1.19e-2 | 8.56e-3 |
| Qwen3 | 0.74 | 0.71 | 0.40 | 0.82 | 9.20e-3 | 1.13e-2 | 1.69e-2 | 7.80e-3 |
| Marin | 0.42 | 0.28 | 0.43 | 0.50 | 9.35e-3 | 1.13e-2 | 1.03e-2 | 8.18e-3 |
| Nanochat | 0.57 | 0.71 | 2.79e-2 | 3.61e-2 | 3.57e-2 | 2.70e-2 | ||
| Regret@3 | NDCG@3 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Environment (unit) | Random | kNN | ridge | Random | kNN | ridge | ||||
| OLMo3-100M (BPB) | 1.32e-2 | 2.80e-3 | 1.81e-3 | 2.14e-3 | 7.00e-4 | 0.28 | 0.70 | 0.78 | 0.72 | 0.84 |
| Diffusion (BPB) | 1.09e-2 | 3.14e-3 | 1.79e-3 | 5.47e-3 | 4.70e-4 | 0.16 | 0.51 | 0.62 | 0.32 | 0.87 |
| Math distillation (pp) | 0.87 | 0.48 | 0.41 | 0.45 | 0.29 | 0.20 | 0.33 | 0.38 | 0.34 | 0.47 |
| Code RL (pp) | 4.16 | 2.20 | 2.12 | 1.54 | 0.83 | 0.42 | 0.59 | 0.64 | 0.73 | 0.81 |
| Inference optimization (tokens/s) | 4.31 | 3.60 | 1.62 | 2.09 | 2.53 | 0.69 | 0.86 | 0.93 | 0.92 | 0.93 |
| Regret@3 (BPB) | NDCG@3 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Target | Random | kNN | ridge | Random | kNN | ridge | ||||
| OLMo3-100M | 1.31e-2 | 4.27e-3 | 4.74e-3 | 4.84e-3 | 2.24e-3 | 0.26 | 0.65 | 0.66 | 0.66 | 0.84 |
| OLMo3-190M | 8.42e-3 | 1.64e-3 | 1.60e-3 | 1.96e-3 | 5.50e-4 | 0.24 | 0.68 | 0.66 | 0.64 | 0.81 |
| Qwen3 | 1.51e-2 | 9.40e-4 | 3.50e-4 | 1.72e-3 | 2.10e-4 | 0.27 | 0.77 | 0.85 | 0.73 | 0.91 |
| Marin | 1.07e-2 | 7.40e-4 | 1.10e-3 | 8.40e-4 | 7.90e-4 | 0.18 | 0.74 | 0.61 | 0.67 | 0.70 |
| Nanochat | 8.27e-3 | 4.94e-3 | 1.53e-2 | 1.51e-2 | 2.22e-3 | 0.14 | 0.32 | 0.12 | 0.06 | 0.43 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Block | Contents |
|---|---|
| Environment | Model and data configuration, resource budget, reference recipe and measurements, evaluation suite, and gain definition. |
| Intervention | Materialized code diff or configuration override and intervention direction. |
| Outcome | Raw evaluation measurements and reference-relative gain when defined by the scoring protocol. |
| Metadata | Execution status and random seed. |
| Target | vs. | vs. kNN | vs. ridge |
|---|---|---|---|
| OLMo3-100M | |||
| OLMo3-190M | |||
| Qwen3 | |||
| Marin | |||
| Nanochat |