RSI-Forge: From Research Papers to Environments for Recursive Self-Improvement
Organizations: Scale AI · University of California, Santa Cruz · University of North Carolina at Chapel Hill
Abstract
Environments are the foundation of recursive self-improvement: they provide the problems agents work on and the feedback used to evaluate progress. Yet constructing challenging research environments with reliable evaluation still depends on domain experts, limiting their scale and disciplinary coverage. We introduce RSI-Forge, a multi-agent pipeline that turns published papers into executable environments for self-improvement. Three agents coordinate construction, reproduction, and review to produce tasks with automated evaluators; each paper's method is independently reimplemented to establish a baseline score. We present 210 environments across 18 fields, including 90 reviewed by independent human domain experts. Both experts and agent judges give high ratings to the potential for improving the provided starting solutions and the evaluators' ability to distinguish solution quality, whereas experts are more critical of shortcut resistance, faithfulness to the source paper, and whether a single idea can exhaust a task. To validate their use for repeated improvement, we evaluate four models over 3 successive attempts on 120 environments, with each attempt inheriting prior code and notes while model weights remain fixed. At least one model improves after the first attempt in 84% of environments. Models also outperform the reproduced paper methods in 68 of the 120 environments, demonstrating room for gains beyond these baselines. Transcript analysis identifies work beyond parameter tuning in 95% of these successful attempts. Analysis of the resulting trajectories shows that models scoring lower on these tasks explore less, more often accept gains smaller than the reported standard error, and rely more heavily on tuning to the development set. RSI-Forge provides a scalable approach to constructing research environments for training and evaluating self-improving agents.
Figures & tables
| Stage | Count |
| Source papers screened | 10,800 |
| Papers eligible for construction | 2,117 |
| Construction attempts begun | 213 |
| Passed Design | 199 |
| Passed Build | 198 |
| Passed Independent Reproduction | 189 |
| Field | Task | Evaluation objective |
|---|---|---|
| Robotics | Choose footholds on a terrain map | Increase distance in six steps |
| Life Sciences | Propose molecules for a binding pocket | Improve docking scores |
| Earth, Climate & Space | Estimate weather from partial observations | Reduce state-estimation error |
| Statistics & Inference | Estimate effects without a control series | Reduce treatment-effect error |
| Software Engineering & Formal Methods | Generate tests within a fixed budget | Trigger more distinct failures |
| Computational Physics | Infer a plasma boundary from magnetic data | Reduce boundary-estimation error |
| Issue identified | Share of repair passages |
|---|---|
| Checks unable to detect the intended failure | 66% |
| Incorrect or incomplete instructions | 63% |
| Checks inspecting a different copy of the code | 54% |
| Score gains without performing the intended work | 37% |
| Median normalized score | Session 3 | ||||
|---|---|---|---|---|---|
| Model | Session 1 | Session 2 | Session 3 | 95% CI | Mean rank |
| Claude Opus 5 | 0.754 | 0.791 | 0.800 | [0.753, 0.867] | 1.67 |
| GPT-5.6-sol | 0.673 | 0.747 | 0.761 | [0.723, 0.824] | 1.73 |
| DeepSeek-V4-Flash | 0.550 | 0.594 | 0.647 | [0.586, 0.711] | 2.97 |
| Claude Haiku 4.5 | 0.129 | 0.277 | 0.332 | [0.203, 0.451] | 3.63 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Field | Count | Field | Count |
|---|---|---|---|
| Robotics | 14 | Security & Cryptography | 10 |
| Life Sciences | 13 | Neuroscience | 9 |
| Earth, Climate & Space | 13 | Optimisation & Discrete Algorithms | 9 |
| Statistics & Inference | 13 | Chemistry & Materials | 9 |
| Networks & Information Th. | 12 | Graphics & Rendering | 8 |
| Software Eng. & Formal Methods | 12 | Quantitative Finance | 7 |
| Dimension | Experts | SD | Panel | SD | Gap |
|---|---|---|---|---|---|
| Research problem | 4.39 | 0.84 | 4.59 | 0.32 | |
| Room above starting solution | 4.81 | 0.51 | 4.89 | 0.31 | |
| Resistance to saturation | 3.09 | 0.96 | 3.98 | 0.84 | |
| Faithfulness to source paper | 3.72 | 0.87 | 4.22 | 0.51 | |
| Metric validity | 4.40 | 0.85 | 4.49 | 0.45 | |
| Single objective | 4.61 | 0.90 | 4.66 | 0.46 |
| Dimension | RSI-Exam | RSI-Forge | Gap |
|---|---|---|---|
| Score discrimination | 2.64 | 4.65 | |
| Development–test distinction | 3.23 | 4.67 | |
| Shortcut resistance | 3.92 | 4.67 | |
| Single objective | 4.64 | 4.66 | |
| Room above starting solution | 4.87 | 4.89 | |
| Research problem | 4.61 | 4.59 |
| Axis | A score of 1 | A score of 3 | A score of 5 |
|---|---|---|---|
| Real problem | a toy with no research content | a real problem, narrowly posed | you would be pleased to see a student work on this |
| Headroom | the starter is already near the ceiling | a competent effort gains something | clearly large, and you can name where it comes from |
| Not saturable | one change takes most of the range | the first change is worth it, more remains | improvement requires several independent ideas |
| Faithful to the paper | unrelated to what the paper does | related, but the hard part has been removed | this is the paper’s problem, honestly posed |
| Metric validity | measures something else | a reasonable proxy | this is the quantity the field would report |
| Single direction | two objectives silently traded | one quantity, the constraint is loose | one quantity, one direction, the constraint binds |
| Versions graded | Output tokens (thousands) | |||
|---|---|---|---|---|
| Model | Median | Mean | Median | Mean |
| Claude Opus 5 | 4.0 | 69.3 | ||
| GPT-5.6-sol | 2.0 | 22.1 | ||
| DeepSeek-V4-Flash | 2.0 | 54.5 | ||
| Claude Haiku 4.5 | 2.0 | 30.2 | ||
| Group | Practice | Required evidence |
|---|---|---|
| Improvement | Check inherited results | Before any edit, rerun the inherited solution to reproduce its recorded score or independently derive its central claim. |
| Use the task specification | Justify the submitted change from a task-document fact or metric property; a parameter-sweep result alone is insufficient. | |
| Improve robustness | Address failures such as degenerate inputs, timing limits, or fallback behavior without targeting the metric. | |
| Exploration | Bound potential improvement | Before searching, compute an upper or lower performance bound and decompose the remaining gap to allocate the budget. |
| Test beyond development data | Evaluate a change beyond the supplied development instances using a stress test or test of invariance. | |
| Compare with error bars | Use an explicit error bar when accepting or rejecting a change; a before-and-after comparison alone is insufficient. |
| Claude Opus 5 | GPT-5.6-sol | |||
|---|---|---|---|---|
| Agent judge | Mean | Mean | ||
| Claude Opus 5 (Claude Code) | 4.37 | 180 | 4.25 | 30 |
| GPT-5.6-sol (Codex) | 4.40 | 180 | 4.27 | 30 |
| Gemini 3.7 Flash (Claude Code) | 4.88 | 180 | 4.74 | 30 |
| Claude Opus 5 | GPT-5.6-sol | |||
|---|---|---|---|---|
| Solver model | Median | Median | ||
| Claude Opus 5 | 0.800 | 120 | 0.784 | 30 |
| GPT-5.6-sol | 0.761 | 120 | 0.736 | 30 |
| DeepSeek-V4-Flash | 0.647 | 120 | 0.564 | 30 |
| Claude Haiku 4.5 | 0.332 | 120 | 0.363 | 30 |