Environments are the foundation of recursive self-improvement: they provide the problems agents work on and the feedback used to evaluate progress. Yet constructing challenging research environments with reliable evaluation still depends on domain experts, limiting their scale and disciplinary coverage. We introduce RSI-Forge, a multi-agent pipeline that turns published papers into executable environments for self-improvement. Three agents coordinate construction, reproduction, and review to produce tasks with automated evaluators; each paper's method is independently reimplemented to establish a baseline score. We present 210 environments across 18 fields, including 90 reviewed by independent human domain experts. Both experts and agent judges give high ratings to the potential for improving the provided starting solutions and the evaluators' ability to distinguish solution quality, whereas experts are more critical of shortcut resistance, faithfulness to the source paper, and whether a single idea can exhaust a task. To validate their use for repeated improvement, we evaluate four models over 3 successive attempts on 120 environments, with each attempt inheriting prior code and notes while model weights remain fixed. At least one model improves after the first attempt in 84% of environments. Models also outperform the reproduced paper methods in 68 of the 120 environments, demonstrating room for gains beyond these baselines. Transcript analysis identifies work beyond parameter tuning in 95% of these successful attempts. Analysis of the resulting trajectories shows that models scoring lower on these tasks explore less, more often accept gains smaller than the reported standard error, and rely more heavily on tuning to the development set. RSI-Forge provides a scalable approach to constructing research environments for training and evaluating self-improving agents.
Figures & tables
Figure 1 : Overview of RSI-Forge and the information available to each agent. Failed checks lead to repair or rejection before an environment is accepted. The output count includes all 210 environments. The performance bars report median scores of final solutions after 3 successive attempts on 120 environments sampled from the main construction run. Scores are normalized so that the starting solution is 0 and the target score is 1 ( Eq. 1 ).
Stage
Count
Source papers screened
10,800
Papers eligible for construction
2,117
Construction attempts begun
213
Passed Design
199
Passed Build
198
Passed Independent Reproduction
189
Table 1: Main-run construction funnel. Counts include environments that passed after repairs. The additional 30 environments built with GPT-5.6-sol are outside this funnel.
Field
Task
Evaluation objective
Robotics
Choose footholds on a terrain map
Increase distance in six steps
Life Sciences
Propose molecules for a binding pocket
Improve docking scores
Earth, Climate & Space
Estimate weather from partial observations
Reduce state-estimation error
Statistics & Inference
Estimate effects without a control series
Reduce treatment-effect error
Software Engineering & Formal Methods
Generate tests within a fixed budget
Trigger more distinct failures
Computational Physics
Infer a plasma boundary from magnetic data
Reduce boundary-estimation error
Table 2: Examples of environments from the main construction run.
Issue identified
Share of repair passages
Checks unable to detect the intended failure
66%
Incorrect or incomplete instructions
63%
Checks inspecting a different copy of the code
54%
Score gains without performing the intended work
37%
Table 3: Most common construction defects identified by model-based coding of repair passages. Percentages are shares of repair passages; categories overlap.
Figure 2 : Quality ratings from human experts and agent judges on the same 90 environments, with one expert and three judges per environment. Points show means on a 1–5 scale (higher is better). Agent ratings are averaged within environments; rows are ordered by expert mean. Only experts received available solver results. Sections B.2.2 and 6 give the rubric and per-dimension standard deviations.
Median normalized score ↑
Session 3
Model
Session 1
Session 2
Session 3
95% CI
Mean rank ↓
Claude Opus 5
0.754
0.791
0.800
[0.753, 0.867]
1.67
GPT-5.6-sol
0.673
0.747
0.761
[0.723, 0.824]
1.73
DeepSeek-V4-Flash
0.550
0.594
0.647
[0.586, 0.711]
2.97
Claude Haiku 4.5
0.129
0.277
0.332
[0.203, 0.451]
3.63
Table 4: Submitted-solution scores on 120 environments, normalized to the starting solution (0) and target (1). Session 3 confidence intervals use bootstrap resampling of environments. Mean rank averages within-environment ranks at session 3 (1 is best).
Figure 3 : Transcript-coded behavior across 360 sessions per model (120 environments, three sessions each). Cell values and color intensity show percentages of sessions assigned to each group by model-based open coding. Groups overlap; “improvement” covers developing and checking solutions, without requiring score gains.
Figure 4 : A three-session Claude Opus 5 chain on a foothold-selection task. Open points are saved solutions graded on the held-out test set; filled points mark final submissions, with normalized scores of 0.15, 0.33, and 0.58. The annotations summarize the changes recorded in each session. The third session adds a ground-support check and reverses the second session’s averaging rule.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Count
Field
Count
Robotics
14
Security & Cryptography
10
Life Sciences
13
Neuroscience
9
Earth, Climate & Space
13
Optimisation & Discrete Algorithms
9
Statistics & Inference
13
Chemistry & Materials
9
Networks & Information Th.
12
Graphics & Rendering
8
Software Eng. & Formal Methods
12
Quantitative Finance
7
Appendix
Table 5: Field composition of the 180 main-run environments.
Dimension
Experts
SD
Panel
SD
Gap
Research problem
4.39
0.84
4.59
0.32
+0.20
Room above starting solution
4.81
0.51
4.89
0.31
+0.08
Resistance to saturation
3.09
0.96
3.98
0.84
+0.89
Faithfulness to source paper
3.72
0.87
4.22
0.51
+0.50
Metric validity
4.40
0.85
4.49
0.45
+0.09
Single objective
4.61
0.90
4.66
0.46
+0.05
Appendix
Table 6: Human-expert and agent-panel ratings on the same 90 environments. Scores are means on a 1–5 scale; SD denotes standard deviation. Agent ratings are averaged within environments; Gap is the panel mean minus the expert mean.
Dimension
RSI-Exam
RSI-Forge
Gap
Score discrimination
2.64
4.65
+2.01
Development–test distinction
3.23
4.67
+1.44
Shortcut resistance
3.92
4.67
+0.75
Single objective
4.64
4.66
+0.03
Room above starting solution
4.87
4.89
+0.02
Research problem
4.61
4.59
−0.02
Appendix
Table 7: Ratings from the same three agent judges on 34 manually constructed RSI-Exam tasks and the 90 RSI-Forge environments in Table 6 . Gap is RSI-Forge minus RSI-Exam; the 8 dimensions exclude faithfulness because this comparison has no common source-paper reference. The task sets are unpaired, so the differences do not isolate the effect of automated construction.
Axis
A score of 1
A score of 3
A score of 5
Real problem
a toy with no research content
a real problem, narrowly posed
you would be pleased to see a student work on this
Headroom
the starter is already near the ceiling
a competent effort gains something
clearly large, and you can name where it comes from
Not saturable
one change takes most of the range
the first change is worth it, more remains
improvement requires several independent ideas
Faithful to the paper
unrelated to what the paper does
related, but the hard part has been removed
this is the paper’s problem, honestly posed
Metric validity
measures something else
a reasonable proxy
this is the quantity the field would report
Single direction
two objectives silently traded
one quantity, the constraint is loose
one quantity, one direction, the constraint binds
Appendix
Table 8 : Descriptions of scores 1, 3, and 5 in the environment-quality rubric, reproduced verbatim. The dimensions are grouped here as in the rating summaries; the human review form presents faithfulness to the source paper last.
Versions graded
Output tokens (thousands)
Model
Median
Mean
Median
Mean
Claude Opus 5
4.0
3.6±1.6
69.3
72.2±22.8
GPT-5.6-sol
2.0
2.7±1.7
22.1
22.6±8.3
DeepSeek-V4-Flash
2.0
2.1±1.1
54.5
53.8±16.3
Claude Haiku 4.5
2.0
2.7±1.9
30.2
29.8±8.8
Appendix
Table 9: Activity per session on the 120 main evaluation environments. Versions are saved runnable solutions that were graded; output tokens are in thousands. Means include one standard deviation. Token volume is not equivalent to monetary or computational cost across providers.
Figure 5 : Transcript-coded mechanisms as percentages of each model’s sessions that surpass the reproduced baseline. The codebook was derived from a sample and fixed before counting. Sessions may receive several labels; nine categories cover work beyond parameter tuning. The labels describe work within the constructed task and do not establish scientific novelty.
Figure 6 : Open-coded families grouped as improvement (21), exploration (27), and shortcuts (14). Cells show percentages of each model’s 360 sessions assigned a family. Labels overlap. Shading is scaled within each group; printed percentages support comparisons across groups. The ungrouped family for submitting the inherited solution unchanged is omitted.
Group
Practice
Required evidence
Improvement
Check inherited results
Before any edit, rerun the inherited solution to reproduce its recorded score or independently derive its central claim.
Use the task specification
Justify the submitted change from a task-document fact or metric property; a parameter-sweep result alone is insufficient.
Improve robustness
Address failures such as degenerate inputs, timing limits, or fallback behavior without targeting the metric.
Exploration
Bound potential improvement
Before searching, compute an upper or lower performance bound and decompose the remaining gap to allocate the budget.
Test beyond development data
Evaluate a change beyond the supplied development instances using a stress test or test of invariance.
Compare with error bars
Use an explicit error bar when accepting or rejecting a change; a before-and-after comparison alone is insufficient.
Appendix
Table 10: The ten fixed practices, distinct from the open-coded families in Fig. 6 and the successful-session mechanisms in Fig. 5 . Labels and operational criteria are paraphrased from the coding scheme; timing requirements and exclusions are retained.
Claude Opus 5
GPT-5.6-sol
Agent judge
Mean
n
Mean
n
Claude Opus 5 (Claude Code)
4.37
180
4.25
30
GPT-5.6-sol (Codex)
4.40
180
4.27
30
Gemini 3.7 Flash (Claude Code)
4.88
180
4.74
30
Appendix
Table 11: Agent-judge ratings by construction backend. Cells average over rubric dimensions, then environments; n counts rated environments. For this comparison, Gemini 3.7 Flash uses Claude Code; Section B.2 compares its ratings under Claude Code and Antigravity.
Claude Opus 5
GPT-5.6-sol
Solver model
Median
n
Median
n
Claude Opus 5
0.800
120
0.784
30
GPT-5.6-sol
0.761
120
0.736
30
DeepSeek-V4-Flash
0.647
120
0.564
30
Claude Haiku 4.5
0.332
120
0.363
30
Appendix
Table 12: Median final submitted normalized scores after three sessions by construction backend. The main-backend column repeats session 3 of Table 4 ; n counts chains with all three sessions scored. Construction sets are unpaired.
Recursive self-improvement requires carrying accepted changes into later improvement cycles, while studying agent-proposed changes also requires substantial research infrastructure. Existing settings often leave agents to rebuild routine infrastructure or restrict exploration to individual components. We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS). RSIGym exposes training, inference, rollout, evaluation, and sandbox execution through reusable services, with shared budget and permission controls supporting Data, Harness, and Joint improvement tracks. This design enables agents to investigate individual interventions and jointly optimize data, training settings, and execution harnesses within the same environment. We define RSI-Index as the mean fraction of the remaining performance gap closed across five benchmarks covering software engineering, terminal interaction, mathematics, scientific reasoning, and skill-based tasks. Comparing six frontier research models in independent Joint runs, Opus 5 achieves the highest RSI-Index of 0.4809 under a $500 platform-service budget per benchmark run. Its selected systems improve all five benchmarks, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. Additional experiments examine DSH-harness refinement, budget variation, and restricted network access, while recorded trajectories reveal how agents diagnose failures and select candidates. We open-source the full RSIGym codebase and results to support reproducibility and further research.
Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by pretrained models. We introduce \textbf{RSIAgent}, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction. RSIAgent coordinates curriculum, actor, and verifier agents to continually explore the environment, validate outcomes, and retain environment-specific knowledge, including reusable causal relationships between actions, conditions, and consequences. It further adopts a \textbf{broad-then-deep} exploration strategy, combining parallel broad recursive self-exploration for discovering diverse environment structures with focused deep self-exploration for uncovering hard cases, hidden constraints, boundary conditions, and previously unknown causal dependencies. The resulting memory is frozen and can be directly reused for downstream tasks without updating model parameters. Experiments on OSWorld-v2 and Agent's Last Exam show that RSIAgent substantially improves strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.
Sibo Zhu, Shicheng Fan, Xinyue Wang +3
Aether AI · University of California San Diego · ‡Work done during internship in Aether AI +1
Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enables regularized experimental actions and maintains persistent, traceable experimental records, and Reviewer-Guided Research Orchestration, which organizes Workers and Reviewers in a dynamically growing research DAG. Workers explore diverse research directions and Reviewers compare evidence across related experiments for subsequent explorations. On PostTrainBench with Qwen3-4B-Base, it averages 54.49 versus 46.53 for the strongest agent baseline, with a 0.0% hacking rate. Scaling to 35B model, RSI-Master surpasses the human-developed Instruct model on LiveCodeBench-v6 (41.21 vs. 37.36) and SciCode, and reaches a nonzero score on HorizonMath, a benchmark of unsolved research problems on which most frontier models score near zero.
Yaxin Du, Xiyuan Yang, Zhifan Zhou +10
Shanghai Jiao Tong University · Carnegie Mellon University · University of Waterloo