Recursive self-improvement requires carrying accepted changes into later improvement cycles, while studying agent-proposed changes also requires substantial research infrastructure. Existing settings often leave agents to rebuild routine infrastructure or restrict exploration to individual components. We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS). RSIGym exposes training, inference, rollout, evaluation, and sandbox execution through reusable services, with shared budget and permission controls supporting Data, Harness, and Joint improvement tracks. This design enables agents to investigate individual interventions and jointly optimize data, training settings, and execution harnesses within the same environment. We define RSI-Index as the mean fraction of the remaining performance gap closed across five benchmarks covering software engineering, terminal interaction, mathematics, scientific reasoning, and skill-based tasks. Comparing six frontier research models in independent Joint runs, Opus 5 achieves the highest RSI-Index of 0.4809 under a $500 platform-service budget per benchmark run. Its selected systems improve all five benchmarks, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. Additional experiments examine DSH-harness refinement, budget variation, and restricted network access, while recorded trajectories reveal how agents diagnose failures and select candidates. We open-source the full RSIGym codebase and results to support reproducibility and further research.
Figures & tables
Figure 1: Overview of RSIGym . Inside a lightweight container, the research agent iterates a research loop: it generates data, trains a checkpoint, evaluates the checkpoint with its harness, and analyzes the results to revise its data, training configuration, or harness. Colored tags mark the services each stage uses. The services themselves, shown at the bottom, run on external providers. All five share one authorization service, which checks every call against the run’s platform key and tracks the run’s budget (Sections 3.3 and 3.4 ). When the run ends, the agent submits its final system, which the same evaluation service scores under a fixed protocol.
Benchmark / Environment
RSI Exploration
System Design
Model Weights
Harness Design
Algorithm Hyperparameters
MLGym a ( Nathani et al., 2025 )
✓
✗
✓
Gym-style container environments
PostTrainBench ( Rank et al., 2026 )
✓
✗
✓
One H100 GPU; agents build the training pipeline
Agent 2 RL-Bench ( Chen et al., 2026 )
✓
✗
✓
Isolated workspaces with a grading API
AI4AI-Bench a ( Chi et al., 2026 )
✓
✗
✓
Code submission, clean reruns, and separate evaluation
AutoHarnessBench b ( Sleiman et al., 2026 )
✗
✓
✗
Harness search with public feedback and held-out evaluation
Table 1: Exploration targets and system designs of benchmarks and environments for AI self-improvement. A checkmark denotes an explicitly supported exploration direction; a cross denotes a direction outside the main compared setting. System design describes how experimentation and evaluation are organized.
Service
Agent submits
Service returns
Backend
Training
Training configuration, dataset, optional loss function
Per-step metrics, checkpoint identifier
Tinker
Inference
Chat completion request to a base model or checkpoint
Completion with tool calls
Tinker
Rollout
Chat completion request to a frontier model
Completion with tool calls
LiteLLM
Evaluation
Job configuration, optional harness archive
Scores, per-task rewards, trajectories, logs
Harbor
Sandbox
Sandbox and image requests via the E2B SDK
Isolated sandboxes, custom images
E2B
Table 2: The five services of RSIGym . Each service accepts a request from the research agent, performs the work on an underlying provider, and returns the results and artifacts listed here.
Track
Research agent
Target model
Initial harness
Benchmarks
Budget
Runs
Joint
all six
Qwen3.5-35B-A3B-Base
minimal
all five
$500
30
Joint
Opus 5, Astra
Qwen3.5-35B-A3B-Base
minimal
SWE, TB2
$1,000
4
Data
Opus 5
Qwen3.5-35B-A3B-Base
minimal
SWE, TB2
$500
2
Harness
Opus 5
Qwen3.6-35B-A3B-Instruct
minimal
SWE, TB2
$500
2
Data, no network
Opus 5
Qwen3.5-35B-A3B-Base
minimal
TB2
$500
1
Harness
Opus 5
Qwen3.6-35B-A3B-Instruct
DSH v0.1.1-rc.1
TB2
$500
1
Table 3: Runs in this study. Each run is an independent search that submits one system. Budget is the per-run allocation for platform services. The first four rows are reported in this section; the last three are the case studies of Section 5 .
Benchmark
Tasks
Trials
Scoring and configuration
SWE-bench Verified
100
300
Fixed seed-23 sample; held-out repository tests
Terminal-Bench 2.0
89
267
All configured tasks; task verifiers
AIME 2024/2025
60
180
Correct final integer answer
GPQA Diamond
100
300
Fixed seed-23 sample; correct option
SkillsBench
37
111
Science/office/finance subset; native task reward
Table 4: Evaluation sets. Task counts are the configured subsets; every final job uses three attempts per task.
Figure 2: Joint track at $500 per run: all six research agents improve Qwen3.5-35B-A3B-Base, and no agent is best everywhere. Each cell gives the official score and, as a bar and a percentage, the share of the gap to a perfect score that is closed. Blue marks the best value in a column and red a score below the initial system. RSI-Index is the mean of the five shares. Spend and time are summed over an agent’s five independent runs; time includes final official evaluation.
Figure 3: Single tracks and a doubled budget. Each row joins two official scores: the initial and the submitted system (top), or the 500andthe1,000 run of the same agent (bottom). Open circles mark the start and filled circles the end; blue rows end higher, red rows lower. Rows are grouped by their starting system. Every run is independent; Δ is computed from unrounded scores.
Figure 4: Training comparisons within Joint runs. Each comparison uses the same harness and task set: a trained candidate against base weights, or a heavier against a lighter update. Each is one evaluation on a small set and indicates direction only. Stars mark submitted checkpoints. a Reconstructed initial-completion votes on a synthetic set. b Approximate two-attempt scores. c Recovered after the agent’s selection.
Figure 5: Research time and normalized gain of every 500Jointrun.Eachbarisonerun’ssession,fromtheagent’sfirsttoitslaststep;theofficialevaluationisexcluded.Thenumberistherun’snormalizedgain,inblueforthebestagentonthebenchmarkandredbelowtheinitialsystem;SWEusesthreedecimalstoseparate0.425from0.433.Rowsaregroupedbymodelfamily:Claude,DeepSeek(bothinClaudeCode),andGPT(inCodex).Thethreeright−handcolumnssummarizeeachagent’sfiveruns:medianminutesuntilitsfirsttrainingrun(Start),mediannumberofrecordsinitssubmittedtrainingset(Data),andthepercentageofthe2,500 allocation spent (Spend).
Figure 6: Three case studies. Blue is above the reference shown in grey, red below; stars mark the submitted system. (a) The Data run without network access: seven checkpoints evaluated on the agent’s own 89-task, one-attempt development evaluation. The grey line shows the initial system’s official avg@3 score (0.1049), a reference from a different evaluation protocol; v2–v7 merge the reasoning into the message content. (b) Harness research starting from DSH, with Qwen3.6-35B-A3B-Instruct fixed; bands join matching categories, not individual tasks. (c) Qwen3.8-27B as both research agent and target. Left: six-task candidate comparison against the initial weights (grey line); 472 and 856 are training-example counts, and low-r. denotes reduced reasoning with the 472-example checkpoint. Right: the selected system on all 37 tasks, with two attempts per task in development and three in official evaluation; the grey line is the initial official score, and −0.0040 is the official score change.
Environments are the foundation of recursive self-improvement: they provide the problems agents work on and the feedback used to evaluate progress. Yet constructing challenging research environments with reliable evaluation still depends on domain experts, limiting their scale and disciplinary coverage. We introduce RSI-Forge, a multi-agent pipeline that turns published papers into executable environments for self-improvement. Three agents coordinate construction, reproduction, and review to produce tasks with automated evaluators; each paper's method is independently reimplemented to establish a baseline score. We present 210 environments across 18 fields, including 90 reviewed by independent human domain experts. Both experts and agent judges give high ratings to the potential for improving the provided starting solutions and the evaluators' ability to distinguish solution quality, whereas experts are more critical of shortcut resistance, faithfulness to the source paper, and whether a single idea can exhaust a task. To validate their use for repeated improvement, we evaluate four models over 3 successive attempts on 120 environments, with each attempt inheriting prior code and notes while model weights remain fixed. At least one model improves after the first attempt in 84% of environments. Models also outperform the reproduced paper methods in 68 of the 120 environments, demonstrating room for gains beyond these baselines. Transcript analysis identifies work beyond parameter tuning in 95% of these successful attempts. Analysis of the resulting trajectories shows that models scoring lower on these tasks explore less, more often accept gains smaller than the reported standard error, and rely more heavily on tuning to the development set. RSI-Forge provides a scalable approach to constructing research environments for training and evaluating self-improving agents.
Renxiong Wang, Darvin Yi, Abril Herrlein +16
Scale AI · University of California, Santa Cruz · University of North Carolina at Chapel Hill
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.
Peng Xia, Rujun Han, Zifeng Wang +11
Google Cloud AI Research · UNC-Chapel Hill · Stanford University +1
Digital agents must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by pretrained models. We introduce \textbf{RSIAgent}, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction. RSIAgent coordinates curriculum, actor, and verifier agents to continually explore the environment, validate outcomes, and retain environment-specific knowledge, including reusable causal relationships between actions, conditions, and consequences. It further adopts a \textbf{broad-then-deep} exploration strategy, combining parallel broad recursive self-exploration for discovering diverse environment structures with focused deep self-exploration for uncovering hard cases, hidden constraints, boundary conditions, and previously unknown causal dependencies. The resulting memory is frozen and can be directly reused for downstream tasks without updating model parameters. Experiments on OSWorld-v2 and Agent's Last Exam show that RSIAgent substantially improves strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.
Sibo Zhu, Shicheng Fan, Xinyue Wang +3
Aether AI · University of California San Diego · ‡Work done during internship in Aether AI +1