Recursive self-improvement requires carrying accepted changes into later improvement cycles, while studying agent-proposed changes also requires substantial research infrastructure. Existing settings often leave agents to rebuild routine infrastructure or restrict exploration to individual components. We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS). RSIGym exposes training, inference, rollout, evaluation, and sandbox execution through reusable services, with shared budget and permission controls supporting Data, Harness, and Joint improvement tracks. This design enables agents to investigate individual interventions and jointly optimize data, training settings, and execution harnesses within the same environment. We define RSI-Index as the mean fraction of the remaining performance gap closed across five benchmarks covering software engineering, terminal interaction, mathematics, scientific reasoning, and skill-based tasks. Comparing six frontier research models in independent Joint runs, Opus 5 achieves the highest RSI-Index of 0.4809 under a $500 platform-service budget per benchmark run. Its selected systems improve all five benchmarks, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. Additional experiments examine DSH-harness refinement, budget variation, and restricted network access, while recorded trajectories reveal how agents diagnose failures and select candidates. We open-source the full RSIGym codebase and results to support reproducibility and further research.
Figures & tables
Figure 1: Overview of RSIGym . Inside a lightweight container, the research agent iterates a research loop: it generates data, trains a checkpoint, evaluates the checkpoint with its harness, and analyzes the results to revise its data, training configuration, or harness. Colored tags mark the services each stage uses. The services themselves, shown at the bottom, run on external providers. All five share one authorization service, which checks every call against the run’s platform key and tracks the run’s budget (Sections 3.3 and 3.4 ). When the run ends, the agent submits its final system, which the same evaluation service scores under a fixed protocol.
Benchmark / Environment
RSI Exploration
System Design
Model Weights
Harness Design
Algorithm Hyperparameters
MLGym a ( Nathani et al., 2025 )
✓
✗
✓
Gym-style container environments
PostTrainBench ( Rank et al., 2026 )
✓
✗
✓
One H100 GPU; agents build the training pipeline
Agent 2 RL-Bench ( Chen et al., 2026 )
✓
✗
✓
Isolated workspaces with a grading API
AI4AI-Bench a ( Chi et al., 2026 )
✓
✗
✓
Code submission, clean reruns, and separate evaluation
AutoHarnessBench b ( Sleiman et al., 2026 )
✗
✓
✗
Harness search with public feedback and held-out evaluation
Table 1: Exploration targets and system designs of benchmarks and environments for AI self-improvement. A checkmark denotes an explicitly supported exploration direction; a cross denotes a direction outside the main compared setting. System design describes how experimentation and evaluation are organized.
Service
Agent submits
Service returns
Backend
Training
Training configuration, dataset, optional loss function
Per-step metrics, checkpoint identifier
Tinker
Inference
Chat completion request to a base model or checkpoint
Completion with tool calls
Tinker
Rollout
Chat completion request to a frontier model
Completion with tool calls
LiteLLM
Evaluation
Job configuration, optional harness archive
Scores, per-task rewards, trajectories, logs
Harbor
Sandbox
Sandbox and image requests via the E2B SDK
Isolated sandboxes, custom images
E2B
Table 2: The five services of RSIGym . Each service accepts a request from the research agent, performs the work on an underlying provider, and returns the results and artifacts listed here.
Track
Research agent
Target model
Initial harness
Benchmarks
Budget
Runs
Joint
all six
Qwen3.5-35B-A3B-Base
minimal
all five
$500
30
Joint
Opus 5, Astra
Qwen3.5-35B-A3B-Base
minimal
SWE, TB2
$1,000
4
Data
Opus 5
Qwen3.5-35B-A3B-Base
minimal
SWE, TB2
$500
2
Harness
Opus 5
Qwen3.6-35B-A3B-Instruct
minimal
SWE, TB2
$500
2
Data, no network
Opus 5
Qwen3.5-35B-A3B-Base
minimal
TB2
$500
1
Harness
Opus 5
Qwen3.6-35B-A3B-Instruct
DSH v0.1.1-rc.1
TB2
$500
1
Table 3: Runs in this study. Each run is an independent search that submits one system. Budget is the per-run allocation for platform services. The first four rows are reported in this section; the last three are the case studies of Section 5 .
Benchmark
Tasks
Trials
Scoring and configuration
SWE-bench Verified
100
300
Fixed seed-23 sample; held-out repository tests
Terminal-Bench 2.0
89
267
All configured tasks; task verifiers
AIME 2024/2025
60
180
Correct final integer answer
GPQA Diamond
100
300
Fixed seed-23 sample; correct option
SkillsBench
37
111
Science/office/finance subset; native task reward
Table 4: Evaluation sets. Task counts are the configured subsets; every final job uses three attempts per task.
Figure 2: Joint track at $500 per run: all six research agents improve Qwen3.5-35B-A3B-Base, and no agent is best everywhere. Each cell gives the official score and, as a bar and a percentage, the share of the gap to a perfect score that is closed. Blue marks the best value in a column and red a score below the initial system. RSI-Index is the mean of the five shares. Spend and time are summed over an agent’s five independent runs; time includes final official evaluation.
Figure 3: Single tracks and a doubled budget. Each row joins two official scores: the initial and the submitted system (top), or the 500andthe1,000 run of the same agent (bottom). Open circles mark the start and filled circles the end; blue rows end higher, red rows lower. Rows are grouped by their starting system. Every run is independent; Δ is computed from unrounded scores.
Figure 4: Training comparisons within Joint runs. Each comparison uses the same harness and task set: a trained candidate against base weights, or a heavier against a lighter update. Each is one evaluation on a small set and indicates direction only. Stars mark submitted checkpoints. a Reconstructed initial-completion votes on a synthetic set. b Approximate two-attempt scores. c Recovered after the agent’s selection.
Figure 5: Research time and normalized gain of every 500Jointrun.Eachbarisonerun’ssession,fromtheagent’sfirsttoitslaststep;theofficialevaluationisexcluded.Thenumberistherun’snormalizedgain,inblueforthebestagentonthebenchmarkandredbelowtheinitialsystem;SWEusesthreedecimalstoseparate0.425from0.433.Rowsaregroupedbymodelfamily:Claude,DeepSeek(bothinClaudeCode),andGPT(inCodex).Thethreeright−handcolumnssummarizeeachagent’sfiveruns:medianminutesuntilitsfirsttrainingrun(Start),mediannumberofrecordsinitssubmittedtrainingset(Data),andthepercentageofthe2,500 allocation spent (Spend).
Figure 6: Three case studies. Blue is above the reference shown in grey, red below; stars mark the submitted system. (a) The Data run without network access: seven checkpoints evaluated on the agent’s own 89-task, one-attempt development evaluation. The grey line shows the initial system’s official avg@3 score (0.1049), a reference from a different evaluation protocol; v2–v7 merge the reasoning into the message content. (b) Harness research starting from DSH, with Qwen3.6-35B-A3B-Instruct fixed; bands join matching categories, not individual tasks. (c) Qwen3.8-27B as both research agent and target. Left: six-task candidate comparison against the initial weights (grey line); 472 and 856 are training-example counts, and low-r. denotes reduced reasoning with the 472-example checkpoint. Right: the selected system on all 37 tasks, with two attempts per task in development and three in official evaluation; the grey line is the initial official score, and −0.0040 is the official score change.